Has the "Efficiency-Centric Approach" of Large Models Triumphed?

08/31 2026 444

Amidst the fierce competition among large model firms centered on computational prowess and parameter counts, MiniMax has charted a uniquely distinct course.

On August 26, MiniMax unveiled its inaugural half-year report since going public, revealing $116.6 million in revenue for the first half of the year, marking a 283.1% year-over-year surge that has already eclipsed the total for 2025. Concurrently, the gross margin has climbed from 12.1% to 17.9%.

Nearly tripling revenue while also enhancing gross margins is an uncommon feat in the large model sector. During the earnings call, founder Yan Junjie attributed this growth to a singular principle: "Minimize Inference Costs, Maximize Intelligence"—a strategy focused on reducing inference expenses while amplifying intelligent capabilities.

More significantly, MiniMax's growth engine is undergoing a transformation. Previously recognized primarily for its C-end offerings like Talkie and Hailuo AI, the company now sees B-end revenue accounting for 63.4% of first-half revenue, with this figure climbing to around 80% in August.

From the ascent of the B-end, driven by enhanced inference efficiency and cost improvements, to the progression of the H3 open-source model and its subsequent iterations, MiniMax is pioneering a path distinct from mere computational power accumulation for growth. It also demonstrates that AI firms can craft superior models at reduced costs, catering to a broader spectrum of individuals, developers, and enterprise users.

I. The Ascendancy of the B-end: MiniMax's Growth Engine Transforms

In recent times, when MiniMax is mentioned, outsiders often first think of its C-end products such as Talkie and Hailuo AI.

Talkie, in particular, was once a cornerstone of MiniMax's overseas market strategy, positioning the company as a benchmark for domestic large model firms venturing into international markets. In the first half of 2025, AI-native products generated $21.22 million in revenue, contributing 69.7% of the company's total.

Beginning in the first half of this year, MiniMax is transitioning from a large model company heavily reliant on C-end revenue to a platform-based entity driven by both C-end and B-end operations.

Financial reports indicate that in the first half of the year, revenue from open platforms and other AI enterprise services reached $73.93 million, a 703.1% year-over-year increase, with its share of total revenue rising from 32.8% for the entirety of 2025 to 63.4%.

Entering the third quarter, this trend is set to continue. Management disclosed during the earnings call that as of August, the company's ARR (Annual Recurring Revenue) had surpassed $800 million, with the ToB business contributing over 80%.

ARR statistical methodologies vary among AI companies, and the industry has yet to establish a unified standard for ARR disclosure. Most companies tend to annualize single-day peak revenue for external PR purposes, but this approach can significantly amplify short-term fluctuations if annualized directly from single-day peaks.

According to MiniMax's disclosure, its August ARR is annualized based on the average daily revenue of the most recent week, rather than simply adopting a single-day peak. This weekly average approach mitigates daily fluctuations to some extent, making it more suitable for gauging the current operational pace of AI companies.

In terms of new revenue streams, this shift is even more pronounced. Compared to the first half of 2025, MiniMax generated approximately $86.14 million in new revenue in the first half of this year, with about $64.72 million originating from open platforms and enterprise services. In essence, about three-quarters of the new revenue came from the B-end.

More importantly, this round of B-end growth has not been accompanied by a significant uptick in sales investment. In the first half of the year, MiniMax's sales and distribution expenses decreased by about 18% year-over-year, while enterprise service revenue grew more than sevenfold. Therefore, MiniMax's growth is primarily driven by enhancements in model capabilities and product competitiveness.

The rise of the B-end does not imply stagnation in the C-end.

In the first half of the year, AI-native product revenue reached $42.64 million, a 100.9% year-over-year increase, outpacing the 82% growth in the fourth quarter of 2025. The financial report attributed this growth to increased user engagement, stronger willingness to pay, and the continued commercialization of products like Hailuo AI, with video generation being a major catalyst.

Thus, MiniMax has established two business lines driven by both C-end and B-end operations.

On one hand, C-end products like Talkie and Hailuo AI directly engage users, generating revenue while also serving as crucial scenarios for model capabilities to be showcased. Especially for multimodal products like video generation, user feedback on effectiveness directly informs product iteration, aiding MiniMax in refining its models and applications.

On the other hand, the B-end involves open platforms that extend model capabilities beyond the company's own products and into the applications of more developers and enterprises. The more these models are utilized by enterprises and developers, the greater the API calls and token consumption, leading to growth in B-end revenue.

As the primary revenue driver shifts from the C-end to the B-end, MiniMax has evolved from an AI application company reliant on hit products to a platform-based company centered on model capabilities.

The former's growth ceiling is contingent on the lifecycle of hit products. The latter embeds model capabilities into enterprise production processes, offering stronger sustainability and scale effects. The former is valued as a product company, while the latter is valued as an infrastructure company. The process of the market revaluing MiniMax has likely just commenced.

II. From Inference Efficiency to Gross Margin Improvement: Technical Routes Begin to Yield Results

The true value of this growth also hinges on cost changes behind the revenue increase.

In the first half of 2026, the gross margin was 17.9%, up from 12.1% in the same period last year, an increase of 5.8 percentage points. Gross profit reached $20.81 million, a 464.8% year-over-year increase.

Generally, revenue growth for large model companies is often accompanied by higher computational consumption. The more models are utilized, the higher the theoretical inference costs, meaning that scaling up does not necessarily lead to improved economics.

However, MiniMax presents a different scenario this time.

In the first half of the year, MiniMax's cost of sales was $95.76 million, a 258.1% year-over-year increase, lower than the 283.1% revenue growth rate. As revenue grew rapidly, costs did not rise at the same pace, and the cost per unit of revenue is declining.

During the earnings call, MiniMax attributed cost efficiency and gross margin improvements to the same factor: the ability to provide the same or even superior model capabilities at reduced costs.

Over the past two months, the throughput per unit of computational power for text models has increased by about threefold. The goal for the next-generation M3.1 is to reduce inference costs to about one-third of those when M3 was first launched.

This also elucidates the source of MiniMax's gross margin improvement.

From the cost structure of large model inference, the final cost is not solely determined by chip prices. Factors such as the computational requirements of the model itself, how many requests a single chip can handle, how many tokens can be generated per second, and the utilization efficiency of the entire cluster directly impact the cost per token.

Therefore, MiniMax is simultaneously enhancing efficiency across models, inference systems, and infrastructure. This represents a shift in the current competition among large models.

In the past, the industry focused more on model capabilities themselves, such as parameter scale and Benchmark scores. However, as model capabilities gradually converge, obtaining the same level of intelligence at what cost is becoming another variable determining commercialization capabilities.

Cost reductions also provide MiniMax with greater operational flexibility. Part of the savings can be passed on to customers through price reductions, lowering the barrier to model usage and attracting more users and enterprises, thereby driving token volume growth. Another part can be used to improve gross margins.

MiniMax has demonstrated that for large model companies, price reductions and gross margin improvements do not necessarily conflict. The key lies in whether the cost per token can decline faster than the price reduction. As long as unit economics continue to improve, greater usage scales can translate into higher gross profits.

This also forms a new growth logic. Improved model capabilities lead to more usage, which increases the utilization efficiency of computational power and infrastructure. Lower unit costs, in turn, provide room for further price reductions, expanded usage scales, or improved gross margins.

III. After the Surge, What Will Sustain Growth?

The rise of the B-end and improved inference efficiency explain the explosion in MiniMax's business. However, the market is more concerned with whether MiniMax can sustain its growth momentum after a strong half-year performance.

We have gleaned some key insights from the company's new product line deployment and the construction of multimodal barriers.

At the end of July this year, MiniMax released the open-source model MiniMax H3. In the past, advanced video generation models were mostly controlled by leading tech companies and provided as closed products or APIs. The open-sourcing of H3 allows MiniMax to directly place video generation models in the hands of developers, giving the company a new ecological entry point in the multimodal field.

On one hand, it validates MiniMax's previous technical accumulations in multimodal capabilities. MiniMax has long advanced language, voice, image, and video capabilities simultaneously, with Hailuo AI providing substantial real user feedback. H3 further attempts to extend language model capabilities into visual generation. If this route proves successful, the advantages of language models in understanding and reasoning could further raise the upper limits of visual generation.

On the other hand, H3 explores a new type of commercial relationship where open-source models, C-end products, and APIs do not necessarily conflict but can instead form synergies.

Developers can start with the open-source model to access MiniMax's technology before moving on to APIs and commercial services. Ordinary users can directly use the model through products like Hailuo AI. For MiniMax, open-sourcing does not mean abandoning commercialization but rather exchanging a larger developer ecosystem for more model usage and feedback.

This is also the value of H3's open-sourcing at this juncture.

When video generation is still in a stage of rapid iteration, with model effectiveness, costs, and application scenarios not yet fully defined, obtaining developer feedback early on represents a competitive advantage. From a commercialization perspective, multimodal models also have the opportunity to enter more complete content production processes in advertising, e-commerce, gaming, design, film, and television.

At the same time, MiniMax is accelerating iterations of its language models, with M3.1 and M3 Pro progressing as planned.

Among them, M3.1 is more directly related to the efficiency issues mentioned earlier. If inference costs can continue to decline while model capabilities keep improving, it will help expand B-end usage scales and further improve unit economics.

M3 Pro corresponds to higher intelligence levels and more complex task demands, expected to meet the rising requirements of enterprise clients for model capabilities. H3.1 continues to extend into video and multimodal capabilities.

Competition in the large model industry can no longer rely on a single hit product to establish long-term barriers. Model capabilities are iterating faster, prices are continuously declining, and today's leading capabilities may quickly become industry standards. The true barrier lies in the ability to sustain a cycle of model R&D—product implementation—user usage—feedback iteration.

Currently, MiniMax has successfully established this path, forming a mutually reinforcing business cycle from model R&D to C-end products and B-end services, and then to feedback from real users and developers. As M3.1, M3 Pro, and H3.1 continue to advance, this capability is expected to keep releasing value.

Solemnly declare: the copyright of this article belongs to the original author. The reprinted article is only for the purpose of spreading more information. If the author's information is marked incorrectly, please contact us immediately to modify or delete it. Thank you.