10/09 2026
538


From Which Model Reigns Supreme to Which One Fits Best
Image Source | Internet (Please contact for removal if infringing) Partially AI-generated
Over the past two years, the most frequently asked question in Silicon Valley's AI circle has been: Which model is the strongest?
But in 2026, this question is being replaced by a more pragmatic one: Is it necessary to use the most expensive model for the same task?
In June of this year, ride-hailing giant Uber revealed a figure that silenced the entire industry: In just the first four months of 2026, the company had exhausted its entire annual AI budget, simply because employees were heavily using AI coding tools.
Amazon immediately halted its internal AI usage leaderboard because employees were deliberately performing unnecessary tasks to boost rankings and increase token consumption. Microsoft, on the other hand, planned to discontinue Claude Code subscriptions for employees in several key product departments.
When even Amazon is starting to worry about token costs, expense management is no longer just a concern for startups.
Consulting firm Bain released a set of analytical data: After enterprise AI cumulative spending surpassed $1 trillion, the actual cost savings brought by AI have generally fallen far short of expectations. 44% of large enterprises are using "unrealized cost savings from the previous round of AI" to justify the next round of AI investments. Bain characterized this practice as "a cyclical bet with structural vulnerabilities."
Under pressure, corporate choices are beginning to change.

From "Strongest" to "Most Cost-Effective"
Data from Citibank shows that the cost per task for open-source models has fallen by 35% weekly to $0.80, with discounts relative to closed models expanding from 40% to 60%.
Forkast's statistics are even more straightforward: The average cost of open-weight models is $0.83 per million tokens, while proprietary alternatives cost $6.03—7.3 times more expensive, saving developers 86% in expenses.
In an extreme comparison, DeepSeek V4-Pro averaged just $0.28 per question in cybersecurity task tests, while closed-source competitor Opus 4.5 cost a staggering $12.50—a gap exceeding 40-fold.
The most dramatic case occurred in September. Xiaomi and SpaceX released their new models on the same day: MiMo-V2.6 series (open source) and Grok 4.7 (closed source).
Both models achieved the same intelligence index score, but MiMo-V2.6-Pro cost just $0.13 per task, while Grok 4.7 cost $3.74—a nearly 29-fold difference.
This is the increasingly clear bill facing Silicon Valley companies.
According to the Financial Times, AlphaSense data shows that in August and September this year, mentions of "open-weight" or "open-source" models by executives in U.S. corporate earnings calls and investor meetings surged sixfold year-on-year.
A more critical signal comes from the usage side. In August, open-weight models accounted for 56% of total token processing through Vercel AI Gateway, up from just 7% in December last year.
The proportion of tokens processed by open-source models on OpenRouter also rose from 34% in January to 65% in June this year.
Telecom giant AT&T is leading the way. Andy Marcus, its Chief Data and AI Officer, revealed that about 40% of AI workloads now run on open-source models, with plans to increase this to 70% within a year.
Facing 45 billion tokens processed daily, AT&T also fine-tunes open-source models with proprietary data to match or even exceed the performance of closed models on specific tasks.
Match Group's Tinder is also worth watching. Its CTO revealed that the company's annualized AI spending skyrocketed from $1 million in January to $10 million in July. To control costs, Tinder has started routing some non-technical queries to open-source models.
Even Chinese models have gained an unexpected entry ticket in this cost war.
According to Ramp's June 2026 software trend rankings, DeepSeek topped the "Trending" list for breakthrough growth, a first for a Chinese model company. Past regulars on this list included Silicon Valley firms like Figma and Fireworks AI.


"Model Assortment": A Simple Business Logic
What truly reshapes AI business models is not an open-source model defeating a closed-source giant but companies abandoning the idea of "one model for all problems."
Tinder's approach is representative: The most expensive cutting-edge models handle complex tasks, while cheaper open-source models take on large volumes of routine work.
High-frequency but low-complexity tasks like information extraction, text classification, simple code modifications, and customer service responses do not require the top-tier models.
The rise of this "multi-model routing" strategy has increased the average number of models used by enterprises from 2.1 in Q1 2025 to 4.7 in Q1 2026.
Companies no longer stick with a single model provider but dynamically allocate different models based on task difficulty, latency requirements, and cost budgets.
Databricks executive David Meyer's assessment confirms this trend: Reinforcement learning-optimized small open-source models can complete specific tasks at lower training costs. The focus of enterprise AI competition is shifting from model scale to actual operational efficiency and unit costs.
In other words, the AI industry is transitioning from "selling the smartest models" to "selling the most suitable computing power."
However, this cost-driven shift has triggered a rare public split in Silicon Valley.
On July 24, NVIDIA CEO Jensen Huang posted on social platform X, stating, "The world needs both cutting-edge closed models and cutting-edge open models." Microsoft CEO Satya Nadella responded just nine minutes later.
Subsequently, Meta, IBM, Hugging Face, a16z, and others joined the chorus. The number of signatories to the open letter, titled "Open Weights and U.S. Leadership in AI," quickly expanded from 33 to 133.
The letter's core appeal is to urge policymakers to avoid prematurely restricting or banning open-weight AI models.
Notably, OpenAI and Google eventually signed the letter under pressure, while Anthropic was absent from start to finish. When Huang further established the "Open Secure AI Alliance," which requires substantial financial investment, OpenAI, Google, and Anthropic all remained unresponsive.
The reason is simple: Their interest calculations differ sharply.
As one angel investor analyzed: The closed-source camp's primary goal is to maintain technological monopoly dividends. OpenAI and Anthropic rely on hundred-billion-dollar-scale computing investments to build top-tier models, with business models entirely dependent on paid access to exclusive models. Only by sustaining technological barriers and closed systems can they preserve high premiums and market pricing power.
The open-source camp, however, is built on industrial expansion logic. Companies like NVIDIA and Meta, which provide infrastructure and ecosystems, do not profit directly from models themselves. Open sourcing significantly lowers AI adoption barriers, instead driving increased computing consumption and cloud service demand.
Huang's strong support for open source essentially aims to maintain the prosperity of NVIDIA's GPU ecosystem. The more open models become, the more developers there are, and the greater the computing demand.


How Close Is Open Source Catching Up?
If open-source models were merely "cheap but underperforming," this shift would not occur. The crucial point is that the performance gap is rapidly narrowing.
Mozilla's latest report shows that the performance gap between Chinese open-source models and closed-source cutting-edge models has shrunk to about 4.4 months.
The API fees companies pay to access top closed-source models typically amount to about five times the deployment cost of equivalent open-weight models, yielding only about a four-month technological lead.
Data from Stanford HAI's 2026 AI Index report is even more striking: As of March 2026, the Elo score gap among the top four closed-source labs—Anthropic, xAI, Google, and OpenAI—on the Arena Leaderboard had narrowed to within 25 points.
When performance convergence becomes reality, models themselves can no longer serve as lasting moats.
In some niche scenarios, open-source models have even begun to surpass closed-source counterparts.
On the Claw-Eval benchmark, Xiaomi's open-weight model MiMo V2.5 Pro outranked GPT-5.4, Meta's Muse Spark, and Gemini 3.1 Pro.
GLM-5.2's self-reported score (62.1) on SWE-bench Pro also exceeded GPT-5.5's 58.6, at about one-sixth the cost.
LeCun praised domestic open-source models on social media, Cursor admitted its proprietary model was built on Kimi K2.5, and Cognition's SWE-1.6 was revealed to be post-trained on GLM models...
These cases collectively illustrate a fact: Open-source models are no longer "alternatives" but the actual foundation for many Silicon Valley products.
Silicon Valley legal AI firm Harvey took an even more direct approach: Its first internal model, Harvey Tenet, was post-trained on Kimi K3. Remarkably, Harvey is a startup backed by OpenAI.
The Commoditization of Cutting-Edge AI Has Just Begun
The deeper logic behind this transformation is that the AI industry is undergoing a shift from "scarce commodities" to "layered commodities."
In September 2026, Silicon Data's LLM Token Spending Index fell below $1 per million tokens for the first time since tracking over 200 models.
Forkast's commentary was blunt: The era of intelligence as a scarce, high-margin commodity has effectively ended.
Meanwhile, prices for cutting-edge models are heading in the opposite direction.
When GPT-6 Astra and Claude Fable 5.1 were released, input prices were set at $10 per million tokens, with output prices doubling to $50.
The industry is "layering": At the bottom lies commoditized, low-cost inference; at the top, expensive autonomous agent capabilities.
As Forkast's analysis points out, cutting-edge labs are no longer trying to offer the cheapest computing but competing for the highest-value autonomous tasks—shifting from "selling tokens" to "selling workflow integration."
This means "open source vs. closed source" is not a binary choice but a strategic question of layered configuration.
For most companies, the answer likely lies neither in fully embracing open source nor clinging to closed-source flagships but, like AT&T, saving money where possible while spending where top-tier capabilities are needed.
Airbnb CEO Brian Chesky shook Silicon Valley months ago by saying, "We rely heavily on Alibaba's Tongyi Qianwen model—it's better and cheaper than OpenAI."
He added that while the company also uses OpenAI's latest models, they are rarely deployed at scale in production due to faster, more economical alternatives.
Chesky and OpenAI CEO Sam Altman are close friends, but when it comes to product integration, "even brothers keep separate accounts."
This perhaps best captures Silicon Valley AI in 2026: Sentiment is sentiment, but the bill is the bill.