Wu Yongming's First Task: Finding 'Jobs' for 20GW of Computing Power

09/24 2026 359

AI Competition Shifts to Computing Utilization Rates

Author|Hua Wen

Editor|Xiaobai

Illustrations|AI-Generated

Produced by|Next Emphasis

At the 2024 Cloud Town Conference on September 22, Alibaba Group CEO Wu Yongming set a series of targets: By 2032, Alibaba Cloud's global data center capacity will exceed 20GW. T-Head's next-gen AI chip Zhenwu V900 will deliver 3x the computing power of the previous-gen M890, with mass production slated for Q1 2027. The parameter scale of subsequent Qwen models will expand to 5-10 trillion.

Parameter scale determines what models can do, while cluster scale determines Alibaba's supply capacity. Together, these factors determine who pays for electricity and depreciation.

Goldman Sachs and other institutions predict Alibaba Cloud's current computing power at approximately 4-6GW. Bloomberg reports Microsoft plans to expand its data center capacity to 38GW by 2032, while AWS previously disclosed adding 3.8GW annually.

Alibaba aims to quadruple its computing power within six years, requiring annual growth rates close to those of the world's top two cloud providers. Wu acknowledged in the same speech that global supply chain shortages for AI data centers continue to constrain computing power growth.

Scale tells only half the story. After chips are purchased, data centers built, and power and networks connected, depreciation, maintenance, and energy costs accrue by the second—regardless of whether computing power is utilized. The key metric is how much of the time these racks operate for external clients.

In his speech, Wu compared future machine intelligence to electricity—available at the flick of a switch. This analogy highlights the other side of the business: A power grid's value depends not on its technological sophistication, but on how many people pay to use it.

Alibaba is not alone in this assessment. One week earlier, Huawei's Full Connect Conference unveiled the Ascend 960 supernode, with a highly consistent technical approach: Given process node limitations, instead of chasing single-chip performance, cluster performance is enhanced through interconnect protocols and system-level innovations to create "one computer."

The divergence lies in monetization strategies. Huawei explicitly focuses on "hardware monetization," treating the supernode as a product for customers to purchase—utilization rates become the customer's concern. T-Head does not pursue standalone chip profits; computing power remains Alibaba's owned asset, with utilization rates directly impacting its own profit margins.

20GW Written Into Cash Flow First

In February 2025, Alibaba announced plans to invest at least RMB 380 billion over three years in AI and cloud infrastructure. By Q1 FY2027 (ending June 30, 2026), approximately RMB 190 billion had already been invested—half the total budget with less than half the time elapsed.

The RMB 380 billion likely represents just the initial commitment. UBS calculates that each additional 1GW of AI data center capacity requires roughly RMB 100 billion in IT equipment CAPEX alone. Expanding from current levels to 20GW implies trillion-level investments.

Investment banks have back-calculated revenue targets from this goal. Citigroup projects 20GW will correspond to external cloud revenue of approximately $160 billion by FY2033, while UBS estimates $170 billion by 2032—both implying ~40% CAGR. Alibaba's current annualized quarterly cloud revenue stands at ~$27 billion, meaning cloud revenue must nearly sextuple within six years.

Quarterly data shows steeper growth. Alibaba's Q2 2024 CAPEX surged 75% YoY to RMB 67.678 billion, far exceeding market expectations of RMB 36 billion. During the same quarter, AI cloud and computing service revenue reached RMB 48.437 billion, with CAPEX at 1.4x this figure.

Cash flow pressures are already visible. Free cash flow outflow reached RMB 44.67 billion this quarter, compared to RMB 18.815 billion inflow YoY. Group operating profit plunged 57% YoY to RMB 15.161 billion.

In August, Alibaba completed a HKD 80 billion rights issue at an 8.4% discount to the closing price—its first since the 2019 HKEX listing. The stock fell 8.54% on the announcement day. Joseph Tsai and Wu Yongming subsequently increased their stakes by approximately HKD 800 million and HKD 400 million, respectively.

Faster-than-planned spending reflects both growing demand and rising chip prices. Management attributed the quarter's elevated CAPEX to hardware delivery schedules, CPU demand growth, and chip price increases.

Alibaba has its own ROI calculations for these investments.

Wu stated on the earnings call that industry consensus sees no ceiling for AI computing demand before 2030. Based on current AI product gross margins, computing investments could recoup costs within three years, potentially shortening to 2.5 or even 2 years as margins improve.

This framing transforms the 20GW vision into a time-bound financial commitment. Three-year payback relies on two Prerequisite s: AI product margins remain stable, and computing power stays fully utilized. The problem is that Alibaba controls neither fully.

T-Head Must Reduce Unit Computing Costs

The Zhenwu V900's role in this ecosystem is more concrete than its title as "China's most powerful AI chip."

According to Alibaba, the V900 features 216GB of memory and 1,200GB/s inter-chip interconnect bandwidth, enabling single clusters to scale up to 500,000 cards.

Its 216GB memory exceeds NVIDIA B200's 192GB, Huawei Ascend 910C's 128GB, and dwarfs Cambricon MLU590's 64GB and NVIDIA H20's 96GB. For the first time, a domestic AI chip surpasses NVIDIA's current-gen products in this core metric.

However, memory ≠ computing power. Alibaba did not disclose complete figures for absolute computing power, power consumption, process node, yield rates, or third-party testing—making "most powerful" currently an internal claim.

The iteration pace reveals additional insights. The Zhenwu M890 debuted just four months ago in May 2024, before the V900's announcement—suggesting competitive pressure. At Huawei's Full Connect Conference five days earlier, the Ascend 960 supernode was unveiled, featuring 4,096 NPU chips interconnected via Huawei's proprietary Lingqu protocol. Huawei committed to annual Ascend chip generations, with the 960DT arriving three quarters early in Q1 2027—nearly coinciding with the V900's mass production.

Both companies bypass process node limitations by focusing on interconnect protocols: Huawei's Lingqu vs. Alibaba's ICN. The next battleground will be ecosystem adoption of these protocols.

Partial cross-validation exists: M890-based Lingjun supernodes already power China's first 100,000-card cloud cluster, supporting training and inference for models like Qwen3.8 and Kimi K3 with over 2 trillion parameters. A new service will launch in Zhongwei, Ningxia, this Q4. The Zhenwu series serves over 650 enterprise clients, but "usable" does not equal "cost-effective."

The same event highlighted more utilization-focused metrics: Next-gen high-performance storage CPFS reduced model startup times by 50%, boosted peak computing utilization by 30%, and cut AI storage costs by 69% in real-world training. On the inference side, Tair KVCM cache scheduling achieved 99% cache hit rates and halved per-token costs. Reduced GPU data waiting times directly increase effective output per unit of computing power: While chip peak performance sets the upper limit, utilization determines the cost baseline.

Wu stated on the earnings call that nearly no Alibaba Cloud servers have idle cards—a claim he must both make and fulfill.

Organizational boundaries have already shifted accordingly.

In June 2024, Alibaba merged its Cloud Intelligence Group and T-Head into an "AI Cloud and Computing Services" segment, while placing model labs, the Qwen consumer group, and QwenWork under a separate "AI Labs and Applications" segment.

This structure directly integrates T-Head's value into cloud operations: A self-developed chip generates value not through external sales but by reducing hourly computing costs and improving Alibaba Cloud's gross margins. Management noted on the call that cloud profitability would be higher without consolidating AI chip results.

Goldman Sachs provided a quantitative benchmark post-Cloud Town Conference: T-Head currently supplies ~10% of Alibaba Cloud's computing power, with a medium-term target of 50%. If achieved, this would rewrite Alibaba Cloud's cost structure—half of computing power priced at internal costs, half procured at market rates—significantly reducing sensitivity to external chip price hikes. For the first time, "self-developed chips reducing costs" has a trackable progress indicator.

V900's success hinges not on chip market share but on how much Alibaba Cloud's per-token costs decline and flow into profit margins.

Notably, NVIDIA, Intel, and AMD all exhibited at the Cloud Town Conference's computing pavilion. Even as Alibaba unveiled full-stack self-developed technologies, it still invited the suppliers it aims to replace. Full-stack does not mean closed-loop; self-developed and external procurement will coexist for considerable time.

AgentCore Transforms Calls Into Long-Term Billing

Alibaba's launch of Agentic Cloud and AgentCore carries even greater commercial significance than releasing another large model.

Alibaba Cloud CTO Li Feifei systematized this approach: Agentic Cloud comprises three layers—Model (providing capabilities), Harness (ensuring stable operation in production environments), and Context (enabling continuous, domain-wide cognitive updates). Alibaba aims to productize the entire chain from agent invocation to operation, rather than competing solely on model capabilities.

Traditional chatbots receive questions and generate answers, billing primarily by tokens. When an agent accepts a task, it may run for minutes or hours, requiring context preservation, external tool calls, database access, code execution, identity/permission maintenance, and retries after failures. A single invocation thus expands into combined consumption of computing power, memory, storage, gateways, networking, and databases.

Billing structures reflect this shift: Beyond model fees, Alibaba Cloud charges for computing units (CUs), gateway calls, storage, dedicated control planes, and public internet traffic, while converting streaming responses, MCP persistent connections, and agent long sessions into invocation counts based on duration. Alibaba seeks revenue beyond model inference, extending to every agent operation, wait, memory update, and tool call.

Officially, Alibaba claims AgentCore boosts task completion rates by 99% and reduces total cost of ownership by 70%. The matching Agent Sandbox can create 100,000 sandboxes per minute, waking from deep sleep in under 600ms.

However, like Zhenwu's "most powerful" claim, these figures currently lack independent customer cases or third-party verification.

CAPEX explanations already reflect this shift: One reason for the quarter's CAPEX surge was anticipation of increased AI agent adoption by customers, prompting in advance increases in CPU computing capacity. While large models primarily consume GPUs, agents require running numerous tools and business logic beyond model calls, reintroducing CPUs into AI infrastructure procurement.

AgentCore allows enterprises to connect multiple model suppliers—Alibaba did not make Qwen the sole option because models do not guarantee customer relationships. While enterprises can switch models, migrating entire identity, permission, credential, tool, memory, and monitoring systems proves far more difficult.

In the cloud era, vendors retained customers through computing, storage, and databases. In the agent era, whoever controls what data agents access, which tools they invoke, and under whose identity they act will dominate the hardest-to-migrate layer of enterprise AI.

Qwen brings customers in; AgentCore keeps the bills coming.

Qwen's Losses: Cloud Acquisition Costs

Alibaba's latest segment data reveals cost allocations for this strategy.

In Q2 2024, AI cloud and computing services generated RMB 48.437 billion in revenue with adjusted EBITA of RMB 5.628 billion (+133% YoY), for a ~12% profit margin. AI Labs and Applications reported RMB 3.338 billion in revenue but RMB 13.861 billion in adjusted EBITA losses—2.5x the profits of the cloud/computing segment. Netting the two, Alibaba's AI segment lost ~RMB 8.2 billion this quarter (~RMB 32.9 billion annualized), accounting for over half of group operating profit.

Alibaba's AI machine currently runs on three ledgers: E-commerce funds it, Qwen burns cash, and cloud operations expand capacity while generating profits. Qwen currently acts as a demand subsidy for Alibaba Cloud, using group profits to cultivate user habits and inference calls while filling pre-built computing capacity.

Wu drew a historical parallel in his speech: In 1882, Edison's Pearl Street Station powered ~400 lamps nearby, initially giving away electricity with lightbulb sales—much like today's agent software bundling tokens.

Building computing capacity to generate profits while using applications to consume that capacity and bear acquisition costs forms Alibaba's current AI cycle. The Qwen App integrates with Taobao, Tmall, and instant retail; QwenWork integrates with DingTalk and enterprise workflows; Qwen Intelligence debuted in Honor's Magic9 on September 28. For application teams, this means user acquisition; for Alibaba Cloud, these entry points continuously generate inference, storage, and tool calls.

Measurable consumption has already emerged on the customer side. Ping An of China disclosed that since partnering with Alibaba, the daily average token consumption has exceeded 300 billion. However, to understand the penetration rate at this stage, we need to look at the national picture: According to the National Data Bureau, as of March this year, the national daily average token invocation volume exceeded 140 trillion; Wang Tao, at Huawei's Full Connect Conference in September, stated that the figure was approximately 500 trillion. Ping An's 3,000 billion represents only about six ten-thousandths of the national total. The 20GW does not correspond to existing demand but to demand that has yet to be created.

Alibaba itself also provided a demand-side slope at the Yunqi Conference main forum: Liu Dayiheng, head of Token Foundry, disclosed that two months after the release of Qwen3.8-Max, real user revenue increased by 8.5 times, and token consumption rose by 12 times. The slope is steep enough, but the base number was not disclosed.

Alibaba's management disclosed that in the last quarter, the annualized revenue (ARR) from AI-related products exceeded 49.5 billion yuan (approximately 7.3 billion USD), with expectations to approach 10 billion USD next quarter. This is currently the most concrete public indicator of "continuous payment from external customers."

However, when placed in the context of the overall ledger, the proportion remains small. The ARR of 49.5 billion yuan is equivalent to only about 18% of the annualized quarterly capital expenditure of 270.7 billion yuan. The quarterly revenue from AI-related products, at 12.376 billion yuan, also accounts for only about a quarter of the revenue from AI cloud and computing services.

Alibaba's AI can already drive growth in cloud services, but current AI revenue is far from sufficient to absorb the 20GW supply by 2032. Internal usage can improve server utilization but also incurs significant inference costs. Only with continuous payment from external enterprises or transaction and service revenue generated by consumer agents can these costs transition from internal circulation within the group to new revenue.

20GW Awaits Payment from External Customers

The Yunqi Conference focused on the technological path to super artificial intelligence, while financial reports discussed depreciation and cash flow. Alibaba Group Chairman Joseph Tsai offered a broader perspective at the same conference: The AI industry is shifting from isolated breakthroughs in cutting-edge technologies to scaling their value in real-world scenarios. He also stated, "We need open collaboration more than ever," a sentiment that, when juxtaposed with the five-layer full-stack self-developed system, highlights the tension in this business. It places Alibaba in a position where it must deliver.

The challenge lies in the timing. Internet companies could previously adjust server purchases based on traffic growth, but 20GW pushes Alibaba into a different operational discipline: build capacity first, then improve utilization, and finally reduce costs through scale. The larger the capacity, the more Alibaba cannot afford idle servers.

The most concerning scenario is not a lack of AI demand but demand growth lagging behind capacity expansion. Once computing supply becomes temporarily excessive, cloud providers must rely on price cuts to improve utilization. Lower prices will extend the capital recovery cycle, pushing back Wu Yongming's three-year payback period.

Alibaba holds several cards. It simultaneously owns cloud services, models, office software, e-commerce transactions, and numerous consumer access points. Internal demand can fill some server racks initially, AgentCore can create a difficult-to-migrate foundation at the operational layer, and T-Head can reduce unit computing costs.

External cloud revenue growing by 45% year-on-year and AI-related product ARR nearing 50 billion yuan indicate that demand is not hypothetical. However, the 13.861 billion yuan in application losses shows that the costs of model and application development currently far exceed the profits left by the cloud segment.

As long as application losses grow faster than cloud profits and e-commerce growth, the so-called full-stack advantage merely shifts costs within the group.

This is what makes this round of AI investment fundamentally different from any previous subsidy by Alibaba. Back then, Taobao burned through traffic costs, which stopped once the campaign ended. This time, it's burning through depreciation—once servers are powered on, costs accrue by the second and cannot be refunded. The 20GW is first reflected in capital expenditures and cash flow. Whether it can translate into AI revenue depends on how many enterprises are willing to let agents operate in their businesses long-term.

Note: The data in this article comes from public sources and does not constitute investment advice.

- END -

Solemnly declare: the copyright of this article belongs to the original author. The reprinted article is only for the purpose of spreading more information. If the author's information is marked incorrectly, please contact us immediately to modify or delete it. Thank you.