How Much Hardware Business Can Be Fueled by the Explosive Popularity of AI Applications?

09/28 2026 400

Doubao, DeepSeek, ChatGPT—these names are becoming increasingly familiar to us. AI applications have swiftly evolved from tools used by a select few early adopters to being widely integrated into various domains, including office work, search engines, programming, and content creation.

But what happens beyond the screens of our phones and computers when an AI application suddenly amasses tens or even hundreds of millions of users? How many orders for servers, chips, memory, and storage will these new users ultimately generate?

TrendForce predicts that by 2026, the combined capital expenditures of nine major cloud service providers—Google, Amazon, Meta, Microsoft, Oracle, ByteDance, Tencent, Alibaba, and Baidu—will exceed US$886.7 billion, marking a year-on-year increase of approximately 90%. Global AI server shipments are expected to grow by nearly 31%.

Capital expenditures related to AI have expanded from model training to encompass chips, servers, and data centers.

AI has already sparked a massive wave of incremental demand in the hardware industry. Large-scale model training initially propelled GPUs into the spotlight, followed by HBM, advanced packaging, and high-speed optical modules. Now, as AI transitions from answering questions to executing tasks, the workload borne by servers has also undergone a transformation.

From GPUs to CPUs, AI is reshaping the distribution of hardware benefits.

01

After ChatGPT gained popularity, the pathway for the first wave of hardware opportunities became clear.

Training large-scale models requires massive matrix operations, for which GPUs, with their numerous parallel computing units, are ideally suited. As models grow larger, the number of GPUs in training clusters increases, necessitating synchronous upgrades to the surrounding hardware infrastructure.

GPUs require high-speed data access, driving rapid growth in HBM demand. Clusters composed of thousands or even tens of thousands of GPUs necessitate constant data exchange, pushing 800G and 1.6T optical modules into expansion cycles. As chips grow larger and integration levels increase, advanced packaging technologies like CoWoS have also become scarce resources.

Thus, the hardware benefits from large-scale model training did not stop at GPUs but extended along the path of computing, storage, and interconnection.

However, with the emergence of Agents, the types of tasks servers need to handle have begun to diversify.

02

The most typical use of ordinary chatbots involves users posing questions, with the model performing inference and returning answers.

Agents, in contrast, must complete a series of actions. When a user asks an Agent to organize emails, search for flights, compare hotels, or analyze dozens of documents, the Agent must first understand the task, then invoke models to formulate steps, subsequently access web pages, query databases, run code, read files, and continue making decisions based on the results.

The rapidly growing Meta Muse provides a fitting example.

Muse achieved approximately 2.8 million downloads within 12 days of launch, enabling capabilities such as sending emails, shopping, and booking travel. Meta also designed an isolated Secure VM runtime environment for Muse, which includes a browser and the data environment required for Agent task execution. Tasks can continue even after users exit the app.

While a single physical server can certainly run multiple virtual machines simultaneously—meaning there is no simple one-to-one correspondence between “one user and one server”—each continuously running Agent requires underlying computing resources. Moreover, it consumes not just GPUs.

03

The division of labor between GPUs and CPUs in AI servers can be roughly likened to the relationship between firepower units and command/logistics systems in an army.

During large-scale model training, GPUs serve as the primary “firepower”; however, Agents introduce significant increases in scheduling, data preprocessing, tool invocation, and system management tasks, substantially boosting CPU workloads.

TrendForce’s research this year suggests that while traditional AI data centers typically configure CPUs and GPUs in ratios ranging from 1:4 to 1:8, some Agentic AI deployments may shift toward 1:1 to 1:2 architectures. Intel’s management has made similar observations: training workloads typically involve one CPU paired with seven to eight GPUs, but this ratio may drop to three to four GPUs during inference phases. More complex Agent and multi-Agent scenarios will further increase CPU proportions.

These figures do not imply that all AI servers will adopt a “one CPU per GPU” configuration. Given the significant variations across models, tasks, and system architectures, what they indicate is that as AI begins to invoke browsers, databases, and various traditional software tools, the proportion of general-purpose computing within the overall system is rising.

From GPUs to CPUs, AI is reshaping the distribution of hardware benefits.

While GPUs remain responsible for the heaviest model computations, the software execution and task scheduling introduced by Agents are elevating the importance of CPUs within the system.

04

As server demand for CPU resources rises, the incremental impact will continue to propagate.

First is server memory. Agents require more CPU resources for task orchestration, data preprocessing, and multi-tool collaboration, and CPU operation relies on DRAM. TrendForce has identified the CPU configuration upgrades driven by Agentic AI as one of the forces behind growing server DRAM demand.

Enterprise-grade SSDs are experiencing even more pronounced growth.

AI inference involves processing increasingly long contexts and storing substantial amounts of KV Cache. When DRAM costs are too high or capacity insufficient, some data can be offloaded to enterprise-grade SSDs. Meanwhile, Agents’ continuous task execution further increases read/write operations on files, databases, and caches.

In the first quarter of this year, revenue for the world’s top five enterprise-grade SSD manufacturers reached US$18.46 billion, up 86.1% sequentially. TrendForce attributes this growth primarily to the rapid adoption of AI Agent services and cloud provider procurement, noting significant supply-demand imbalances in the market during that period.

At least where enterprise-grade SSDs are concerned, the incremental demand driven by Agents is already evident in revenue and supply-demand data.

As AI transitions from model training to continuous task execution, cloud providers are expanding their procurement lists.

05

Recently, capital markets have extended this demand chain to IC substrates.

High-performance CPUs and GPUs cannot directly connect to server motherboards for all connections; instead, they require packaging substrates to bridge the chip’s highly dense internal circuitry with larger-scale PCBs. Server CPUs commonly use advanced ABF substrates.

If CPU demand continues to rise, substrates will naturally be incorporated into market demand projections.

However, this does not fully explain today’s tight ABF supply.

Currently, high-end ABF substrates are already in short supply, primarily driven by demand from GPUs, ASICs, and HPC. Unimicron noted in September that high-end ABF capacity is extremely tight, while Kinpo Electronics plans to increase its monthly ABF production capacity by 25% by 2027 compared to the end of 2026.

Thus, Agents are not the starting point of today’s ABF shortage. However, if Agents continue to drive up server CPU configurations, ABF will face additional demand from the CPU side.

While it would be unfounded to directly extrapolate ABF sales from Muse’s download numbers—given variables such as user engagement, task duration, virtualization efficiency, CPU configurations, and cloud provider procurement cycles—what the industry truly cares about is how much computing demand will ultimately be transmitted to servers if Agents expand from millions to hundreds of millions or even billions of users.

06

More than 150 years ago, British economist William Stanley Jevons observed while studying steam engines that as machines became more coal-efficient, Britain’s coal consumption did not decrease. Improved efficiency lowered usage costs, enabling steam power to penetrate more scenarios and driving total consumption higher. This phenomenon later became known as “Jevons’ Paradox.”

AI may not fully replicate this history, but similar dynamics are already visible today.

Models like DeepSeek continue to reduce inference costs, while chip computing efficiency improves. By simple logic, completing the same AI task should require fewer resources over time. However, when AI becomes cheap enough, people tend to use it more.

Instead of asking AI a dozen questions a day, users might soon have several Agents working continuously for hours. While individual tasks become cheaper, usage frequency and task duration increase simultaneously.

This is the key difference between Agents and previous AI applications.

When ChatGPT first emerged, the industry’s primary concern was how many GPUs would be needed for model training. Now that Agents are entering scaling phases, the question has shifted: How much general-purpose computing, memory, storage, and networking resources will be required if an AI works on behalf of hundreds of millions of people daily?

While no standard answer exists yet, capital expenditures and server procurement are already moving ahead of schedule.

As AI applications compete for users, server manufacturers are already recalculating their procurement lists.

Solemnly declare: the copyright of this article belongs to the original author. The reprinted article is only for the purpose of spreading more information. If the author's information is marked incorrectly, please contact us immediately to modify or delete it. Thank you.