08/26 2026
549
This is my 431st column article.
Recently, I had an in-depth conversation with Tao Jianhui, the founder of Taos Data. He shared a vivid anecdote: when visiting a cigarette factory and chatting with on-site engineers, one of them glanced at him and said, “You don’t even smell like tobacco.”
This remark applied to him—but it also describes today’s AI large models. No matter how remarkable their linguistic abilities, once inside a factory, they’re just “new employees.” They don’t understand what “the outlet temperature of Furnace No. 3 reads 620 degrees” actually means, which production line the furnace belongs to, what its normal fluctuation range is, or how it interacts with upstream and downstream processes.
The industry has long debated the dilemma of how large models can be integrated into factories. During our conversation, I compiled data from the prospectuses and annual reports of five leading industrial internet companies, totaling over 1,700 pages.
By comparing these five financial reports and archives with overseas industrial digitization samples from the past decade, we can answer a question that predates—and is more urgent than—“how large models can enter factories”: Why did the previous generation of industrial internet platforms, which “paved the way” for large models, fail one after another?

The Collapsed Platform Myth: A 14-Year “Experimental Report”
Let’s first look overseas. In 2012, GE introduced the concept of the industrial internet globally and formally launched its Predix platform in 2013. GE invested over $4 billion in its digital transformation, with its digital division peaking at nearly 28,000 employees. However, after GE restructured its digital business in 2018, the grand narrative of its “horizontal platform” ended. Siemens’ MindSphere made a high-profile debut in 2016 but was quietly renamed Insights Hub in 2023 and downgraded to a product component. In the same year, Google IoT Core and IBM Watson IoT platform were quietly shut down, while SAP’s Leonardo brand had already faded away.
Now, let’s shift our focus to China. The Ministry of Industry and Information Technology’s list of “cross-industry, cross-domain platforms” expanded rapidly from 10 in 2019 to 51 in 2023 before contracting, with the evaluation system beginning to tier platforms. In the capital market, Rootcloud withdrew its IPO from the Science and Technology Innovation Board in 2023, marking a turning point.
According to Frost & Sullivan, in China’s platform-based industrial data intelligence solutions market, the top-ranked company holds a mere 1.7% market share. After over a decade of grand industry narratives and seven years of competition among the 51 cross-industry platforms, the “industry leader” that emerged controls less than 2% of the market.
This isn’t just a company’s report card—it’s a brutal industry experimental report.
The conclusion is clear: In the industrial sector, there is no such thing as “platform economics.”

The term “platform” was borrowed from the consumer internet. Its core premise is “network effects”: each additional user increases the value of the network for others. But industrial data inherently lacks this property. When Factory A connects its production line to a platform, it doesn’t enhance the value of Factory B’s connection—the cross-client externality is nearly zero. Without network effects, there’s no winner-takes-all dynamic.
The second dilemma lies in the marginal cost of services: people.
Consumer platforms can serve 100 million users with roughly the same server costs as serving 1 million. But industrial platforms must assign a dedicated project team for each new client. GE’s digital division once ballooned to nearly 28,000 employees. Rootcloud’s workforce surged from 500 to 1,500 in two years. In 2021, its per-capita revenue was 340,000 RMB, but per-capita losses exceeded that figure. The revenue generated by each employee couldn’t even cover their share of the financial losses.
The third dilemma is misaligned pricing anchors.
The previous generation of platforms aimed to charge entrance fees: connection fees, platform usage fees, and subscription fees. But throughout the past 50 years of industrial history, customers have only willingly paid for two things: validated, deterministic outcomes (hence the prevalence of project-based models) and “annuities” with prohibitively high switching costs, exemplified by OSIsoft (PI System). This company, with annual revenue of approximately $400 million, was first acquired by AVEVA for $5 billion, and Schneider Electric later privatized AVEVA at a valuation exceeding $10 billion. What did the giants buy? Decades of data accumulation from over 20,000 global sites and the massive applications built upon them. Customers couldn’t switch, so maintenance fees were paid year after year.
In contrast, the previous generation of platforms lacked both: they couldn’t charge based on results, nor did they possess locked-in assets that customers couldn’t abandon.
The Turning Point: How AI Unravels “2.5” Dilemmas
The turning point has arrived. AI directly undermines “two and a half” of the three major dilemmas.
First, consider marginal costs. Tao Jianhui told me that over the past year, he has used AI almost daily to write code and documents, boosting internal R&D efficiency at least fivefold. Software production costs have plummeted. More critically, service costs—deployment, troubleshooting, reporting, and dashboards—which once required a project team per client, can now be handled by AI agents within the system. For the first time in industrial software history, marginal service costs have been exponentially compressed.
Next, network effects—half-solved by semantics. Standardizing data and building information models were essentially “public goods” for the next project. The current project benefited the next, but no client wanted to pay for it alone. As a result, standardization never advanced, and the 50 cross-industry platforms developed hundreds of models.
For example, the number “620” means nothing in isolation. It could represent temperature, rotational speed, or even a random code. “Semantics” are the instructions attached to this number: it’s a temperature in Celsius, measured at the outlet of Furnace No. 3, belonging to Production Line 1 in Workshop 2. Its normal fluctuation range is 580–650 degrees, and an alarm must trigger above 680 degrees. Its upstream process is batching, and its downstream is cooling. Only with these labels can an AI large model truly understand what “620 degrees” means and its relevance. This is akin to compiling a “dictionary” for all factory data, allowing AI to look up the true meaning of each term.
AI large models may transform the cost-benefit structure of this effort. The quality of the underlying semantic framework directly determines whether AI can function effectively. Only when each object has an identity, each relationship is explicitly defined, and each data point carries business semantics can a large model interpret “620 degrees.”
For the first time, semantics have a direct buyer—the “dictionary” may become a core product. Whose semantic model becomes the de facto standard will likely control the entry point for industrial AI.
The third dilemma, pricing anchors, is the most deeply disrupted, as the basis for valuation has fundamentally changed.
AI large models are inherently probabilistic and uncertain, while industrial sites demand absolute certainty. Thus, basic “functions” can be free—AI has drastically reduced the production costs of code and features. But “liability” cannot be free. What enterprises are truly willing to pay for is “insurance” to ensure the system’s flawless operation, including service guarantees, compliance audits, role-based access control, and human-in-the-loop oversight.
This forms a solid pricing wall between commercial and free versions. In short: functions are free, but liability costs.
A new concept must be introduced here. The PI System’s annuity model was anchored in “data immobility,” but AI is continuously eroding migration costs. Code can be generated, data can be transferred, and even ontology models can be exported with one click. Assets that were “immovable” are rapidly depreciating. The annuity model won’t disappear, but its anchor may shift: from passive binding due to “data immobility” to active service based on “liability must be assumed.”
After discussing 2.5 dilemmas, the remaining half is that AI has not yet created network effects at the client data layer. Factories still resist sharing data, and single-factory closed loops remain an ironclad rule in industry…
But curiously, network effects may be sprout (budding).
A New Growth Flywheel: Competing for “Model Memory” and “AI Sovereignty”
The first true growth flywheel in industrial software history may emerge from the “memory of AI large models.”
In the past, software companies expended significant sales resources to persuade CIOs to purchase products. In the AI era, however, code writing, system integration, and technology selection are increasingly automated by AI agents. The decision-makers for software procurement are shifting from “humans” to “models.”
If your industrial software is open-source and frequently appears in AI training corpora and tool registries, AI will prioritize your tools when generating solutions. Deployment volumes increase, generating more data, creating a virtuous cycle.

The previous generation of platforms competed for CIO budgets; this generation competes for “model memory.”
Meanwhile, Palantir, one of the most successful overseas data companies, emphasized another selling point in its earnings report: “AI sovereignty.”
Palantir reported $4.475 billion in revenue for fiscal 2025, up 56%; in Q2 2026, revenue reached $1.94 billion, up 93%, with a single-quarter net profit of $1.06 billion and a net profit margin of 55%. In his 2025 annual report, CEO Alex Karp dubbed this trend “the commoditization of cognition.” Six months later, he pushed the narrative further to “AI sovereignty,” repeatedly assuring clients that their proprietary factory data and process advantages would never train any public models.
On one hand, open-source tools gain AI’s favor and distribution by “entering the corpus”; on the other, Palantir sells trust by “isolating the corpus.”
These seemingly contradictory strategies actually form a complete industrial AI commercial landscape: as an infrastructure “tool,” the more public and widely used, the better, as this determines how extensively AI can call upon it. Conversely, a client’s specific “process context” must remain private, as this defines the factory’s core competitiveness.
What is “context”? Context is the data’s backstory—the global state of the site at that moment. Knowing “Furnace No. 3’s outlet temperature is 620 degrees” is insufficient. What product is being manufactured on this line? Is the furnace in the heating phase after ignition, or has it been running steadily for eight hours? Was a new batch of raw materials introduced last night? Does this factory’s veteran operator habitually keep temperatures 10 degrees lower than peers? At 620 degrees, this reading might be normal during heating but a fault precursor during stable production. The full contextual background needed to judge the data’s validity is the context.
In short: semantics teach AI to “read,” while context teaches it to “understand.” Semantics label individual data points; context describes the entire site’s ecosystem.
The cigarette factory engineer’s remark at the beginning—“you don’t smell like tobacco”—exposes AI’s lack of context. The large model recognizes the word “temperature,” but without immersive experience in the factory, it doesn’t understand local rules, habits, or emergency (unexpected situations).
Understanding this layer clarifies why semantics and context must be treated separately commercially.
Dictionaries (semantics) can and must be standardized. Temperature should uniformly be called “temperature,” not “Temp” in one factory and “T01” in another. Whichever dictionary format becomes the universal standard wins the ecosystem, making it inherently suited as public infrastructure.
But context harbors each factory’s “secret sauce”—formulas, parameters, and decades of a veteran operator’s intuition. These are core assets that give a company an edge over competitors and must never become public knowledge, let alone training data for others’ models.
As a tool, semantic standards should be as open and widely adopted as possible. As a core asset, on-site context should be as private and valuable as possible. The entire business of this generation of data infrastructure revolves around doing both these opposing things simultaneously.
Two paths diverge from here:
One is Palantir’s “artisanal, high-price” route: building private models for clients + engineer-intensive services + assuming responsibility for outcomes. This path is extremely expensive and selective. Its 55% quarterly net profit margin proves its high ceiling, but it’s essentially a premium project-based model that bypasses platform economics.
The other is the “volume route”: open-source underlying engines, free basic platforms, and bets on scale and ecosystem. This is the path the previous generation of industrial internet platforms failed to build, gambling that AI can truly reduce marginal delivery costs.
Epilogue
Looking back at the past 14 years, the pioneers of industrial internet platforms worked tirelessly, but their attempt to force-fit the consumer internet’s “platform economics” onto industrial complexity ultimately succumbed to the gravitational pull of business model fundamentals.
Today, the rise of large models isn’t merely about providing a smarter chatbox—it’s restructuring the cost framework and value anchors of industrial software.
When the marginal cost of code generation plummets, when semantic models make data truly readable, and when business models shift from “selling tickets” to “charging for liability,” the gears of industrial digitization finally begin to engage.",