The Edge AI Battle Has Officially Kicked Off | Business Wave

09/17 2026 531

Editor | Yang Xuran

Shifting focus from the fiercely competitive, resource-heavy super cloud-based large models to the development of edge models or intelligent terminals powered by edge AI has emerged as a key strategy for many companies diving deep into the AI wave, particularly hardware firms.

Li Dahai, CEO of ModelBest, predicts that 2026 will herald the 'first year of commercialization for edge AI.' Despite not yet securing relevant qualifications for AI phones, the company has attracted substantial financing thanks to its innovative concepts and products.

In July this year, China's Cyberspace Administration mandated the systematic registration of AIGC services, adding seven large models to the registry, including Apple Intelligence, Huawei Xiaoyi, and Xiaomi HyperAI.

Following this, the Doubao AI phone NaviX Ultra made its official debut on September 16. Apple Intelligence, coupled with the new Siri AI, was unveiled at WWDC26, while the newly launched Xiaomi 18 Fold boasts a built-in Xiaomi MiMo edge large model.

In the smart car arena, StepFun, under the leadership of Yin Qi, is making strategic bets on both automotive and mobile phone markets. It is partnering with Geely to penetrate the smart car space and has also launched its first native Agentic Phone, the STEPX Neo.

Across the ocean, tech behemoths like Meta, NVIDIA, Google, and OpenAI are also venturing into edge R&D or collaborating with device manufacturers to create innovative AI hardware.

The battle for edge dominance has effectively commenced, with the smart hardware industry bracing for a massive wave of upgrades.

This article, a deep-dive analysis from the Business Wave content team, explores this trend in detail. Follow us across multiple platforms for more insights.

Achieving More with Less: The Edge AI Advantage

Compared to renowned cloud-based large models like DeepSeek and ChatGPT, edge models are more compact, with fewer parameters, enabling them to run directly on device-side computing power without cloud dependency.

The parameter count stands as the most significant differentiator between AI large models and edge large models. AI large models often boast trillions of parameters, with Kimi K3 (Yuezhi Darkside) reaching 2.8 trillion (2.8T) MoE, and Anthropic's Mythos hitting 10 trillion parameters, trained on a staggering 300 trillion tokens of data.

These massive parameter counts empower large models to deliver superior performance, capturing complex data patterns and excelling in diverse tasks. However, their reliance on cloud-based inference introduces latency and privacy concerns, rendering them unusable when disconnected or facing unexpected issues.

Edge devices, constrained by limited chip performance and storage space, necessitate the compression of edge large model parameters. Techniques like knowledge distillation, pruning, and quantization enable edge large models to reduce their parameter counts to the billion (B) scale.

Currently, many mainstream model providers, including DeepSeek, Meta Llama, and IBM Granite, have introduced smaller model versions that excel in specific tasks.

This means that large foundational models can be scaled down into more agile versions, achieving faster inference speeds, lower memory usage, reduced power consumption, and maintaining high performance levels. This makes them ideal for deployment on end devices like smartphones, PCs, cars, and robots.

AI large models cannot achieve cost reductions through economies of scale like the internet, posing a significant bottleneck for their commercialization.

Training a super large model consumes vast resources; the estimated training cost for Mythos is a staggering $10 billion. Once deployed, the model's inference costs escalate with increasing daily active users and frequency of use.

Guolian Minsheng Securities estimates that, including Doubao's free consumer services, ByteDance's daily costs range from 132 million to 240 million yuan. The exorbitant costs of cloud-based inference will severely hinder the sustainability of large model commercialization.

Edge AI, on the other hand, can 'achieve more with less,' offering high immediacy, strong privacy, and enabling more customized innovations with local data, further reducing costs for developers and promoting the commercialization of edge AI.

Driven by the 'Moore's Law' for chips and the 'Density Law' for edge AI, edge intelligence is accelerating its commercialization.

In late 2024, Tsinghua University's Sun Maosong and Liu Zhiyuan teams, along with the open-source community OpenBMB, proposed the 'Capability Density' concept to measure the intelligence level per unit parameter of large models.

The Density Law posits that 'the capability density of large models grows exponentially over time, doubling approximately every 3.3 to 3.5 months,' meaning 'every hundred days, a model with half the parameters can achieve the current optimal performance.' Combined with Moore's Law in semiconductors, where chip computing power at the same price doubles approximately every two years, the capability density of models doubles roughly every hundred days. This synergy between software and hardware is propelling AI large models toward end devices.

Currently, AI phones and AI PCs are rapidly gaining traction (though some data may exhibit measurement inconsistencies). The penetration rate of generative AI phones reached 36% in 2025 and is expected to exceed half by 2027, with edge AI transitioning from a high-end selling point to a standard product feature. The shipment share of AI PCs climbed to 31% in 2025 and is set to become the norm by 2029.

Coupled with innovations in smart cars, embodied AI, and even smart hardware forms, the era of edge AI is dawning.

A Turning Point: The Rise of Edge Intelligence

The emergence of a 'duck' is reshaping the evolution of global physical AI and edge AI.

On September 3, NVIDIA acquired Hugging Face, the world's largest AI open-source platform, for $12.93 billion. Just a week earlier, Hugging Face's Pollen Robotics unveiled a $399 bipedal robot, Microduck.

Standing just 25 centimeters tall and weighing less than 800 grams, this robotic duck may have triggered the 'ChatGPT moment' for edge AI—pre-orders surpassed $2.6 million within 24 hours, with an average of one unit sold every four seconds during peak periods. Consumers ordering now face a wait of over six months for delivery.

Microduck features an open-source design, allowing users to retrain it themselves, meaning this robotic duck is not just a toy but capable of continuous self-learning. Players can upload their training results to the Hugging Face Hub, where data generated by countless devices in real-world environments will drive algorithmic evolution at nearly zero marginal cost.

Since ChatGPT's debut, the AI narrative has revolved around larger models, more powerful GPUs, and vast computing clusters.

Microduck, however, eschews high-end chips from NVIDIA, AMD, or Qualcomm, relying instead on a mid-range SoC RK3566 developed by Rockchip in 2020.

Compared to humanoid robots costing tens of thousands of dollars, such a low-cost, fully open-source, and ecologically rich 'duck' significantly lowers the entry barrier for edge intelligent devices.

As Hugging Face CEO Clem Delangue remarked, Microduck's launch signifies the arrival of an 'affordable open-source robot era,' aiming to enable more people to participate in the development of physical AI and world models.

In the AI PC space, chip giants like AMD, Intel, and NVIDIA are accelerating their strategic positioning. In January, AMD launched its Ryzen AI 400 and Ryzen AI PRO 400 processors for Windows 11 AI PCs, delivering up to 60 TOPS of computing power to run large models locally.

Intel unveiled its third-generation Core Ultra processor at CES 2026, the first computing platform built on Intel's 18A process.

In June, Jensen Huang introduced the RTX Spark superchip for AI PCs at the Taipei GTC conference, featuring 1 Petaflop of computing power and up to 128GB of memory, capable of running 120 billion-parameter large models locally. This fall, major PC manufacturers like ASUS, Dell, HP, and Lenovo will successively release products powered by this chip.

Driven by factors such as cost, commercial viability, and technological evolution, the future evolution of AI large models may gradually divide into three categories:

The first category comprises well-known cloud-based large models, handling complex reasoning, long-form text/video generation, and Agent capabilities, defining the upper limit of AI capabilities;

The second category consists of industry-specific models, excelling in vertical domains like finance, healthcare, and justice, serving professional scenarios and addressing AI implementation challenges;

The third category, and our current focus, is edge models, designed for high-frequency, real-time, privacy-preserving, and offline tasks, determining the future reach and breadth of AI.

The hierarchical collaboration of 'cloud defining the upper limit, industry defining the depth, and edge defining the entry point' is becoming increasingly evident.

China's Edge AI Boom: Key Players and Strategies

Under the evolving trend of AI large models, China's edge AI is also reaching a tipping point.

Among the key players, ModelBest positions itself as a pure edge model infrastructure provider, adhering to the 'achieving more with less' philosophy and attempting to manage all intelligent hardware with its models; Doubao Phone, co-developed by ByteDance and ZTE Nubia, aims to seize the super entry point for edge-cloud collaboration; StepFun defines itself as a native multimodal terminal solution provider, primarily building AI brains for cars and phones; intelligent driving solution providers like Horizon Robotics serve as integrated intelligent platforms, with barriers in automotive-grade chips and software-hardware decoupling capabilities.

Among these, ModelBest, co-founded by Liu Zhiyuan, is a standout player. The company made a 'counter-consensus' strategic decision in the second half of 2023, abandoning hundred-billion-parameter cloud-based infrastructure to fully commit to edge models and originally proposed the 'Density Law.' Leveraging its MiniCPM 'pocket rocket' open-source model, it has built an edge model ecosystem.

On September 8, it open-sourced the new-generation 'pocket rocket' MiniCPM5-2B, with a parameter scale of 2B, designed for local operation on end devices like phones and PCs while possessing strong Agent capabilities.

According to the company, in the latest AA benchmark evaluation, it achieved a comprehensive score of 23, ranking first among global open-source infrastructure models with under 4B parameters. This means it achieved the capabilities of a 4B-scale model with only half the parameters.

In terms of commercialization, ModelBest positions itself as an independent third-party provider, currently focusing on phones, cars, and embodied AI. Among the seven registered mobile edge AI models announced by the Cyberspace Administration, Samsung's 'Galaxy AI' edge capabilities are powered by MiniCPM. Entering Samsung's supply chain means it is not only the first domestic edge AI supplier to enter an international phone giant but also the only independent and pure edge model company among the first batch of registered mobile model enterprises.

Previously, the company's automotive edge intelligent cockpit assistant, cpmGO, was mass-produced and deployed with Changan Mazda's MAZDA EZ-60; in the embodied AI direction, MiniCPM-Robot, with 1.5B parameters, has been deployed in scenarios like exhibition hall guidance and campus inspections in collaboration with UBTECH Robotics.

After Yin Qi took the helm at StepFun, the company accelerated its commercialization. StepFun seems to lean toward a 'software-hardware integrated' full-stack approach.

On July 13, StepFun released its first large model-native AI terminal brand, STEPX, and simultaneously unveiled its purported native Agentic Phone, STEPX Neo. Yin Qi's goal is to integrate the infrastructure model, Agent system, and hardware terminals into a complete chain, realizing 'AI + terminals.'

Previously, StepFun, Qianli Technology, and Geely have formed a strategic alliance. Qianli Technology has also collaborated with StepFun to develop the vehicle version of Step AOS, touted as the native intelligent driving foundation for vehicles. It deeply integrates on-device large models, world models, and autonomous driving, with Geely Automobile highly likely to be the first to implement it.

Among large model manufacturers, Doubao has been the most aggressive in developing on-device solutions, being the earliest to extend its reach into the mobile phone sector.

After testing the waters with the first-generation engineering model Nubia M153, Doubao, in collaboration with ZTE Nubia, adjusted its technical approach from the first-generation cross-App solution to a second-generation on-device MCP, aiming to create a truly AI Native smartphone.

The Nubia NaviX Ultra is hailed as the "world's first mass-produced AI-powered agent smartphone," setting itself apart from merely downloading an app like Doubao. By leveraging voice commands to operate its built-in Agent, users can effortlessly hail a taxi or order milk tea with a single spoken sentence. Furthermore, the device endeavors to endow AI with "eyes," enabling it to perceive and interact with the physical world.

However, the phone's ecosystem primarily revolves around ByteDance's suite of products, with limited integration of select third-party applications such as Caocao Chuxing. Due to security considerations, its Agent functionality remains disconnected from widely-used payment platforms like Alipay or WeChat Pay.

In the fiercely competitive landscape of consumer-facing AI applications, Doubao has successfully amassed a substantial user base. Through its collaboration in developing AI-powered smartphones, ByteDance aims to further solidify Doubao's foothold among mobile users and capture a larger share of the on-device AI market.

Among providers of smart car solutions, Horizon Robotics zeroes in on on-device deployment for two key scenarios: intelligent driving and embodied intelligence. Its competitive edge stems from its integrated hardware-software solutions, encompassing the Journey intelligent driving chips, Starry cabin chips, and the HSD (Horizon's full-scenario assisted driving system). Additionally, the company is actively developing the HoloMotion and HoloBrain models to advance embodied intelligence capabilities.

Desay SV champions an "end-cloud collaboration" approach. Presently, the high costs associated with integrating large-scale models into vehicles pose a significant barrier to widespread adoption. Consequently, Desay SV tailors its solutions to meet diverse customer requirements by matching varying levels of chip computing power. This strategy includes offering lightweight on-device models and evolutionary plans for iterative upgrades of the full-modal Omni model.

On September 1st, iFLYTEK open-sourced two on-device large models, Spark X2.5-4B/1.7B, both capable of supporting up to 1 million Token contexts. Their implementation zeroes in on "AI-powered office" applications, facilitating not only "offline conversations" but also enhancing the ability to handle complex tasks, with the aim of penetrating the enterprise productivity sector.

In Conclusion

As the competition among large AI models transitions from a "hundred-model battle" to a "contest among the elite," and as leading large model companies capitalize on both financial and market opportunities, on-device solutions have emerged as another focal point for AI model companies.

The low cost, lightweight design, offline usability, and enhanced privacy and security offered by on-device models present significant advantages, paving the way for a potential commercial explosion in the near future.

To date, on-device AI has yet to fully demonstrate its unique value and significance. However, as previously mentioned, this will gradually become evident as the functions and roles of large models continue to diversify.

Nevertheless, an on-device model company that operates independently of all third parties and views all hardware companies as potential clients is likely to face disappointment. This is because these potential clients will inevitably explore developing their own small-scale on-device models. From this perspective, the ultimate growth prospects of "pure on-device intelligence" companies, exemplified by FaceMind, are inherently limited.

End-cloud collaboration is poised to become a prevailing trend in the future. Large hardware institutions will delve into model compression and the continuous iteration of NPU chips, with on-device solutions gradually evolving to support Omni full-modal capabilities. This will facilitate the widespread integration of AI into various intelligent hardware devices (meaning "AI will permeate various smart hardware"). Following such a comprehensive downscaling of artificial intelligence capabilities, the era of AI-driven inclusive intelligence will draw ever closer.

Solemnly declare: the copyright of this article belongs to the original author. The reprinted article is only for the purpose of spreading more information. If the author's information is marked incorrectly, please contact us immediately to modify or delete it. Thank you.