The Brief
Chinese artificial intelligence compute providers are transitioning from traditional hardware capacity rentals to token-based service delivery in response to exploding inference workloads. According to National Data Administration figures reported by People's Daily, daily national token calls reached over 140 trillion in March 2026—a thousandfold surge compared to early 2024. Industry experts highlight that competition is moving away from raw GPU cluster size toward per-watt token productivity, driving operators to integrate custom hardware with specialized software scheduling engines.
Why it matters
The shift from leasing GPU card-hours to delivering utility-metered tokens changes how AI infrastructure is monetized and optimized. As inference workloads overtake model training in volume, compute providers must move beyond basic hardware aggregation toward full-stack efficiency—combining liquid cooling, specialized chips, and dynamic software optimization to improve margins.
China context
The rapid growth in token consumption reflects the widespread integration of generative AI into commercial software across China. Chinese data infrastructure initiatives increasingly focus on maximizing energy efficiency and utilization rates across both domestic and overseas computing clusters to sustain high-volume demand while keeping operational costs manageable.
Editor's View
EDITOR'S VIEW — Analysis and inference, not factual reporting.
Transitioning to 'Token-as-a-Service' converts AI compute from a capital asset lease into a managed utility service. Crucially, long-term market leadership in this space will not belong to companies that simply hold raw GPU inventories, but to operators capable of organizing hardware procurement, dynamic cooling, enterprise storage, and real-time inference engines into a reliable, cost-effective delivery ecosystem.
What to watch
- Adoption rates of TaaS billing models among major Chinese public cloud providers and independent data centers.
- Technological advancements in custom ASIC hardware, liquid cooling systems, and enterprise SSD storage tailored for inference.
- Industry efforts to establish standardized token metering and service-level agreements across computing networks.
Key Takeaways
- 1Daily national token calls in China exceeded 140 trillion in March 2026, up over 1,000-fold from early 2024, according to official data.
- 2Industry competition is shifting from cluster scale to per-watt token production efficiency, driving the rise of Token-as-a-Service (TaaS).
- 3Infrastructure providers are deploying full-stack solutions—combining liquid cooling, enterprise SSDs, ASICs, and inference engines—to lower operational costs.
China’s artificial intelligence infrastructure sector is undergoing a structural shift, pivoting from leasing physical hardware by the GPU hour to delivering compute measured directly in tokens, according to a report by People's Daily.
At the World Artificial Intelligence Conference 2026, "Token factories" emerged as a dominant focal point as industry demand expanded beyond model training into large-scale inference deployment. Figures from China's National Data Administration show that daily token calls across the country exceeded 140 trillion in March 2026, representing a more than 1,000-fold increase from roughly 100 billion in early 2024.
Zheng Weimin, an academician at the Chinese Academy of Engineering, stated that tokens have become the primary production metric for the AI sector. According to Zheng, market competition is rapidly shifting from Model-as-a-Service (MaaS) to Token-as-a-Service (TaaS), moving the central benchmark from overall compute cluster scale to per-watt token efficiency.
This operational evolution requires service providers to execute tighter software-hardware integration. On the hardware side, operators are incorporating liquid cooling systems, enterprise solid-state drives (SSDs), and application-specific integrated circuits (ASICs) across compute clusters. On the software side, providers rely on specialized inference optimization engines and intelligent load-balancing tools to smooth traffic peaks, driving higher average server utilization rates.
Industry insiders describe TaaS as an end-to-end chain spanning hardware procurement, software scheduling, capital backing, and client integration. For example, compute provider Xingyun Technology recently introduced a token-based cloud billing service backed by its proprietary AlayaJet inference engine. According to People's Daily, the company reported first-half 2026 revenue of 254 million RMB ($35.3 million), up 497.11% year-on-year, alongside a net profit of 12.02 million RMB, achieving a turnaround to profitability.
As compute transitions from hardware asset availability to managed service delivery, the capacity to organize fragmented logistics, deployment, and ongoing maintenance into a stable pipeline is fast becoming the key operational test for Chinese infrastructure providers.