The Brief
China's National Data Bureau plans to accelerate the construction of a unified national data market by linking data resources with artificial intelligence adoption, bureau director Liu Liehong announced at the China International Big Data Industry Expo. China had built over 126,000 high-quality datasets encompassing more than 1,815 petabytes by August 2026, up over 89 percent from March. An official industry report revealed that AI software and tokenized data services generated 4.99 trillion yuan, accounting for 73.6 percent of national data software and resource output.
Why it matters
Beijing's focus on structured datasets and tokenized services addresses persistent hurdles around data valuation, distribution of proceeds, and digital rights monetization. By treating data tokens as invocable smart service units, policymakers and tech firms aim to establish standardized pricing mechanisms for raw inputs feeding large language models. The sharp surge in high-quality dataset creation shows that state-level policy is directly catalyzing data availability to feed domestic AI infrastructure.
China context
The National Data Bureau has anchored digital governance around a targeted cycle: scenarios guide data collection, data trains foundation models, models power enterprise applications, and applications yield economic value. Rather than approaching data governance purely from an administrative or defensive security lens, Beijing treats data as a core factor of production. Under twin state campaigns—'AI Plus' and 'Data Element X'—authorities are compelling state-owned enterprises, regional governments, and private businesses to digitize legacy assets.
Editor's View
EDITOR'S VIEW — Analysis and inference, not factual reporting.
Liu Liehong's emphasis on token-based data services signals a significant technical shift in how Chinese regulators conceptualize data trading. Traditional data exchanges faced illiquidity because raw data assets are notoriously difficult to appraise, secure, and clear. By framing data transactions around callable tokens rather than bulk transfers, China is aligning factor-market monetization directly with generative AI consumption patterns. However, scaling this model will test whether public-data governance and intellectual property protections can keep pace with fast-growing commercial dataset pipelines.
What to watch
- Implementation rules for the National Data Bureau's six action campaigns on high-quality dataset construction.
- Institutional frameworks establishing accountability and revenue-sharing mechanisms for public data resources.
- Commercial models and regulatory guidelines governing tokenized data products in enterprise AI deployments.
Key Takeaways
- 1National Data Bureau Director Liu Liehong outlined plans for an integrated national data market and a public data responsibility mechanism.
- 2China built more than 126,000 high-quality datasets totaling over 1,815 petabytes by August 2026, rising over 89 percent since March.
- 3Data tokenization is being advanced as a technical framework to solve valuation and revenue-sharing bottlenecks in AI training.
- 4AI-enabled data software and tokenized products generated 4.99 trillion yuan, comprising 73.6 percent of relevant industry output, according to an official report.
- 5Registered data enterprises reached 482,000 by late 2025, with fundraising clustering in foundation models, synthetic data, and embodied AI.
China's National Data Bureau plans to accelerate the construction of a unified national data market and establish an accountability system for developing public data resources, agency director Liu Liehong said at the opening ceremony of the China International Big Data Industry Expo in late August 2026, according to a report by People's Daily.
Speaking at the event, Liu detailed plans to advance six dedicated campaigns targeting high-quality dataset development, focusing on foundational capacity expansion and dataset annotation. The bureau aims to foster an operational cycle where real-world operational scenarios generate targeted data, refined data trains artificial intelligence models, models power specialized applications, and applications create tangible economic returns.
Liu emphasized that artificial intelligence has become a decisive mechanism for unlocking the latent value of data assets. Emerging technological tools—including data fabrics and ontology engineering—are reconfiguring conventional data production pipelines, expanding supply volume, improving data cleanliness, and prompting market demand. Under national initiatives promoting AI and data factor integration, Chinese enterprises have begun mining decades of accumulated proprietary records to power domain-specific AI workflows.
State-backed dataset initiatives have grown rapidly. According to data shared by the bureau, China had established more than 126,000 high-quality datasets by August 2026, representing an aggregate data volume exceeding 1,815 petabytes. That total reflects an increase of more than 89 percent compared to levels recorded in March 2026.
Liu also pointed to data tokenization as an emerging pathway for commercializing data assets. By converting raw proprietary data into callable smart service units, tokenization offers a structural solution to long-standing industry challenges regarding how to quantify data value and distribute economic returns across the supply chain. Market demand for token invocations is experiencing rapid expansion as use cases multiply.
The expo also saw the release of the China Data Industry Development Report (2026). The study found that data software products incorporating AI applications and data resource offerings featuring token services generated a combined output value of 4.99 trillion yuan ($700 billion), representing 73.6 percent of the surveyed industry segments. By the end of 2025, China's registered data enterprises reached 482,000, up 17.9 percent year on year, with venture financing heavily concentrated in foundation models, synthetic simulation, and embodied robotics services.