EN

Beijing moves to export the data that trains global chatbots

Khaled Aziz

Key Points

  1. China's data agency set a plan to make the country a data powerhouse by end-2028.
  2. Beijing pledged at Shanghai's AI conference to share data sets with developing countries.
  3. Analysts warn exported Chinese data could carry party narratives into foreign AI models.

The latest:

China wants its own data — not just its cheap models — inside the systems training the world’s chatbots, and its National Data Administration has published a blueprint to turn the country into a data powerhouse by the end of 2028, according to The New York Times. The plan calls for high-quality data sets in more than two dozen strategic fields, and for sharing them abroad.

Details:

  • The blueprint: The National Data Administration’s plan, unveiled earlier this year, proposes building high-quality data sets across more than two dozen strategic fields including scientific research, industrial manufacturing and autonomous vehicles, and calls on China to share those sets worldwide, according to The New York Times.
  • The Shanghai pledge: At the World Artificial Intelligence Conference in Shanghai last month, China pledged to share data to help dozens of attending developing countries build their own AI systems. Xi Jinping used the same conference to cast China as a champion of an open approach to artificial intelligence.
  • What is already out: State labs and state-owned media have released large data troves for download globally, including free access on GitHub and Hugging Face. The largest, WanJuan, was created by the state-backed Shanghai AI Laboratory and covers history, sports, law, current events, medicine and literature.
  • Language reach: WanJuan is designed to align with mainstream Chinese values and is available in Arabic, Korean, Russian, Thai and Vietnamese besides Chinese, positioning it as a starting point for developers building or fine-tuning models.
  • The stated motive: Yu Xiaohui, president of the state-affiliated China Academy of Information and Communications Technology, wrote last month on the data administration’s website that competition in the AI era is not only about models and computing power but also about high-quality data supply.
  • The domestic problem: China holds vast data from mass surveillance and its largest tech platforms, but it sits fragmented in silos across departments and companies. Xiaomeng Lu of Eurasia Group said resolving domestic data-flow hurdles is Beijing’s top priority; the plan calls for breaking those silos open.
  • Distillation: Lu said scarce usable data is one reason Chinese labs lean on distillation, collecting outputs from powerful systems to build their own models. US firms including Anthropic complain their Chinese competitors are unfairly copying their technology.
  • The US gap: American providers such as Mercor and Scale AI now recruit mathematicians to annotate proofs and lawyers to mark up briefs. Beijing’s blueprint mandates a shift from cheap manual labeling to expert-type annotation, asking universities to build annotation courses and steering graduates toward the field.
  • The research finding: A recent study in Nature found Chinese state narratives have seeped into data training American models including ChatGPT and Claude. Asked whether China is an autocracy, or whether Xi is a good leader, the chatbots answered far more favorably to Beijing in Chinese than in English.
  • The warnings: Alex Colville of the Australian Strategic Policy Institute said the downside is greater power for authoritarian states to dictate a chatbot’s values. Kenton Thibaut of the Atlantic Council said the overarching goal is “to make the world safer for the party.”

Background:

Beijing Institute of Technology researchers published findings in 2023 that ChatGPT misidentified basketball star Yao Ming as the first Chinese woman to play professionally in the US, confused two classic Chinese novels, and generated biased commentary about China. Training data remains overwhelmingly English.

Between the lines:

Princeton sociologist Brandon Stewart, a co-author of the Nature study, argued AI disconnects the messenger from the message, and that readers would react differently knowing an answer came from People’s Daily. Chinese chatbots such as DeepSeek’s already evade questions on Xi and zero-Covid. Thibaut said cheap, capable Chinese models create technological lock-in that data sets deepen.

What’s next

Watch whether the data administration meets its end-2028 targets, which developing countries adopt Chinese data sets after the Shanghai pledge, and whether US labs move to screen Chinese-language training data.

Source:

What to read next

WSJ: China’s Hengli is top buyer of sanctioned Iranian crude

This Week’s Top Five Stories in AI