SBS News

Hiding 'Sinocentrism' and Spreading Globally: China's Next Strategic Move Beyond 'Distillation'


Add SBS News to Google preferred sources
Show video

China has begun supplying large-scale datasets for AI training to the world.

Analysts suggest this is a strategy to narrow the AI technology gap with the United States while expanding Chinese influence across global AI ecosystems.

The New York Times reported that China's National Data Administration has unveiled a blueprint to make China a data powerhouse by the end of 2028.

The plan involves building high-quality datasets across more than 20 strategic sectors, including scientific research and industrial manufacturing.

Datasets refer to massive collections of texts, images, and videos used to train artificial intelligence.

China has previously been evaluated as lacking high-quality data for AI training compared to the United States.

Although massive amounts of data were accumulated through the Chinese government's extensive surveillance systems and private platforms like Alibaba and Tencent, limitations remained in effectively utilizing this data for AI development because it was scattered across various institutions and companies.

In the United States, even professionals such as mathematicians and lawyers are being mobilized to produce high-value data necessary for AI training.

In contrast, Chinese AI companies have relatively relied more on the so-called "distillation" method, training their own models using the outputs of powerful foreign AI models.

To narrow this gap, China plans to establish data annotation training courses at universities and directly secure professional training data.

The data built this way will also be supplied outside of China.

At the World Artificial Intelligence Conference held in Shanghai in July, China announced it would share data to support developing countries in building their own AI.

Large-scale datasets created by China are already being made publicly available for free on global platforms such as GitHub and Hugging Face.

A representative example is the Wanjuan dataset created by the Shanghai Artificial Intelligence Laboratory.

It contains information across various fields, including history, law, current affairs, and medicine.

However, concerns are also being raised that if Chinese-made data is utilized for AI training worldwide, the perspectives and values of the Chinese government could spread along with it.

Wanjuan itself explicitly states that it contains information aligned with "China's mainstream values."

Experts believe that beyond competing in the performance of AI models themselves, China has fully jumped into the "data competition" that determines what AI learns and what information it bases its answers on.

Reported by Kim Minjeong | Video by Lee You-jin | Graphics by Lee Jung-joo | Produced by SBS Digital News

※ Please note: This article was translated by AI and may contain errors.
Copyright Ⓒ SBS & SBSi. All rights reserved.
Copying, redistribution, and unauthorized use in AI training are strictly prohibited.
Kim Minjeong View More Articles
AD
AD
AD
AD