The next phase of the global AI race may not be decided by who builds the biggest model.
It could be decided by who controls the data those models learn from.
China is moving aggressively to become a major supplier of AI training data, with Beijing setting a goal of transforming the country into a global data powerhouse by the end of 2028.
The strategy goes beyond helping Chinese companies build better AI. It could also give China greater influence over what AI systems around the world know and how they respond to questions about the country.
The Race For Better Data
China's National Data Administration has called for high-quality datasets across areas including scientific research, manufacturing and autonomous vehicles.
Beijing also wants those datasets shared internationally, including with developing countries building their own AI systems.
The economic logic is straightforward. Better data can produce better models, while wider distribution of Chinese datasets could encourage countries and companies to build parts of their AI infrastructure around Chinese technology.
But there is another dimension.
China's AI companies operate under rules requiring them to follow official government narratives. That creates concerns about what happens when information produced by China's state-controlled media becomes part of datasets used to train global AI systems.
When Data Carries A Narrative
Recent research cited by The New York Times found that Chinese state narratives are already appearing in the training data of American AI models.
Researchers found that some models gave substantially more favorable responses about China's institutions and leaders when questions were asked in Chinese rather than English.
That does not mean every Chinese dataset is propaganda, or that every AI system trained on Chinese data will produce politically biased answers.
It does demonstrate something more important: training data is not neutral simply because it is data.
The source, language and institutions behind that data can influence what an AI system ultimately produces.
China's WanJuan dataset illustrates the scale of the strategy. The state-backed collection covers areas including history, medicine, law, literature and current events and is designed around what its creators describe as "mainstream Chinese values."
What This Means For Miami
Miami is increasingly positioned between the United States, Latin America and the Caribbean, making the city an important hub for companies operating across different technology and information environments.
That makes the global AI data race relevant locally.
Miami startups building AI products for international markets may increasingly need to understand not only which models they use, but where those models' training data comes from and what assumptions are embedded within it.
For investors, the issue creates another layer of AI infrastructure to watch. Data ownership, provenance, quality and geographic origin could become strategic considerations alongside computing power and model performance.
Miami's role as an international technology and business hub could make it a natural base for companies working on trusted data infrastructure, model evaluation and AI governance.
The AI race is becoming a race over information itself. For Miami, that could create opportunities well beyond building another chatbot.
Reporting Source: This article builds upon reporting from The New York Times and adds analysis of what the development means for Miami and South Florida.
