China can't out-spend American AI labs, and it doesn't have access to the most advanced Nvidia chips. It's winning market share anyway, on price.
Chinese developers including DeepSeek, Alibaba and Moonshot have focused on models that perform close to leading U.S. systems without needing the most powerful hardware, and on releasing those models as open weights that anyone can run and modify. That's a direct challenge to the American approach of charging premium prices for closed, proprietary models.
The Price Gap Is Enormous
The numbers make the strategy obvious. A million output tokens costs roughly $50 on Anthropic's Fable model, according to Fortune. The same volume of output costs about 87 cents on DeepSeek-V4-Pro and around $4.40 on Z.ai's GLM-5.2.
On input tokens, DeepSeek V4 Flash lists at 14 cents per million, against $1.75 for GPT-5.2. Some head-to-head comparisons run close to 98% cheaper on the lowest-cost pairings.
American Companies Are Already Switching
This isn't staying confined to developer forums. Reporting from CNBC found U.S. firms routed as much as 46% of their traffic on the OpenRouter platform to Chinese open models at points in 2026, up from roughly 11% over the prior year and just 4.5% in the first half of 2025.
The switching is showing up at real companies, not just in aggregate statistics. Coinbase CEO Brian Armstrong said publicly that pushing employees toward Kimi and GLM roughly halved the company's AI spending. Cursor built its Composer 2 coding model on top of Moonshot's Kimi. Airbnb has used Alibaba's Qwen for customer service.
"Price is doing the work here," said Harpreet Arora, head of agentic infrastructure at Vercel.
China's Bet on Giving the Software Away
The strategy behind the discount is deliberate. Chinese developers are treating AI software as something to commoditize rather than protect, reducing their dependence on the advanced chips U.S. export controls have kept out of reach.
Moonshot pushed that bet further in July, releasing Kimi K3, a 2.8 trillion parameter model the company describes as the largest open-source model available, at roughly $3 per million input tokens and $15 per million output tokens, still a fraction of top U.S. pricing.
The open-weight structure matters as much as the price. A company that downloads DeepSeek's or Qwen's weights isn't locked into a single vendor's API, pricing changes, or usage policies the way it would be with a closed model. That portability is part of the appeal independent of cost, particularly for companies wary of building critical infrastructure on top of a single U.S. lab's roadmap.
The Security Question Hasn't Gone Away
None of this erases the concerns that have followed these models since their release. U.S. federal agencies and several states have banned the DeepSeek app specifically on government devices over data-security concerns, and those restrictions extend to some other countries as well.
The bans target the China-hosted consumer apps, not the open-weight models themselves, which businesses can download and run entirely on their own infrastructure. That distinction is exactly what's fueling enterprise adoption, self-hosted deployments keep data off Chinese servers while still capturing the cost savings, though it does nothing to resolve the broader national security debate over relying on Chinese-engineered AI systems, the same tension at the center of the Trump family's WorldClaw venture.
What This Means for Miami
Miami's cost-conscious startup and enterprise AI community has direct exposure to this shift. Companies here already managing tight AI budgets, a pressure MAIN has covered as the gap between what AI actually costs to run and what businesses can afford, now have a genuine lower-cost alternative for workloads that don't require frontier-level performance.
The tradeoff is compliance risk, not just technical risk. Miami businesses in finance, legal or healthcare considering Chinese open-weight models should treat self-hosting as a requirement rather than an option, and should expect clients and regulators to start asking which models power their AI tools, not just how well those tools perform.
