China’s leading AI models have narrowed their reported performance gap with US rivals to about 3%, while several Chinese models are now appearing near the top of global rankings. The shift is not just about DeepSeek: Alibaba, Moonshot AI and Z.ai are also gaining visibility, and DeepSeek V4.1 Flash recorded 85 trillion tokens of use on OpenRouter over 30 days, according to China Daily’s account of Bloomberg Intelligence data and other sources.
Watch Desk analysis
What happened
China Daily says Bloomberg Intelligence measured the gap at about 9% in May and 15% earlier this year. The article links the narrowing to DeepSeek’s September release of V4.1 Flash and says models from Alibaba, Moonshot AI and Z.ai are also placing near the frontier on Artificial Analysis tests. Some scored above DeepSeek on that platform’s composite tests.
The article also cites OpenRouter usage figures: DeepSeek V4.1 Flash handled about 85 trillion tokens over the 30 days through Wednesday, making it the platform’s most-used model. Z.ai’s GLM 5.3 Flash ranked third with about 55.2 trillion tokens, while other Chinese models also appeared among the top ten. China Daily’s report attributes the performance-gap figures to Bloomberg Intelligence.
Why it matters
A narrower benchmark gap and substantial use on a model-aggregation platform point to two different kinds of momentum: measured capability and developer uptake. Neither settles the larger contest. The article also notes weaknesses in demanding reasoning and tool use, alongside challenges adapting models to China’s diverse chip and software systems.
Nor does token volume tell us how many people use a model, whether they stick with it, or how well it performs on a particular job. But the figures make the AI race look less like a single challenger chasing a fixed US lead and more like a widening field of Chinese developers competing for users.
Our read
The headline number is striking, but the more interesting development may be the breadth behind it: DeepSeek is no longer the only Chinese name near the top, and usage is spreading across several models. Treat the 3% figure as a reported comparison, not a universal scorecard. Benchmarks depend on what is measured; usage depends on who is counting what.
What to watch
- How Bloomberg Intelligence defines and updates its performance-gap comparison.
- Whether Chinese models sustain their rankings and OpenRouter usage over time.
- How reasoning, tool use and hardware-software compatibility affect real-world adoption.
Discussion spark: When judging the AI race, should benchmark performance or actual developer usage carry more weight, and what evidence would change your mind?
Sources and evidence
- China narrowing AI gap with US – chinadailyhk (9 October 2026, 02:38 UTC)
Watch Desk is operated by WittyWires as an independent cross-cutting AI news tracker. It does not speak for the organisations or people it covers.