AI RadarWe read first, then explain what changed
China's model price war

DeepSeek has not fallen behind. China's model race is crowded, and GLM's sharpest weapon is price

Last updated 2026-08-27Editorial synthesis: signals connected before judgementNot a wire dump; facts, judgement, and unknowns are separated
Original illustration of a GLM-5.3-Flash model tower connected to image, PDF and agent capabilities in a crowded model market
AI Radar original AI illustration: the image represents GLM in a crowded field; current benchmark, price and speed figures are documented in the article rather than implied as a permanent ranking.
Bottom line

In the current like-for-like Artificial Analysis comparison, GLM-5.3-Flash scores 57 at about $0.10 per million blended tokens; DeepSeek V4 Pro scores 53 at about $0.69, while V4 Flash scores 52 at about $0.23 and approaches 119 tokens per second. DeepSeek has not fallen behind and GLM has not won every dimension. Quality, speed, modality and price have become four separate races.

Bottom line

DeepSeek has not fallen out of the front rank. What changed is that the gaps among leading Chinese models have compressed into a few index points, a few tenths of a dollar and a few dozen tokens per second. The old promise that one model could be the obvious answer for every job is becoming harder to defend.

In the current like-for-like comparison from Artificial Analysis, GLM-5.3-Flash scores 57, DeepSeek V4 Pro 0813 scores 53 and DeepSeek V4 Flash 0731 scores 52. The interesting part is not merely 57 versus 53. It is that each model makes a different trade: GLM accepts images and costs the least in the measured mix, DeepSeek Pro generates faster, while DeepSeek Flash runs at nearly two and a half times GLM's output speed.

Three numbers before declaring a winner

Illustrated comparison of GLM-5.3-Flash, DeepSeek V4 Pro and DeepSeek V4 Flash on index score, cost and speed
Editorial note: 57, 53 and 52 are aggregate index scores, not tokens per second. The illustration's roughly 125 t/s is an earlier rounded snapshot; the article uses the current measured figure of about 119 t/s.

Under the same Artificial Analysis methodology, GLM-5.3-Flash has an Intelligence Index of 57, measured output of roughly 50 tokens per second and a blended price of about $0.10 per million tokens. DeepSeek V4 Pro records 53, about 68 tokens per second and $0.69. DeepSeek V4 Flash records 52, about 119 tokens per second and $0.23.

That $0.10 is not simply GLM's public input price. Artificial Analysis calculates a blended cost with a 7:2:1 mix of cached input, uncached input and output. GLM's listed input and output prices are approximately $0.15 and $0.50 per million tokens. A workload with long answers or few cache hits can therefore cost materially more than the headline blended figure.

A four-point advantage cannot honestly be translated as domination, just as 119 tokens per second cannot be translated as the best answers. Quality, speed, modality and price answer four separate questions.

GLM's sharper weapon is the cost curve

GLM-5.3-Flash has 320 billion total parameters while activating 18 billion for each token. It combines image input, a one-million-token context window and MIT-licensed open weights in the same Flash model. Z.ai says the model was pretrained on 30 trillion multimodal tokens and uses sparse plus linear attention to reduce the cost of long contexts.

Those architectural descriptions still come from the vendor and should not be treated as independent proof. The outside comparison does, however, establish a useful current snapshot: for the providers and mixed workload tested, GLM obtained a higher aggregate score at roughly one seventh of DeepSeek Pro's blended price. For teams running large volumes of agent, document and image work, that economic difference matters more than four positions on an index.

First reversal: DeepSeek Flash is more than twice as fast

If the workload is code completion, bulk classification, structured extraction or any interaction in which a person is waiting at the screen, DeepSeek V4 Flash remains difficult to dismiss. Its aggregate score is five points below GLM and its blended cost is higher, but its measured output speed is about 119 tokens per second, roughly 2.3 times GLM's rate.

That is why the claim that GLM won and DeepSeek became obsolete does not survive contact with the data. GLM wins the current combination of aggregate score and measured cost. DeepSeek Flash wins on waiting time. In a frequently used interactive product, latency is itself a feature rather than an implementation detail.

Second reversal: DeepSeek Pro has the harder case

DeepSeek V4 Pro is more awkward to explain. It scores only one point above its own Flash sibling, while measured speed falls from roughly 119 to 68 tokens per second and blended price rises from $0.23 to $0.69. It may be more reliable on particular difficult jobs, but the public aggregate score does not show a universal advantage proportional to the threefold cost.

This does not make Pro useless. It means buyers should no longer pay for the word Pro without evidence. A team already using it should run the same repository tasks, long-running jobs and known failure cases against Flash. Product naming is a poor substitute for an evaluation and an especially easy way to overspend.

The front rank is no longer a fixed seating chart

Original illustration of Claude Opus 5, GPT-5.6 Sol, Claude Fable 5 and Grok 4.6 as neighboring model towers
Original AI illustration: this represents a crowded frontier, not a fixed order. Rankings move with benchmark versions, reasoning settings and harnesses.

The upper part of the current Artificial Analysis leaderboard also includes Claude Opus 5, Claude Fable 5, GPT-5.6 Sol, Grok 4.6 and Kimi K3. Results move with the benchmark version, reasoning setting, agent harness and serving provider. A gap of four or five points is far less stable than the kind of generational gap that once separated model families.

An aggregate leaderboard also compresses coding, mathematics, knowledge and agent work into one number. A legal-retrieval team, a frontend agent and a Chinese customer-support operation do not inhabit the same ranking. Front rank is useful as the name of a competitive group. It is much less useful as a permanent order from first to tenth.

How users and developers should choose

For images, PDFs, long documents and low cost, GLM-5.3-Flash is the first candidate to test. For high-throughput text work and shorter waits, begin with DeepSeek V4 Flash. Teams already paying for DeepSeek Pro should put Flash into the same production-shaped evaluation and ask whether a one-point aggregate difference is worth roughly three times the blended price.

An enterprise evaluation can be simple without being casual. Assemble 20 to 50 representative tasks and preserve the prompts, expected outcomes, human rework time, time to first token, total duration and actual bill. Public rankings should reduce the field. Only an internal test should decide what gets purchased or routed into production.

Our judgement

The most important part of GLM-5.3-Flash is not that it pushed DeepSeek off the table; it did not. The change is that the boundary of what buyers expect from a cheap model has moved again. Native multimodality, a one-million-token context and open weights are beginning to arrive together with extremely low API prices.

DeepSeek remains in the front rank, and the speed advantage of its Flash model is substantial. What it no longer has is an answer so obvious that users can stop comparing. The Chinese model contest has moved from who can catch up to a more practical fight: once several systems are good enough, which one is cheaper, faster and easier to deploy for a specific workload?

What remains to be seen

First, we need to know how long GLM's low price can be sustained and whether it applies consistently across major providers and peak periods. Second, self-hosters need real numbers for memory, parallelism and engineering overhead around the 320B-total, 18B-active architecture. Third, multimodal-agent success rates, recovery from failures and stability on long tasks need replication by independent teams.

The next DeepSeek release or price adjustment could quickly redraw the table. Every figure in this article is a verifiable snapshot dated August 27, 2026, not a permanent ranking.

Sources and evidence boundaries

Parameter counts, licensing and architectural descriptions come from Z.ai's launch material and the Hugging Face model card. Aggregate scores, output speed, context windows and blended costs come from the current Artificial Analysis model pages and comparison methodology. Vendor benchmark claims are treated as product claims rather than substitutes for independent measurement.

Prices can change with provider, cache ratio, input-output mix and promotion. Aggregate scores also move when a leaderboard is updated. The recommendations above are editorial judgements intended to guide testing, not guarantees for every workload.