AI RadarWe read first, then explain what changed
AI leaderboard watch

K3 106, Grok 111: this AI leaderboard is starting to contradict itself

Last updated 2026-08-20Editorial synthesis: signals connected before judgementNot a wire dump; facts, judgement, and unknowns are separated
Original diagram comparing K3, Grok, and DeepSeek across harnesses, real tasks, IQ, time, and cost
Editorial diagram: Codex Radar's IQ is a custom display scale; the community benchmark keeps refreshing, and public-beta scores do not enter formal recommendations.
Bottom line

Codex Radar compares more than models: it exposes harnesses, reasoning levels, real-task completion, time, and cost. K3, Grok, and DeepSeek scores change live, and beta numbers should not be treated as final crowns.

Codex Radar compares more than models: it exposes harnesses, reasoning levels, real-task completion, time, and cost. K3, Grok, and DeepSeek scores change live, and beta numbers should not be treated as final crowns.

This essay is currently written in Chinese.

Evidence and originals