Parallax · tech desk · — 28

Open models are four months behind, not years

Kimi K3 took 84 days to reach the frontier OpenAI had in April. You can download it free. Almost nobody has.

Two measurements, one answer

Four months behind, and the gap has stopped shrinking.

What most people think

Open models are hobby projects. The paid ones are a generation ahead, and closing that gap will take years.

What the data shows

4months

Epoch AI, a research group that scores models on hard tests, measures the best open model four months behind the best closed one.

A second measure asks which model people prefer. The closed lead there grew from 0.5% to 3.3%.

WHAT YOU ASSUME

Two readings of the same gap, seven months apart. In between, the day an open model caught April's best closed one.

2025-10-30
Epoch AI reads the gap
2026-04-23
GPT-5.5 · ECI 158.22
2026-05-29
Epoch AI reads it again
2026-07-16
Kimi K3 · ECI 158
April's frontier, reached 84 days later
2026-09-03
GPT-6 Astra · ECI 166

FIVE DATES

Eight points ahead on the index, four months ahead in time. Those are two readings of the same gap: the points say how far behind, the months say how long the catch-up took. Only these four tests score both models on the same terms.

Near parity on graduate science. On a puzzle game the open model scores 26% against 84%. Epoch AI warns this reading may understate the gap, because open models are tuned hard on public tests.points
Mystery Game Puzzles Open 26% · closed 84% 58 Fifty-eight apart on a puzzle game
Chess Puzzles Open 39% · closed 72% 33
SimpleQA Verified Open 51% · closed 76% 25
GPQA Diamond Open 93% · closed 96% 3 Three points apart on graduate science

INSIDE THE AVERAGE

Those gaps run from three points to fifty-eight. So who releases these open models? Three Chinese labs: DeepSeek, Moonshot, which makes Kimi, and Zhipu, which makes GLM. Each has one column below, and a line joins every model to the one built from it.

Eight of these ten releases are open weights, and the newest landed in July.history
  • deepseek-v3-2 DeepSeek · V3.2 open
  • deepseek-v4-pro DeepSeek · V4-Pro, MIT licence open
  • deepseek-v4-flash DeepSeek · V4-Flash, MIT licence open
  • kimi-k2-thinking Moonshot · Kimi K2 Thinking open
  • kimi-k2-6 Moonshot · Kimi K2.6 open
  • kimi-k3 Moonshot · Kimi K3, 2.8T open
  • glm-5 Zhipu · GLM-5 open
  • glm-5-1 Zhipu · GLM-5.1 open
  • gpt-5-5 OpenAI · GPT-5.5 closed
  • gpt-6-astra OpenAI · GPT-6 Astra closed

WHO IS ACTUALLY THERE

Eight of those ten releases are free to download. Nothing under them is, so the rest of the stack stays somebody else's bill, and that kharcha does not come with the file.

Training data is almost never released, and the compute to serve a model is not downloadable at any licence.architecture
Weights Downloadable today. Kimi K3, DeepSeek V4, Qwen. The layer everyone means.
Licence Permissive for 81% of Chinese releases above 20B. Only 29% of American ones.
Code Inference code usually. Training code rarely.
Training data Almost never released. You cannot see what the thing was trained on.
Compute and serving Not downloadable at any licence. The cluster is the real bill.

THE WHOLE STACK

The one layer you can rent has a price. Here it is, per million tokens, which means the small chunks of text a model reads and writes.

A million output tokens costs $50 from the closed frontier and $15 from the best open model.

GPT-6 Astra, the closed frontier

$50per million output tokens

Ask the same question in an Indian language and it costs about five times the English price.

  • The same million from the best open model, $15 Less than a third of the frontier price

  • From DeepSeek V4-Flash, $0.28 The cheapest open model on the list

WHAT IT COSTS

That's the short version.

Full issue · 8 sections · 13 sources, all linked

Read the full issue →