Parallax · tech desk · — 18

The Bill Came Due in April

AI keeps getting cheaper per token, and that is exactly why the bill goes up. The price per unit fell, the units per task exploded, and Uber burned its whole 2026 AI budget in four months.

Everyone agrees on the sentence: AI is getting cheaper. Look closer and it is true at the unit level and false at the invoice level, and that gap is the whole story. A scissor has two blades. Watch both.

The blade everyone watches
Price per token keeps falling.
Frontier capability that cost dollars per million tokens in 2023 now costs cents. Claude 3 Haiku launched at $0.25 / $1.25 per million input/output tokens; today's small models sit lower still — GPT-5-mini at $0.05 / $0.40, Grok 4 Fast at $0.20 / $0.50. The rate card has fallen more than an order of magnitude in three years. The corollary feels obvious: an AI budget should be shrinking.
The blade that cuts the budget
Tokens consumed per task climb faster.
The tools stopped being things you talk to and became things that work for you. A coding agent re-reads the whole codebase on every step, calls tools, and runs reasoning loops, so one task can burn well over a million tokens where a chat reply burns a few thousand. Multiply a unit price that halves each year by a unit count that climbs an order of magnitude, and the product, the actual invoice, goes up. Uber budgeted 2026 against last year's cheaper-per-token mental model and exhausted the entire annual figure in four months.

THE SCISSOR

Around November 2025, coding agents crossed from often-work to mostly-work and became daily drivers for highly paid engineers. The shift from a tool you talk to into a worker that runs for you is what changed the token math, because an agent that works autonomously burns vastly more tokens than a chat that answers a question.

THE CROSSOVER

The other blade, plotted. One of Willison's GPT-5 Codex tasks consumed 169,818 input, 17,112 output, and 1,176,320 cached tokens: well over a million for a single task, and the only verified point on this chart, sitting at the top. The vertical axis is logarithmic because that anchor sits three orders of magnitude above the bottom. The lower points trace the shape from a single chat turn upward and are illustrative of task class, not measured.

Tokens per task · chat turn → agentic session tokens consumed (one task)

How to read this — The five points climb in an even line because each step up this axis is a multiplication rather than an addition. Press Linear for the proportional view, where the lower four crowd into the bottom fifth while the measured task stands far above them; both plot the same five numbers.

3,192 2e4 8e4 4e5 2e6 0.68 1.84 3 4.16 5.32 0 3.7e5 7.4e5 1.1e6 1.5e6 0.68 1.84 3 4.16 5.32 one chat reply*short Q&A thread*single-file edit*multi-step refactor*GPT-5 Codex task (measured) one chat reply*short Q&A thread*single-file edit*multi-step refactor*GPT-5 Codex task (measured) task autonomy → tokens consumed (one task) (log)

Log y-axis · equal steps are equal multiples. Top value is 273× the bottom.

TOKENS PER TASK

Read top to bottom, the sequence explains itself. The tools got good, so the pricing flipped to metered usage; once usage was metered, the frontier price started rising; and an enterprise budget set against the old mental model ran out a third of the way into the year.

Aug 7 2025
GPT-5 sets the frontier price.
Nov 24 2025
The November inflection.
Apr 2 2026
OpenAI moves Codex to token metering.
Apr 23 2026
GPT-5.5 ships at 2× the price of GPT-5.4.
~Apr 2026
Uber's 2026 AI budget is exhausted.
Jun 2 2026
Uber caps every engineer at $1,500 / month per tool.

HOW THE BILL CAME DUE

The rate card is one number; the realised cost is another. Over thirty days, at standard interface rates, Willison ran up more than two thousand dollars of tokens. He paid two hundred. That gap is what keeps the true cost off the individual's own statement, which is precisely why finance is the last to find out. The dashed line marks what he actually paid.

One heavy user · 30-day token cost at API rates vs. subscription price paidUSD · 30 days
what he actually paid (subscriptions) · 200
Tokens billed at API rates (total) 2180.16
Anthropic Claude Code 1199.79
OpenAI Codex 980.37
Actually paid (Max + Pro) 200

THE REAL PRICE OF A SESSION

Uber is the first named company to run into the scissor in public. The cap is a ceiling, drawn around a line item that, left uncapped, ate a year's budget by April. The figures, set side by side.

UBER · AI CODING SPEND · 2026
$1,500 /mo
Per engineer, per agentic tool
Separate budgets per tool. Applies to Cursor and Claude Code — agents, not chat assistants.
4 months
To exhaust the entire 2026 AI budget
Set against last year's cheaper-per-token, lighter-usage reality.
~11 %
Of a median engineer's total comp
Two tools ≈ $36,000/yr against ~$330,000 comp. Willison's arithmetic on Levels.fyi data, not an Uber disclosure.
$2,180
Tokens one heavy user burned in 30 days
At API rates, for a $200 subscription. The cost the individual never feels.

$1,500 / MONTH

That's the short version.

Full issue · 7 sections · 7 sources, all linked

Read the full issue →