Parallax Changelog
Issue 05 · live Subscribe
Menu
← Changelog № 05 · June 04, 2026

LLM PRICING

Uber's year of AI money lasted four months

You keep hearing AI is getting cheaper. Uber burned its whole 2026 AI budget in four months, and capped every engineer.

The primer

A token is the unit AI bills by: a piece of a word. Its price keeps falling, but coding tools now burn millions of tokens per task. Here is where the money went.

PublishedJune 04, 2026
Reading time4 min
Sources7 cited
DeskChangelog
The dekThe price per unit fell. The units per job exploded.
— 01
THE ARITHMETIC

Cheaper tokens, bigger bills

Uber exhausted its entire 2026 AI budget in four months and capped each engineer.

What most people think

AI gets cheaper every year, so a company's AI bill in 2026 should be smaller than its bill in 2025.

What the data shows

4months

Uber set aside a whole year of AI money for 2026 and ran through all of it by April, then capped every engineer's spending.

Cheaper per unit, many more units per job. That hisaab only goes one way.

Uber exhausted its entire 2026 AI budget in four months and capped each engineer.

In plain terms Two panels: on the left what most people assume, on the right what the numbers show, with the one figure that settles it.Source · Bloomberg (Natalie Lung) via Simon Willison's Weblog · 3 June 2026

— 02
THREE WORDS

Only one of these is your bill

Look at your electricity bill. There is a rate per unit, and there are units used.

token
A piece of a word. It is the unit AI companies price their work by.
per-token price
What one of those pieces costs. In 2023 the cheap models started at $0.25 per million, and capability that cost dollars then now costs cents.
cost per task
The price of one piece times how many pieces the job used. On your electricity bill, this is the bottom line.

Source · Simon Willison's Weblog · 4 June 2026

— 03
TOKENS PER TASK

One job, over a million tokens

So how many of those pieces does one job use? A chat reply is a few thousand of them. One measured coding job used over a million.

How to read this — Each step up the side is ten times the step below it, so a small-looking gap on this chart is a large one. Read the height of each dot. Press Linear for the evenly spaced view.

One measured coding task used 1,363,250 tokens. A chat reply uses a few thousand. tokens consumed (one task)

How to read this — Each step up the side is ten times the step below it, so a small-looking gap on this chart is a large one. Read the height of each dot. Press Linear for the evenly spaced view.

3,192 2e4 8e4 4e5 2e6 0.92 1.21 1.5 1.79 2.08 0 3.7e5 7.4e5 1.1e6 1.5e6 0.92 1.21 1.5 1.79 2.08 one chat reply*GPT-5 Codex task (measured) one chat reply*GPT-5 Codex task (measured) task autonomy → tokens consumed (one task) (log) The only measured point here. Onetask, 1.36 million tokens. The only measured point here. Onetask, 1.36 million tokens.

Log y-axis · equal steps are equal multiples. Top value is 273× the bottom.

One measured coding task used 1,363,250 tokens. A chat reply uses a few thousand.

In plain terms Each dot is one kind of task, placed by the number of tokens it used.Source · Simon Willison's Weblog, llm-pricing. Lower point illustrative, not measured · 4 June 2026

— 04
WHAT CHANGED

The software changed jobs

In November 2025 the software stopped answering questions and started doing whole jobs. A job is a loop.

1
The tools got good

In November 2025, two coding models shipped a week apart. Coding agents went from often working to mostly working.

2
A job is a loop

The agent reads your code, tries a change, runs the tests, reads the failure, and tries again.

3
It re-sends everything

The developer Simon Willison, who tracks AI pricing, writes that agents "maintain state by replaying entire conversations with each new prompt".

Source · Simon Willison's Weblog · 19 and 27 May 2026

— 05
INSIDE ONE TASK

Most of it was re-reading

So the loop re-sends everything it already has, on every step. Here is one real coding task, split three ways. Cached means the pieces the agent had already sent once.

One measured task used 1,363,250 tokens. Of those, 1,176,320 were cached and 17,112 were new code.tokens · one task
cached tokens 1176320 Most of the task was context re-sent, not new code.
input sent 169818
output written 17112

One measured task used 1,363,250 tokens. Of those, 1,176,320 were cached and 17,112 were new code.

In plain terms Each bar is one slice of a single task's token bill, longest first.Source · Simon Willison's Weblog, llm-pricing · 4 June 2026

— 06
HOW IT HAPPENED

Billing changed, prices rose, budget gone

Almost all of that task was context the agent re-sent. Now watch what the billing did next. Frontier means the newest and most capable model on sale.

Nov 24 2025
Coding agents start mostly working.
Nothing got more expensive here. The tools just started working.
Apr 2 2026
OpenAI starts billing Codex by the token.
Priced on token usage now, not per message.
Apr 14 2026
Anthropic starts billing the seat by usage.
A seat is one person's licence. About $20 a month, plus what you use.
Apr 23 2026
New frontier model, twice the price.
At the frontier, the per-token price is now rising.
~Apr 2026
Uber's 2026 AI budget is exhausted.
Four months in, the whole year's money is gone.
Jun 2 2026
Uber caps engineers at $1,500 per tool.

In plain terms The key events in order down a rail, with coloured dots flagging the pivotal moments.Source · Simon Willison's Weblog · 19 May to 3 June 2026

— 07
THE CAP

One engineer, one tool, one month

So the cap arrived on 2 June. It covers only the tools that write and run code on their own.

The cap covers agent tools like Cursor and Claude Code, not chat assistants.

Per engineer, per tool

$1,500/month

A median Uber US engineer's total pay is around $330,000 (about ₹3.1 crore), an outside estimate, not a company figure.

  • About ₹1.4 lakh a month, for one person and one tool.

  • A year of that cap, on two tools, is roughly $36,000. Assuming an engineer uses two tools.

The cap covers agent tools like Cursor and Claude Code, not chat assistants.

In plain terms One number, large, and beside it the everyday things it equals, so the size can be felt rather than read.Source · Bloomberg (Natalie Lung) via Simon Willison's Weblog. Levels.fyi for the pay estimate. Converted at about ₹95 to the dollar, September 2026 · 3 June 2026

— 08
THE BILL ELSEWHERE

Where else the money moved

So one engineer with two tools can spend $36,000 a year. Here are four more readings, taken off other meters.

The token bill, four readings
$2,180
Tokens one heavy user burned in 30 days
Valued at the advertised price. The subscriptions behind it cost $200.
Newest top model's price against the last one
Shipped 23 April 2026. At the top end, prices are rising.
~1.4×
Cost rise with the price per piece unchanged
The April update counts the same text as more pieces.
~11 %
Of a median engineer's estimated total pay
Two tools a year, against an outside estimate, not Uber's figure.

The token bill, four readings

In plain terms The headline numbers from the story, each with its unit and a one-line note.Source · Simon Willison's Weblog. Levels.fyi for the pay estimate · 20 April to 3 June 2026

— 09
WHAT ACTUALLY MOVED

The work changed shape

Those four numbers point the same way. One price rose in April, at the top of the market. Where the price per piece held still, the bill rose, because the work changed shape, and a business pays for work. Think of it like your electricity bill again. The rate per unit fell, which means very little when one job went from a few thousand units to over a million. Writing new code is close to free now. Getting good code is still expensive, and that is what the meter counts.

Source · Simon Willison's Weblog

Reader signal

How did this land?

A private signal to the editor. Pick as many as apply — nobody else sees it, and it shapes what runs next.

Letters

Sixty days of argument.

Longer responses from readers. The editor reads each one before it appears here.

No letters yet Newest first

Be the first to write one. Letters open for sixty days after an issue is published.

Write a letter

0 / 2000 · reviewed before it appears

Sources

7 cited.

Every number above traces to one of these. Primary sources first.

  1. 01 Uber Caps Usage of AI Tools Like Claude Code to Manage Costs simonwillison.net Simon Willison's Weblog (relaying Bloomberg / Natalie Lung) · primary
  2. 02 I think Anthropic and OpenAI have found product-market fit simonwillison.net Simon Willison's Weblog · primary
  3. 03 The last six months in LLMs in five minutes simonwillison.net Simon Willison's Weblog · primary
  4. 04 Simon Willison — llm-pricing tag index simonwillison.net Simon Willison's Weblog · primary
  5. 05 Claude Token Counter, now with model comparisons simonwillison.net Simon Willison's Weblog · primary
  6. 06 How coding agents work (Agentic Engineering Patterns) simonwillison.net Simon Willison's Weblog · analysis
  7. 07 Writing code is cheap now (Agentic Engineering Patterns) simonwillison.net Simon Willison's Weblog · analysis
0% · 4 min left
Press S

One email. Two things.

The dispatch when each issue ships — and a free shelf that saves your place. One click in your inbox does both.

Enable JavaScript to subscribe.