Learner LabsLearner 1.0Weight-based learning without forgetting

Most of the bill is re-reading

Ninety-six percent of the input tokens a coding agent pays for are context it has already read. Learning the material instead shrinks the prompt by a median of 52 times. The bill falls by about eight, and the reason that gap exists is arithmetic rather than engineering.

A language model that answers a question about your codebase has to be told about your codebase first. Today that telling happens on every single turn. The material is retrieved, or pasted, or replayed from a cache, and it arrives in the prompt again, and again, and again. In the best measured trace of real agentic coding that exists in public, 4,265 developer sessions and 357,161 model steps, the total came to 54.90 billion input tokens against 186.9 million output tokens. That is 293.7 tokens in for every token out, and 52.56 billion of the input tokens, 95.7 percent of them, were context being re-read rather than anything new.

This is not a bug. It is how the interface works: a model has no memory between calls, so the only way to give it your material is to send your material. But it is the largest single line in the bill, and it is the line a model that can actually learn the material should be able to remove. We build such a system, so we have an interest in the answer being large. This piece works out how large it honestly is, using published figures where they exist and saying so where they do not.

The question is: how much of what the industry spends on inference is spent re-reading things the model has already been told, and how much of that does learning the material instead actually save?

Conclusions

  1. The re-reading claim is true of tokens, not of revenue. Ninety-five point seven percent of the input tokens in the measured trace are re-reads. The share of the actual bill is 59.5 percent, because caching already discounts most of them. Both are worth knowing and they are not the same number. See the section.
  2. Two figures in circulation do not survive checking. There is no sourced two hundred billion dollars of laboratory revenue: the only audited annual figure in the sector is OpenAI's 13.07 billion for 2025, and even the naive sum of every lab's run-rate, a metric that overstates, reaches about 120 billion. There is no sourced six trillion of valuation: the private labs come to about 2.2 trillion of paper value, and NVIDIA alone is 5.16 trillion. See the section.
  3. For a question whose answer has been learned, the prompt is 52 times smaller. Median over 68 previously graded questions against a retrieval pipeline on the same source, with a distribution-free interval of 42 to 56. Against the whole document in context it runs from 29 to 2,133, depending on document size. See the section.
  4. The bill cannot fall as fast as the prompt, and the ceiling is about nine times. Output is 11.2 percent of an agentic bill and no input-side technique touches it. Our 52 converts to about 7.8 times on a whole bill, which is 87 percent of the theoretical maximum for any technique of this kind. See the section.
  5. So the answer to "a hundred times cheaper" is no. Reaching a hundred would need output tokens to fall by a further factor of eleven. We have not measured output tokens, which is the largest gap here and the one measurement we specify before publishing anything stronger. See the section.
  6. Learning pays for itself in tens of questions against a large context, and in hundreds to low thousands against a good retrieval stack. Teaching one 706-word document costs about fifty cents to two dollars of compute. Break-even is the number a buyer actually needs, and it is not flattering in every regime. See the section.

How big the pot is, when you only count what is disclosed

Before asking what share of the money is re-reading, it is worth establishing how much money there is, because the figure most often repeated is not a revenue figure. Figure 1 puts the two quantities side by side.

OPENAI 2025, BILLIONS OF DOLLARS 5 10 15 20 0 13.07 audited 2025 revenue ~20 end-2025 annualised run-rate the run-rate ran ~50% ahead of what was booked
The number usually quoted runs about fifty percent ahead of the number the same company actually booked. OpenAI's audited 2025 revenue was 13.07 billion dollars. Its end-of-2025 annualised run-rate, one recent month multiplied by twelve, was about 20 billion. Both are true statements about different things, and almost every laboratory revenue figure quoted in public is the second kind. Anthropic has no audited annual figure at all, for any year.

What the data says

The largest laboratories state their size as annualised run-rates: above 65 billion dollars for Anthropic at the end of July 2026 and above 40 billion for OpenAI in August. Both figures reached the press through investors rather than through filings, and the two companies may compute the metric differently, so the ordering between them is softer than it looks. The only fully audited annual figure in the sector belongs to OpenAI: 13.07 billion dollars of 2025 revenue against a 20.92 billion dollar operating loss. The revenue figure is about a third below its own end-of-year run-rate, which is the gap Figure 1 draws. A quarterly comparison of booked revenue also circulates, but we could not trace it to a source either company stands behind, so we do not use it. Adding every laboratory's run-rate reaches roughly 120 billion, a sum that mixes revenue booked gross with revenue booked net, recurring revenue with audited revenue, and February vintages with August ones, and should not be used. The best-documented large number in the sector sits on the other side of the ledger: Microsoft's fiscal-2026 annual report discloses 24.1 billion dollars of revenue from OpenAI, including revenue-sharing. What a laboratory pays its cloud provider is now audited. What any laboratory earns is not.

Two figures set the scale. Private valuations: Anthropic 965 billion dollars post-money, OpenAI 852 billion, xAI 250 billion, DeepSeek 52 billion, Moonshot 35 billion, Mistral at a last confirmed 11.7 billion euros. That is about 2.2 trillion dollars of paper value across the independent labs, against 5.16 trillion for NVIDIA alone. And 2026 capital expenditure guided by the four largest cloud providers comes to 720 to 745 billion dollars, which is about seven times the two largest labs' combined run-rate.

Why this is hard to pin down

Only NVIDIA has genuinely separable AI revenue. Gemini has no revenue line anywhere. Meta discloses two segments, neither of them AI, and is spending 130 to 145 billion dollars against zero disclosed AI revenue. xAI's revenue sits inside a segment its own filing defines as "our AI compute, Grok, and X", which in the prior full year was 57.6 percent advertising. Alibaba's model appears in filings only as a cost. Any table with a clean AI revenue column across all the majors is manufacturing several of its cells.

What it means for building

Size a market from the two companies that disclose, and treat the rest as unbounded. Be careful which total you quote: 73, 105 and 120 billion are all defensible and they describe different quantities.

Also true

The picture is circular. Analysts put Anthropic and OpenAI at roughly 73 percent of Amazon's AI revenue and above 70 percent of Microsoft's, while Amazon invested 50 billion dollars in OpenAI, Microsoft counts OpenAI's Azure spend inside its own AI run-rate, and NVIDIA invested 30 billion in OpenAI. Adding NVIDIA's data-centre revenue to cloud AI revenue to laboratory run-rates counts the same dollars two and three times.

How much of the compute is inference, and why five answers disagree

Every one of the figures below is real and none of them can be averaged with another, because each has a different denominator. We print the denominator in its own column, which is the only honest way to show them together.

SourceFigureDenominatorBasis
Gartnerinference 23.3 billion dollars against training 19.0 billion; 55 percent rising to 59 percentAI-optimised infrastructure spend onlyprojected
Deloitteabout two-thirds, up from a third in 2023all AI computeprojected
NVIDIAabout 40 percentNVIDIA's own data-centre revenuereported
Epoch AI1.8 billion against 5 billion, about 26 percentOpenAI's 2024 compute dollarsestimated

What the data says

2026 is the first year inference spending is projected to exceed training spending on the narrowest denominator, and the direction is the same on all of them. The most quotable of the four is the weakest: NVIDIA's 40 percent dates from May 2024 and has never been restated, and when the chief executive was asked to update it on the November 2025 earnings call he declined, saying it is hard to know what the percentage will be at any given moment.

Why this is hard to pin down

No published split of cloud capital expenditure into training and inference exists. One research firm models exactly this and keeps the output behind a paywall. The International Energy Agency states that none of the companies disclosing per-query metrics has offered a comparable breakdown, and its own report carries no training-versus-inference electricity split either.

Also true

Inference is real cost of goods. OpenAI's gross margin is estimated at 33 percent with inference costs of 8.4 billion dollars in 2025 rising to a projected 14.1 billion in 2026, and Anthropic cut a projected gross margin by about ten points after inference costs ran roughly 23 percent above plan. Both are third-party estimates. Neither company publishes the split.

How much of that is re-reading

This is the load-bearing section. The best measurement available is a trace of real agentic coding published in June 2026: 4,265 sessions from 43 developers, 357,161 model steps and 432,510 tool calls, recorded across two commercial coding agents and worth about 40,400 dollars of API-equivalent spend. It splits input tokens into two kinds. Append tokens are genuinely new material. Prefix tokens are the conversation so far, sent again. Figure 2 shows where the money goes once caching is applied.

WHERE ONE DOLLAR OF AGENTIC CODING SPEND GOES 29.2% 59.5% 11.2% genuinely new input context sent again output the part no input-side change can touch, which sets a ceiling of about 8.9× on the whole bill the largest single line in the bill
Six in every ten dollars of an agentic coding bill buy context the model has already read. The split is reported directly by the trace, after caching: 29.2 percent genuinely new input, 59.5 percent re-read context, 11.2 percent output. The re-read share is this large even though the trace achieves a 95.7 percent cache hit rate and cached reads are billed at about a tenth of the input rate.

What the data says

Of 54.90 billion input tokens, 52.56 billion were prefix and 2.34 billion append. That is 95.7 percent of input tokens being re-read, 293.7 input tokens for every output token, and only 12.5 to 1 counting genuinely new input alone. The reported medians per step are about 119,000 prefix tokens, 875 append tokens and 214 output tokens: roughly two orders of magnitude more prefix than append.

Two independent measurements land in the same place. A published account of a large porting project, 6,778 commits over eleven days with about 64 agents running concurrently, reports 5.9 billion uncached input tokens, 72 billion cached input reads and 690 million output tokens for about 165,000 dollars: 113 to 1, with 92.4 percent of input being re-read. And a token-marketplace study reports more than 85 percent of agentic token burn originating in cached prompts. Figure 3 puts those beside our own.

SHARE OF INPUT TOKENS THAT IS CONTEXT SENT AGAIN 0%25%50% 75%100% agentic coding trace our repository session token marketplace 4,265 sessions · measured 20 tasks, 400 files · measured platform-wide · estimated 95.7 84.3 85
Three measurements made by three different parties on three different workloads agree that most input tokens are re-reads. Our own is a 20-task simulated session over a 400-file open-source repository: 1,843,698 input tokens of which 1,469,096 were re-sends, 79.7 percent overall and 84.3 percent once the session reaches steady state, with only 6,261 genuinely novel tokens across all twenty tasks.

Why this is hard

Caching is supposed to solve exactly this, and it does not, because caching makes the re-read cheaper rather than unnecessary. Cache reads cost about a tenth of the input rate at OpenAI, Anthropic and Google, and the trace above already achieves a 95.7 percent hit rate. The re-read line is still 59.5 percent of the bill, because at 281 cached input tokens per output token a tenfold discount still leaves the largest line.

What it means for building

If your agent's bill is uncomfortable, the first thing to look at is not which model you pick but the fraction of your prompt that is unchanged from the previous call. One published case study moved a single dynamic identifier out of a cacheable prefix, took the hit rate from 7 percent to 74 and later 84, and cut cost by 59 percent without changing the model at all.

Also true

Agentic coding is the extreme end of this distribution, not the average. General traffic through a large model marketplace ran at roughly 15 prompt tokens per completion token in late 2025, a twentieth of the agentic ratio. And agentic coding is not all of AI revenue: token-metered API is estimated at 15 to 20 percent of OpenAI's revenue and 70 to 75 percent of Anthropic's. So the 59.5 percent cannot be lifted to an industry-wide share of revenue, and since no provider publishes an input-to-output split, that industry-wide number is not derivable from anything public. We do not derive it.

What we measured: the prompt for a question whose answer has been learned

Learner 1.0 takes a frozen base model, Qwen3.6-27B in bf16, adds about 1.14 billion trainable parameters, runs a short warm-up phase, and then trains only those parameters on whatever a user teaches it. The mechanism is proprietary and this piece describes none of it. It describes a measured consequence: once the material is learned, the prompt for a question about it is the question. Figure 4 is 68 previously graded questions from three demonstrations, each asked three ways.

PROMPT TOKENS, CONTROL ÷ LEARNED · SAME QUESTION · LOG SCALE 10×30×100× 300×1000×3000× handbook, 706 words reference page, 2.4 kB two corpora, 22,000 words each 24 questions12 questions32 questions 38.5 89.8 25.5 29.1 59.2 2133 28.1× live 47.7× live against retrieval against whole document
The learned prompt does not grow with what was learned, so the ratio grows with the size of the material. Bars are minimum to maximum across the questions in each lane, dots are the lane median, x axis is logarithmic because the quantity is a ratio. The two dashed lines are earlier single-question measurements taken live, at 28.1 and 47.7 times, both of which fall inside the retrieval band. The retrieval control is a BM25 pipeline taking the top four chunks of the same source. The whole-document control was computed from the corpora rather than served, because a 22,000-word corpus does not fit our 8,192-token serving window.

What the data says

Across all 68 questions the retrieval control pays a median of 51.9 times more prompt tokens, with an interquartile range of 32.4 to 61.7 and a distribution-free interval on the median of 42.0 to 56.0. Per lane the medians are 38.5 for the handbook, 25.5 for the reference page and 59.2 for the invented-language corpora. The whole-document control runs from a median of 29.1 on the smallest source to 2,132.9 on the largest, and we report it per lane rather than pooled, because the pooled interval runs from 89.8 to 1,987.2 and a range that wide is not a summary of anything.

14median prompt tokens after learning a 706-word handbook
26median prompt tokens after learning 22,000 words
590median prompt tokens for the retrieval control on the handbook
57,910median prompt tokens to put the 22,000-word corpus in context

Those four numbers are the whole of the ratio. The learned prompt is 11 to 37 tokens across every lane and does not grow: 14 after a 706-word handbook, 26 after 44,000 words of invented language. Every context-based route pays for the material on every question. Learning pays for it once.

The statistical statement, said precisely

This is a census of 68 previously graded questions over three corpora that we chose, not a random sample from a population of workloads. All 68 ratios exceed one, so a sign test returns a p-value near ten to the minus twenty-one, which is true and tells you nothing: the direction was never in question, because one arm's prompt is the question and the other's is the question plus the material. The quantity in doubt is the magnitude, and for strongly right-skewed ratios the right summary is the median with a distribution-free interval, the 42.0 to 56.0 above. It licenses a statement about these three corpora. It does not license one about workloads in general, or about dollars.

Replicate it

Teach the same handbook, ask the same question, compare the prompt you sent: teach_my_doc(document=...) then ask_learner(question=...). The document, the 68 questions, the three token counts each and the counter itself are in the data repository under token-efficiency/.

Also true

These are prompt tokens only. Completion tokens were not counted, on the assumption that they are similar across arms and would dilute every ratio equally. That assumption is reasonable and it is not a measurement, and the conversion from a token ratio to a bill ratio depends on it, which makes it the largest gap in this piece. The measurement that closes it is specified below.

At equal accuracy?

A token ratio means nothing if the cheap arm answers worse. This is the question we can answer least completely, so it is split into three buckets that must not be mixed: one live paired measurement of ours, our own pass rates, and the published literature on the alternatives. Figure 5 is the third bucket only.

PUBLISHED RATES · EACH POINT IS ITS OWN BENCHMARK, NOT A COMMON SCALE 0%25%50% 75%100% closed book, accuracy closed book, wrong answers 36 models, net positive score retrieval off, then on fabrication with document given purpose-built legal retrieval 55% 40% 3 of 36 28.6% 75.6% 1.8% to 24% 17% and 33% accuracy error before and after retrieval
Nobody has solved recall, in either direction. The best model on a short-fact benchmark without the web reaches 55 percent accuracy while giving a wrong answer 40 percent of the time. On a 6,000-question benchmark that penalises wrong answers symmetrically, only 3 of 36 models score above zero. Retrieval is a large improvement, moving one controlled study from 28.6 to 75.6 percent, and it leaves a stubborn residue: models still fabricate in 1.8 to 24 percent of summaries when the source document is supplied and internal knowledge is forbidden, and two commercial legal research tools built on curated databases still returned incorrect or unsupported citations 17 and 33 percent of the time in a human-adjudicated study. Each row is a different benchmark with a different definition, which is exactly why they are drawn as separate rows and never averaged.

Bucket one: our one live paired measurement

On one question, both arms were served and graded live on the same day. The retrieval control answered correctly using 858 prompt tokens. The learned arm answered correctly using 18, a ratio of 47.667. That is the only point in the whole distribution where control-arm correctness was measured rather than assumed, and n is one. It is reported precisely because it is the shape of measurement the rest of the distribution lacks.

Bucket two: our own pass rates, each on its own metric

What was taughtWhat was measured
a 706-word company handbook4 of 8 propositions answered on the first wording, 12 of 24 across three independent wordings. Every extracted row trained, so the misses were extraction rather than training. Policy and numeric facts passed both verbatim and paraphrased. The four that tie a name or a year to a specific organisation failed. The extraction fix shipped on 2026-08-25 and, on the identical document, extracted all four and trained them.
nine facts the base model believes are false9 of 9 confirmed unknown to the base model beforehand. 7 of 9 served the taught value afterwards, after a re-grade against ourselves. Real-world control questions held at 3 against 3. Unlearning one reverted exactly that answer and nothing else.
twelve facts, then one unlearned11 of 12 answered before the unlearning. The unlearned one reverted. Of the 10 others that were known, 9 held and 1 moved. A separate run with no unlearning at all showed the same one-row movement, so it is retrain-to-retrain variation rather than damage from the unlearning.
three topics taught in sequenceThe third topic scored 4 of 4 after its own teaching. The two earlier topics moved from 4 of 4 to 3 of 4, each by one row, inside the measured variation.
two invented languages, one after the otherHeld-out validation loss dropped from 4.54 to 2.72 and from 4.27 to 2.11 nats. Cross-language interference was 0.00 in every measured cell in both directions. English controls held 4 of 4 at every stage. A standard adapter of 0.97 times the same trainable parameter count, on identical bytes in identical order, went 4.345 to 2.218 to 6.001 nats on the first language, which is worse than never having learned it, with 96 percent of its first-language generations coming out in the second language.

One number governs how all of those are read. Repeating a training run with nothing changed moves about one row of 8 to 12, in either direction. That is a measured floor, established in a run where nothing was unlearned, and every strict per-row count above sits inside it. We publish the strict count anyway.

Bucket three: what the literature says about the alternatives

The cleanest controlled comparison toggles browsing on and off with the same models, prompts and grader. On biographies, which are long-tail entity facts, error rates fall by factors of 4.25 and 7.39 when browsing is enabled. On broad conceptual questions retrieval barely helps, and for one model it made things slightly worse. That is the most useful shape in this literature: retrieval's benefit is concentrated almost entirely on the long tail. There is also a theoretical floor. A 2025 analysis proves a base model's error rate on facts is lower-bounded by the share of those facts appearing exactly once in its training data.

Also true

Bucket three bounds the alternatives and says nothing about us: we have not run our system on any of those benchmarks. The phrase "hallucination rate" also means at least five different things across those sources, so no two of them appear in one sentence here. Abstention silently drives all of them, since a model halves its rate by refusing more, having learned nothing.

From a token ratio to a bill ratio

Here the honest number is much smaller than the exciting one, for an old reason. If a technique improves one part of a system, the total improvement is capped by how large that part was. Output tokens are 11.2 percent of an agentic coding bill and nothing on the input side touches them. Figure 6 draws the consequence.

HOW MUCH CHEAPER THE WHOLE BILL GETS, AS THE INPUT SIDE SHRINKS 10× 10×100×1000× factor by which the input side shrinks (log scale) ceiling 8.9×, set by the 11.2% output share our 52× → 7.8× 20× → 6.4× a hundred times on the whole bill would sit far above this chart, and above the ceiling
The bill cannot fall as fast as the prompt, because output tokens are untouched. The curve is one divided by ((0.292 plus 0.595) divided by R, plus 0.112), where R is the factor by which the input side shrinks. It flattens quickly: a tenfold input reduction gives 5.0 times on the bill, our measured 52 gives 7.8, and an infinite reduction gives 8.9. The dashed line is that ceiling. Log x, because the input ratio spans three decades.

What the data says

Work it through at a checkable pace. The bill splits 29.2 percent new input, 59.5 percent re-read, 11.2 percent output. Shrink all input by a factor R and the remaining bill is 0.887 over R, plus 0.112. At R of 10 that is 0.201, or 5.0 times cheaper. At our measured 52 it is 0.129, or 7.8 times. At infinity it is 0.112, or 8.9 times and no more. So our 52 already captures 7.75 of the 8.93 available, which is 87 percent of everything an input-side technique could ever win. The remaining headroom is small, and that is the useful thing to know.

The conservative reading is lower and belongs beside it. If learning displaces only the re-read portion and leaves genuinely new input alone, which is the honest assumption since new code and new user text are new, the bill goes to 0.292 plus 0.112, or 2.5 times cheaper. The defensible band on an agentic coding bill is therefore roughly 2.5 to 7.8 times, and where you land depends on how much of your prompt is material that could have been learned in advance.

So, a hundred times cheaper?

No, and the reason is arithmetic. To reach a hundred times on a whole bill the total has to fall to 1 percent of what it was. Output alone is 11.2 percent. So even with input driven to zero you would still need output tokens to fall by a further factor of 11.2. We have not measured output tokens at all, so we claim none of that. The honest bounded statement is this: on prompt tokens, for a question whose answer has been learned, 52 times, median, over 68 questions on three corpora. On a whole agentic bill, projected from a measured cost split, between 2.5 and 7.8 times, with a hard ceiling of 8.9 that no input-side technique can pass.

Why this is hard

The temptation is to quote the largest true number, and the largest here is 2,649 times, a real ratio for a real question against a real control. It is also the ratio for putting 22,000 words in a context window on every turn, which nobody sensible does, and it converts to under 8.9 times on a bill just like every other ratio above about a hundred. Quoting it would be true and would mislead.

Also true

This is a projection, not a measurement. It composes two measured quantities, a cost split from somebody else's agentic coding trace and a prompt ratio from our own single-document question answering, and the composition is untested. Whether a 52 times prompt ratio survives on a repository-scale coding workload is open, and the honest experiment is a large one.

What it costs to learn instead

A ratio per question hides a cost per document. Teaching is a training run, it takes minutes and it costs money, and it happens once. So the buyer's question is not the ratio, it is how many questions it takes to get the money back. Figure 7 answers it at published list prices.

QUESTIONS BEFORE TEACHING PAYS FOR ITSELF · LOG SCALE · LIST PRICES, UNCACHED 1101001000 1302 521 260 603 241 121 735 294 147 20 8 4 706-word docvs retrieval 706-word docvs whole document 22,000-word corpusvs retrieval 22,000-word corpusvs whole document tens of questions $2 / MTok input $5 / MTok $10 / MTok
Against a large context the payback is tens of questions. Against a good retrieval stack on a small document it is hundreds to low thousands. Break-even is the measured teach cost divided by the per-question saving in input tokens at each price. Log y, because the four cases span two decades. With cached input billed at a tenth of the list rate, every one of these multiplies by about ten.

What the data says

Measured teach costs, in our own compute: nine structured facts took 1,120.6 GPU-seconds for about 62 cents. One 706-word document took about 50 cents on its own and 1.50 to 2.00 dollars through the full demonstration path. Two 22,000-word corpora together took 6,071 GPU-seconds for about 4.60 dollars, so about 2.30 dollars each. The per-question saving is the token difference multiplied by the price: 576 tokens for the handbook against retrieval, 1,243 against the whole document, and 57,884 tokens for the 22,000-word corpus against the whole document. Divide and you get the figure.

What it means for building

Learning is worth it when a document will be asked about many times, or when it is large enough that shipping it is expensive. A reference manual a whole team queries daily clears break-even in a week. A document uploaded once, asked about twice and never seen again does not, and retrieval is the right tool for that. Any honest system does both, and chooses.

Also true

Those are our measured compute costs on current hardware and settings, not a customer price, and they will move. The break-even calculation counts input tokens only, for the same reason the ratio does.

The optimum, and where it actually lands

The best case is simple to state and we have measured it. If the material is learned, the prompt is the question: 11 to 37 tokens across every lane, flat in the size of what was learned. There is nothing beyond that, since you cannot ask a question in fewer tokens than the question.

Practice sits below that optimum for four reasons, all worth naming.

The long tail still wants retrieval. The controlled browsing comparison above puts retrieval's benefit almost entirely on rare entity facts. Material that is genuinely new to the world sits in that regime until somebody teaches it, and something has to fetch it in the meantime.

Learning is a decision taken before the question. Retrieval is one taken during it. A system that only learns cannot answer about a document uploaded a second ago, so the honest architecture keeps both routes.

The teach cost has to be amortised, which is the whole of the previous section.

Our serving path has an 8,192-token context window. A real constraint, and it cuts both ways: it is why the whole-document control was computed rather than served for the large corpora, and it means that for us the alternative to learning a 22,000-word corpus is not a long prompt but retrieval, or nothing.

One more reason not to assume a bigger window solves this: a 2025 study measured accuracy as input length grows, with retrieval held perfect and irrelevant tokens stripped out, and found drops of 13.9 to 85 percent across five models. Length itself costs accuracy even when retrieval is flawless. Cost moves the same way. One forecaster projects token prices falling about 95 percent by 2030 while inference cost per agentic workflow rises more than fivefold through 2028, because tokens per task grow faster than price per token falls.

What a token costs

Published list prices per million tokens, observed on 2026-08-25. Promotional rates change and at least one row below is explicitly temporary.

ProviderModelInputCached inputOutput
OpenAIgpt-5.6-sol$4.00$0.40$20.00
OpenAIgpt-5.6-terra$2.00$0.20$12.00
OpenAIgpt-5.6-luna$0.20$0.02$1.20
AnthropicClaude Fable 5$10.00$1.00$50.00
AnthropicClaude Opus 5$5.00$0.50$25.00
AnthropicClaude Sonnet 5$2.00$0.20$10.00
AnthropicClaude Haiku 4.5$1.00$0.10$5.00
GoogleGemini 3.7 / 3.6 Flash$0.75$0.075$3.75
GoogleGemini 3.1 Pro Preview$2.00$0.20$12.00
xAIgrok-4.6, 500k context$2.00$0.50$6.00
xAIgrok-4.3, 1M context$1.25$0.20$2.50
DeepSeekv4-pro, off-peak / peak$0.66 / $1.32$0.022 / $0.044$1.98 / $3.96

Three things move real cost more than the headline row. Anthropic's newer tokenizer produces about 30 percent more tokens for the same text, so a price cut across a model generation is not the whole story. OpenAI, Google and xAI roughly double the input rate above about 200,000 tokens of context, while Anthropic bills a 900,000-token request at the same rate as a 9,000-token one. And Google alone bills cache storage by the hour, from 50 cents to 4.50 dollars per million tokens per hour, on top of the cache-read rate.

What matters more than any row above is what output actually costs once you pay for the input that produced it. Under the measured agentic mix of 12.5 fresh plus 281.2 cached input tokens per output token, the effective cost per million output tokens is 456 dollars for Claude Fable 5 against its 50-dollar list price, 228 for Opus 5, 182 for gpt-5.6-sol, 172 for grok-4.6 and 93 for Gemini 3.1 Pro. The list price of output is not what output costs, and the difference is the re-reading.

The honest gaps, and the measurement that closes the biggest one

Four things here are weaker than they look, and one is fixable this week.

We did not count completion tokens. The 68-question distribution counts prompt tokens only, assuming completions are similar across arms. That is a reasonable prior and not a measurement, and the bill ratio depends on it, so it is the gap that matters most. Three outcomes are possible: the learned arm answers more tersely, because it is not paraphrasing a retrieved passage, and the bill ratio beats 7.8. Or the two match, and 7.8 stands. Or the retrieved text constrains the control into shorter answers, and the ratio is worse.

The measurement that closes it is small and specific. Serve all three arms live on the same checkpoint in one window, on the same 68 questions so every new number pairs to an existing one. Count prompt and completion tokens from the served usage fields rather than re-counting. Grade every arm with the same strict grader, which also turns control-arm correctness from one point into 68. Then report the paired total-token ratio, the subset where both arms are correct, a Wilcoxon signed-rank test on the paired difference and an exact McNemar test on the correctness disagreements. The bar is written down in advance: if the median paired ratio on the both-correct subset falls outside a factor of 1.5 of the prompt-only ratio, the arithmetic here is re-derived before publication rather than after.

The bill projection composes two measurements never made together: somebody else's agentic coding cost split and our own single-document ratio. Whether the ratio survives on a repository-scale workload is unmeasured, and the honest experiment is a large one.

Control-arm correctness has one live point. One question, both arms correct, 858 tokens against 18. The re-measurement above turns that into 68.

There is no energy number here, because there is no energy number to have. No laboratory publishes energy per token. Every official figure is per query, and they are not comparable to one another. The published spread from a simple text query to an agentic reasoning query is roughly a thousandfold, from about 0.05 to about 50 watt-hours on one workload-level estimate, which is both the most decision-relevant figure in that literature and the reason a token-to-joule conversion would be fiction. The International Energy Agency notes that even at a generous blended average of one watt-hour per query, ten billion queries a day comes to roughly 3.6 terawatt-hours a year, under one percent of what data centres already consume, so published per-prompt energy cannot explain the build-out.

Method, at the level we can state it

Confidentiality. The learning mechanism inside the product is confidential and is reserved under NDA pending patent prosecution. This piece describes the frozen base model, the count of trainable parameters, the training and serving protocol, and the measured behaviour. It never describes the mechanism. In every comparison table the method cell reads proprietary (details withheld). Every number traces to a data file named at the end.

The base model is Qwen3.6-27B in bf16, frozen. About 1.14 billion parameters are trainable. Teaching runs a short warm-up phase and then a single continuous optimizer over the material, with no task identifiers and no boundary signal of any kind. Serving uses an 8,192-token context window. All three prompt variants were counted with one counter identical across arms, chosen over any per-model tokenizer precisely so the arms stay comparable. The learned arm's prompt is the question with an empty context, exactly as the two live single-question measurements ran it. The retrieval control is BM25 over the top four chunks of the same source. The whole-document control is the question plus the entire source, computed rather than served.

How these answers were produced. Every answer quoted on this page is raw model output, shown exactly as it was served. There is no system prompt, no prompt engineering, no retry, and no reranking. Each question is asked once in a plain wording and graded by exact substring match, so an answer that is correct but phrased differently is scored as a miss. A harness that tunes the prompt, samples more than once, or grades more generously would score higher than what you see here. We publish the strict number.

Open questions

  1. Does a single-document prompt ratio survive on a repository-scale agentic workload? The projection in this piece assumes it does and does not test it.
  2. Do completion tokens differ across arms, and in which direction? Unmeasured. The design is above.
  3. What share of all AI revenue is re-read context? Not derivable from anything public. It needs an input-to-output split published by a provider and no provider publishes one.
  4. What is the energy ratio? Not derivable. No laboratory publishes energy per token.
  5. How does the ratio behave as one learner accumulates fifty documents rather than one? Every measurement here is one document or one pair of corpora.
  6. What is the retrieval control's accuracy across the whole distribution, rather than at one point?
  7. Does the argument survive a million-token context window? The published finding that accuracy falls with input length even under perfect retrieval suggests a larger window changes less than its size implies, but we have not tested it.
  8. Four of eight propositions in the teaching demonstration failed, all of them ties between a name and a specific organisation. The extraction fix shipped and extracted all four on a replay. Whether they then answer correctly at serving time, in several wordings, is the next measurement in that line.

Data and protocol

Everything below is published so the numbers can be checked, and the training documents are synthetic and invented precisely so that they can be.

WhatWhere
the 68 questions, three prompt variants each, with token counts and gradestoken-efficiency/distribution.json
the two earlier live single-question measurementstoken-efficiency/live-points/
our re-read share measurement over a 400-file repositoryreread-audit/
the source documents: a company handbook, an invented API reference, two invented languagescorpora/
the retrieval control, exactly as runcontrols/bm25/
measured teach costs in GPU-secondsspend/
every demonstration transcript, verbatim and untruncateddemos/

Citation

@misc{learnerlabs2026tokeneconomics, title = {Most of the bill is re-reading}, author = {Ganguli, Anurup}, year = {2026}, month = {August}, note = {Learner Labs, draft} }

Sources

Every industry figure above is numbered here with its source and the date it was observed. Figures we measured ourselves are in the data repository instead.

  1. Anthropic run-rate: CNBC, 2026-08-17; TechCrunch, 2026-08-17. Anthropic has published no audited annual revenue figure for any year.
  2. Microsoft revenue from OpenAI, 24.1 billion dollars inclusive of revenue-sharing: Microsoft FY2026 Form 10-K, SEC EDGAR, filed 2026-07-29.
  3. OpenAI run-rate and enterprise crossover: CNBC, 2026-08-14; Reuters via Investing.com, 2026-08-13.
  4. OpenAI audited 2025 revenue and operating loss: Where's Your Ed At, 2026-06-16, verified by the Financial Times.
  5. The two companies count revenue differently: Forbes, 2026-03-25.
  6. Valuations: Anthropic Series H; OpenAI financing announcement; CNBC on the xAI merger; Caixin on DeepSeek. Market capitalisations at close 2026-08-25: stockanalysis.com.
  7. Segment disclosure limits: Alphabet Q2 2026 8-K; Meta Q2 2026; SpaceX S-1; Alibaba 6-K.
  8. Inference share of compute: Gartner, 2026-08-10; Deloitte TMT 2026; Epoch AI; NVIDIA's 40 percent and the declined restatement, NVIDIA Q3 FY2026 call transcript.
  9. The agentic coding trace: Zhu, Jacob, Ma, Pan, Wang, Krishnamurthy and Kasikci, TraceLab: Characterizing Coding Agent Workloads for LLM Serving, arXiv:2606.30560, 2026-06-30.
  10. The porting case study: Bun, 2026-07-08. General traffic ratios: OpenRouter State of AI. Cached-prompt share of agentic burn: ppc.land on OpenRouter data, 2026-08-25. Caching case study: ProjectDiscovery, 2026-04-10.
  11. Prices observed 2026-08-25: OpenAI, Anthropic, Google, xAI, DeepSeek. Caching terms: Anthropic, OpenAI.
  12. Recall and hallucination: GPT-5 System Card; SimpleQA; AA-Omniscience; Why Language Models Hallucinate; Large Language Models Struggle to Learn Long-Tail Knowledge.
  13. Retrieval helps and leaves a residue: FreshLLMs; Vectara hallucination leaderboard, updated 2026-05-11; Stanford RegLab on legal research tools; ReDeEP.
  14. Long context is not free: Context Length Alone Hurts LLM Performance Despite Perfect Retrieval; Lost in the Middle; Context Rot.
  15. Cost per agentic workflow rising while token prices fall: Gartner, 2026-08-17.
  16. Energy: Google on inference energy and its technical paper; Microsoft Research in Joule; Epoch AI; IEA, Key Questions on Energy and AI, 2026-04-16.
  17. Token volumes: Google I/O keynote, 2026-05-19; Menlo Ventures on OpenRouter, 2026-05-26.

Figures marked as measured come from a trace or a benchmark. Figures marked as estimated come from an analyst or a third party. Figures marked as projected are forward-looking. Where a quantity does not exist publicly, this piece says so rather than approximating it: no provider publishes an input-to-output token split, no laboratory publishes energy per token, no published split of cloud capital expenditure into training and inference exists, and no current bottom-up study attributes enterprise value to AI at the product level.