Learner LabsLearner 1.0Foundation models that continually learn

89% of the bill goes to input tokens. What if the model could remember?

What happens when foundation models can continually learn?

Disclaimer. This is back-of-the-envelope math with reasonable assumptions. The assumptions are stated so they can be pressure-tested.

LET'S DEFINE THE VARIABLES FIRST

Cost of input tokens (to customer)  = I  ($/M tokens)
Cost of output tokens (to customer) = O  ($/M tokens)
General assumption: output tokens cost 5× input tokens, i.e. O = 5·I.1
  1. Total AI-lab revenue (what customers are paying them) = $ paid for input tokens + output tokens.
  2. Input tokens = cached + new, where cached bills at 10% of the input price.2
  3. Output tokens always bill at 100% of O (so no discount).
  4. From publicly available data: a trace of real agentic coding published in June 2026: TraceLab, 4,265 sessions from 43 developers, 357,161 model steps and 432,510 tool calls, worth about $40,400 of API-equivalent spend:
    Total input tokens = 54.9 B Cached input tokens = 52.56 B Genuinely new input tokens = 54.9 − 52.56 = 2.34 B Total output tokens = 186.9 M
  5. So money spent by you, the customer:
    = (52.56 × I × 10%) + (2.34 × I) + (0.1869 × O) = (52.56 × I × 10%) + (2.34 × I) + (0.1869 × 5 × I) = (5.256 + 2.34)·I {input} + 0.935·I {output} = 7.596·I {input} + 0.935·I {output} As a share of the total (7.596 + 0.935 = 8.531): 89% of the money is input tokens · 11% is output tokens
  6. For every $100 of modeled spend in this trace, about $89 is input-token charges, including $61.60 for cached input.
EVERY $100 OF MODELED SPEND IN THIS TRACE, SPLIT BY WHAT IT BUYS · TRACE PRICED AT O = 5·I, CACHE = 10% $61.60Cached input tokens61.6%$27.40genuinely new input27.4%$11.00 · output11.0%
Most of the money buys tokens the model has already read. The split follows directly from the trace: cached re-reads bill at a tenth of the input price and still come to 61.6% of the total, because there are 22× more of them than genuinely new tokens. Output, the only part that is actually new work, is 11%.

So what happens if a continually learning foundation model needs to learn a dataset only once, and then never consumes context again for that data, because it answers from weights instead?

What we measured: the prompt for a question whose answer has been learned

Learner 1.0 adds about 1.14 billion trainable parameters, and then trains only those parameters on whatever a user teaches it. Once the material is learned, the prompt for a question about it is the question itself. The figure below is 68 previously graded questions from three demonstrations, each asked three ways.

PROMPT TOKENS, CONTROL ÷ LEARNED · SAME QUESTION · LOG SCALE 10×30×100×300×1000×3000×28.1× live47.7× livehandbook, 706 words24 questions38.589.8reference page, 2.4 kB12 questions25.529.1two corpora, 22,000 words each32 questions59.22,133 against retrieval (BM25, top-4 chunks) against the whole document in context
The bigger the material, the bigger the saving. The learned prompt stays small no matter how much was learned. Bars are minimum to maximum across the questions in each lane, dots are the lane median, and the x-axis is logarithmic because the quantity is a ratio. The two dashed lines are earlier single-question measurements taken live, at 28.1× and 47.7×, both of which fall inside the retrieval band. The retrieval control is a BM25 pipeline taking the top four chunks of the same source. The whole-document control was computed from the corpora rather than served, because a 22,000-word corpus does not fit our 8,192-token serving window.

Across all 68 questions the retrieval control pays a median of 51.9× more prompt tokens, with an interquartile range of 32.4 to 61.7 and a distribution-free interval on the median of 42.0 to 56.0. Per lane the medians are 38.5 for the handbook, 25.5 for the reference page and 59.2 for the invented-language corpora. The whole-document control runs from a median of 29.1 on the smallest source to 2,132.9 on the largest, and we report it per lane rather than pooled, because the pooled interval runs from 89.8 to 1,987.2 and a range that wide is not a summary of anything.

14
median prompt tokens after learning a 706-word handbook
26
median prompt tokens after learning 22,000 words
590
median prompt tokens for the retrieval control on the handbook
57,910
median prompt tokens to put the 22,000-word corpus in context

Those four numbers are the whole of the ratio. The learned prompt is 11 to 37 tokens across every lane and does not grow: 14 after a 706-word handbook, 26 after 44,000 words of invented language. Every context-based route pays for the material on every question. Learning pays for it once.


Conclusions

This is tested on a handful of experiments, so it does not claim broad generality, but the physics and the equations should hold.

Continual learning targets the repeated-context portion of the input bill.

One-time training cost, anticipated for large corpora and not optimized: $6 per million tokens. So for a large enterprise with 5 billion tokens of content, the cost is roughly $30,000, one time.

The accuracy of our Learner 1.0 foundation model can be read in depth in the demonstrations. The numbers today are likely near the floor and not the ceiling: room for optimization in the data preprocessing and other aspects of the training pipeline is completely untouched.

A few more facts

The weights are owned by you, the customer, and are deleted when you delete the checkpoint. No backup is maintained at this time. Every time you start a session, your unique weights are loaded onto a GPU and served to you. This currently adds latency to the start of the experience, due to data movement, and will be optimized in the near future. Once you are idle for 15 minutes, your GPU is automatically released after your weights are transferred to persistent memory. This is done to prevent ongoing billing of expensive GPU time.

Today, most of what the industry bills you for is re-reading what you already told it. A model that learns your material once answers from the weights, and that line of the bill goes away. That is what we are building.


1. Published price sheets run about 5–6×, e.g. GPT-5.6 at $4 in / $20 out, Claude Sonnet 5 at $2 in / $10 out per million tokens (prices observed August 2026).

2. Cached-input rates at OpenAI and Anthropic are about 10% of the input price.

3. The trace: Zhu, Jacob, Ma, Pan, Wang, Krishnamurthy and Kasikci, TraceLab: Characterizing Coding Agent Workloads for LLM Serving, arXiv:2606.30560.

4. The 68 questions, the token counts, the grades, the source documents and the retrieval control are published in the data repository.

89% of the bill goes to input tokens · Learner Labs