Learner LabsLearner 1.0Foundation models that continually learn

Demonstrations · Continual Learning Bench

Twelve years of sales forecasts, one model learning as it goes

On the sales prediction task of Continual Learning Bench, our Learner 1.0 model forecasts furniture sales for twelve consecutive years and is trained after each year. In one complete run it scores 8.44 of 12 where the same model with learning off scores 5.76, and it is ahead in 10 of the 11 years after the first.

Model Learner 1.0

Benchmark Continual Learning Bench, sales prediction, the task's one fixed year order

Years 12, each a five-year forecast of 75 product, city and year sales figures

Runs one complete run (n = 1) and one stateless baseline

Learning after every year, one pass, one example per update, no replay

What the learner sees only the current year, exactly as the benchmark presents it

Teacher a teacher LLM writes the training examples after each year

This is one run. The database exploration and cohort studies results each use three runs. More runs of this task are in progress, and this page will be updated when they finish.

Learner 1.0 is the only system on this benchmark that learns strictly through its weights. Its context is reset before every year, so there is no in-context learning between years: anything it carries from one year to the next is in its weights. Every other system on the public leaderboard carries its experience forward in its context, a notepad, a memory store or a coding agent’s workspace.

1. The question and what was measured

The question: can our continual learning agent learn from its own experience by updating its weights, year after year? We use only weight-based learning and reset the context between years, so there is no in-context learning between years. If the agent improves over the baseline, it must be due to online parametric continual learning.

Each year the agent receives a data room of past furniture sales for a retailer that is expanding across cities, and it must forecast the next five years of sales for five products in three cities. It works in a shell, inspecting files and running analysis scripts, and then submits one structured forecast. Reward for a year is the task's composite score of that forecast against the actual sales, between 0 and 1. The years come in one fixed order, and every year starts with an empty workspace.

Reward is the sum over the twelve years (at most 12). Gain is reward minus the reward of the same model playing the same years with learning off, which the benchmark calls the stateless baseline. Gain is the benchmark's learning metric.

Reward rises as the learner trains after each year measured Running average of reward per year, 12 years in order reward per year = the task's 0–1 composite score of a five-year forecast. Mean over the years played so far 0.3 0.4 0.5 0.6 0.7 0.8 0.9 1 1 2 3 4 5 6 7 8 9 10 11 12 year (position in the run) running average reward Learner 1.0 0.703 learning off 0.480 One complete run of the learner (n = 1) and the same model with learning off and the same decoding, each year played on its own. Year 1 is played before any training. Source: repo clbench/sales_prediction (outcomes.jsonl).
The learner ends with a running average of 0.703 per year against 0.480 for the same model with learning off. Higher is better: 1 is a perfect five-year forecast.

The learner totals 8.44 against 5.76 for the baseline. The first year is played before any training, so the two start from the same untrained model. Year 1 scores 0.469 for the learner and 0.539 for the baseline, a difference that comes from run-to-run variation alone.

Gain over the learning-off baseline measured Running average gain per year gain per year = reward of the learner − reward of the learning-off model in the same year -0.1 0 0.1 0.2 0.3 0.4 0.5 0.6 1 2 3 4 5 6 7 8 9 10 11 12 year (position in the run) running average gain Learner 1.0 +0.224 Line: running average gain. Dots: the gain in each single year. Gain is the CL-Bench learning metric: each year is paired with the same year played by the same model with learning off. Source: repo clbench/sales_prediction (outcomes.jsonl).
The total gain over the twelve years is +2.68. The learner is ahead of the learning-off model in 10 of the 11 years after the first.

2. Setup

Model. Learner 1.0, the same model as in the ten-skill experiment, with a fixed number of trainable parameters set before the first year. The mechanism is proprietary and is not described here.

What the learner sees. At each year the learner receives the benchmark's own prompt for that year and nothing else: no notes from earlier years and no retrieved examples. Its workspace is wiped before every year. Anything it carries from one year to the next is in its weights.

The loop. The learner plays a year. The benchmark scores it. A teacher LLM reads the learner's commands and their results for that year and writes lessons: what the year needed and which commands get there. The recipe turns each lesson into training examples built on states the learner actually reached, and the learner is trained on them before the next year. The teacher never forecasts and never acts in the benchmark.

Training. One pass, batch size one, no replay of earlier examples, no optimizer resets between years. Across the run the learner trained on 73 examples and 32,729 supervised tokens.

Runs. The task has one fixed year order, so runs differ only by seed. The stateless baseline plays every year on its own, with nothing carried over.

3. Every step of every year

The figure draws every step of every year, in order, for the baseline and for the learner. Each row is a year and each cell is one step, colored by what the step does.

Every step of every year: learning off and learning on measured Learning off year 1 0.54 year 2 0.57 year 3 0.51 year 4 0.60 year 5 0.49 year 6 0.45 year 7 0.45 year 8 0.44 year 9 0.37 year 10 0.48 year 11 0.49 year 12 0.37 reward Learner 1.0 year 1 0.47 year 2 0.68 year 3 0.66 year 4 0.58 year 5 0.70 year 6 0.74 year 7 0.77 year 8 0.89 year 9 0.94 year 10 0.66 year 11 0.80 year 12 0.54 reward list the data room read a file or header write and run an analysis script command returned an error submit the forecast One cell per step, in order, left to right. A step is one shell command, or the final structured forecast. Steps are classed from the command text only. Source: repo clbench/sales_prediction (trajectories.jsonl).
Over the twelve years the learning-off model takes 169 steps and the learner 177. The learner does not score higher by working faster: it reaches better forecasts in about the same number of steps.

Over the twelve years the baseline takes 169 steps and the learner 177. The learner's better scores do not come from doing more or less work. They come from what the work produces: in year 9, for example, the learner reaches a score of 0.94 in 9 steps, where the baseline reaches 0.37 in 13.

4. What it learned

After each year the teacher writes a short rationale for every lesson, and each training example carries one. A few of them, verbatim, with the first command of the example:

trained after yearlesson (verbatim)first line of the target command
2catalog keys known from prior experience -> look up target IDs by the real name key directlypython3 -c "
3Pull target-product history from each store file using that store's own id/year/quantity column names; filter on catalog IDs, not names.python3 << 'EOF'
92035: only Chicago changed -> inspect only its headerhead -20 data/sales_chicago.csv
11stdlib-shadowing import error -> remove shadowing file, then do the needed inspectionrm /app/inspect.py && head -20 data/sales_new_york.csv

The three store files of this task name the same columns differently, and the catalog lists products by name with a numeric ID. Early lessons teach the learner to read each store's own column names and to look products up by ID. Later lessons teach it to read only the store file that changed, and to recover from errors it caused itself, such as a script that shadowed a standard library module.

5. Learner 1.0 learns efficiently online from very few examples

How little it trains on measured Supervised tokens trained after each year one teach after each of years 1–11. One pass, one example per update, no replay 0 2,000 4,000 6,000 8,000 1 2 3 4 5 6 7 8 9 10 11 after year supervised tokens 7 8 7 11 10 8 1 5 10 6 Bar labels: number of training examples. The lessons written after years 7 and 8 were not confirmed in time and were trained after year 9 with that year's lessons. Source: repo clbench/sales_prediction (training_examples.jsonl).
Over the whole run the learner trains on 73 examples, 32,729 supervised tokens, about 7 examples after each year. Each example is trained once.

The forecast is never a training target. Every training example is a step of the learner's own attempt at a year: the prompt, the commands it ran so far and their real results, with the next command as the target. None of the 73 examples has the forecast itself as its target: every target is a command. Training happens only after a year has been played, and each year is played once.

6. Results

laneyearsreward (of 12)mean reward per yeargain over baselinemean gain, years 2–12years ahead (2–12)
Learner 1.0128.440.703+2.68+0.25010 of 11
Stateless baseline125.760.480–––

The gain grows over the run: +0.150 per year on average over years 2–6 and +0.334 over years 7–12.

yearyears forecastLearner 1.0stateless baselinegain
12027–20310.4690.539-0.070
22028–20320.6830.573+0.111
32029–20330.6630.514+0.149
42030–20340.5790.600-0.020
52031–20350.7040.487+0.217
62032–20360.7420.448+0.295
72033–20370.7680.454+0.314
82034–20380.8940.442+0.453
92035–20390.9410.367+0.574
102036–20400.6590.482+0.176
112037–20410.8010.485+0.316
122038–20420.5370.367+0.169

7. Against the published leaderboard

Learner 1.0 is the only system on this leaderboard that learns strictly through its weights. Its context is reset before every year, so no in-context learning is allowed between years. Every other system carries its experience forward in its context, a notepad, a memory system or a coding agent's workspace.

The benchmark publishes results for twelve systems on this task. It ranks them by normalized reward, the total reward measured against GPT-5.4's stateless baseline, and by normalized gain, each system's gain over its own stateless baseline divided by its headroom up to the reference score.

Sales prediction: our run against the published leaderboard measured Normalized reward 0 0.2 0.4 0.6 0.8 ICL Notepad · Claude Sonnet 4.6 0.699 ICL · Claude Sonnet 4.6 0.690 Claude Code · Sonnet 4.6 0.621 ICL · GPT-5.4 0.567 ICL · Claude Opus 4.7 0.538 ICL Notepad · GPT-5.4 0.451 ICL · Gemini 3 Flash 0.450 Learner 1.0 (ours, 1 run) 0.444 Codex · GPT-5.4 0.425 ICL Notepad · Gemini 3.1 Pro Preview 0.392 Mem0 · GPT-5.4 0.346 ICL · Gemini 3.1 Pro Preview 0.136 ACE · GPT-5.4 0.081 Normalized gain 0 0.2 0.4 0.6 0.8 ICL Notepad · Claude Sonnet 4.6 0.750 ICL · Claude Sonnet 4.6 0.709 Claude Code · Sonnet 4.6 0.651 ICL · Claude Opus 4.7 0.611 ICL · GPT-5.4 0.567 ICL Notepad · Gemini 3.1 Pro Preview 0.560 ICL Notepad · GPT-5.4 0.532 ICL · Gemini 3 Flash 0.477 Codex · GPT-5.4 0.461 Learner 1.0 (ours, 1 run) 0.430 Mem0 · GPT-5.4 0.423 ICL · Gemini 3.1 Pro Preview 0.316 ACE · GPT-5.4 0.134 Published systems: CL-Bench leaderboard data of 2026-07-18, sales_prediction task, 5 runs each. Ours: 1 complete run and one learning-off baseline with the same decoding. Normalized as the leaderboard does: reward against GPT-5.4's learning-off baseline, gain against each system's own learning-off baseline, each divided by the headroom up to the reference score. Sources: data/leaderboard_data_2026-07-18.json, repo clbench/sales_prediction.
Our run has a normalized reward of 0.444 (8 of 13) and a normalized gain of 0.430 (10 of 13).
systemrunsnormalized rewardreward ranknormalized gaingain rank
ICL Notepad · Claude Sonnet 4.650.69910.7501
ICL · Claude Sonnet 4.650.69020.7092
Claude Code · Sonnet 4.650.62130.6513
ICL · GPT-5.450.56740.5675
ICL · Claude Opus 4.750.53850.6114
ICL Notepad · GPT-5.450.45160.5327
ICL · Gemini 3 Flash50.45070.4778
Learner 1.0 (ours)10.44480.43010
Codex · GPT-5.450.42590.4619
ICL Notepad · Gemini 3.1 Pro Preview50.392100.5606
Mem0 · GPT-5.450.346110.42311
ICL · Gemini 3.1 Pro Preview50.136120.31612
ACE · GPT-5.450.081130.13413

On normalized reward we rank 8 of 13, and on normalized gain 10 of 13, from one run against five for each published system.

8. The recipe

After each year. The teacher receives the learner's commands and results for the year just played, the benchmark's released feedback for it, the learner's trajectories on earlier years as evidence, and the lessons it wrote before. It returns lessons in a fixed format: what the year needed, the states that lead to a forecast, and the exact next command at each state. It never writes the forecast.

From lessons to examples. The recipe builds each example from a state the learner actually reached: the conversation up to that point, with the recorded command results, and the next command as the target. It keeps only examples grounded in recorded results and adds a few fixed examples about reading the data room. Every example is trained once.

Acting. Thinking when stuck, used on the database exploration task, was not used here.

9. Interpretation, and how to check it

Training after each year changed how the learner works through this task's data, and its forecasts scored higher than the same model's with learning off in 10 of the 11 years after the first.

To check: every trajectory of both lanes, the per-year scores and every training example are in the replication repository, and the figures regenerate from them.