Demonstrations · Continual Learning Bench
Twelve years of sales forecasts, one model learning as it goes
On the sales prediction task of Continual Learning Bench, our Learner 1.0 model forecasts furniture sales for twelve consecutive years and is trained after each year. In one complete run it scores 8.44 of 12 where the same model with learning off scores 5.76, and it is ahead in 10 of the 11 years after the first.
Model Learner 1.0
Benchmark Continual Learning Bench, sales prediction, the task's one fixed year order
Years 12, each a five-year forecast of 75 product, city and year sales figures
Runs one complete run (n = 1) and one stateless baseline
Learning after every year, one pass, one example per update, no replay
What the learner sees only the current year, exactly as the benchmark presents it
Teacher a teacher LLM writes the training examples after each year
This is one run. The database exploration and cohort studies results each use three runs. More runs of this task are in progress, and this page will be updated when they finish.
Learner 1.0 is the only system on this benchmark that learns strictly through its weights. Its context is reset before every year, so there is no in-context learning between years: anything it carries from one year to the next is in its weights. Every other system on the public leaderboard carries its experience forward in its context, a notepad, a memory store or a coding agent’s workspace.
1. The question and what was measured
The question: can our continual learning agent learn from its own experience by updating its weights, year after year? We use only weight-based learning and reset the context between years, so there is no in-context learning between years. If the agent improves over the baseline, it must be due to online parametric continual learning.
Each year the agent receives a data room of past furniture sales for a retailer that is expanding across cities, and it must forecast the next five years of sales for five products in three cities. It works in a shell, inspecting files and running analysis scripts, and then submits one structured forecast. Reward for a year is the task's composite score of that forecast against the actual sales, between 0 and 1. The years come in one fixed order, and every year starts with an empty workspace.
Reward is the sum over the twelve years (at most 12). Gain is reward minus the reward of the same model playing the same years with learning off, which the benchmark calls the stateless baseline. Gain is the benchmark's learning metric.
The learner totals 8.44 against 5.76 for the baseline. The first year is played before any training, so the two start from the same untrained model. Year 1 scores 0.469 for the learner and 0.539 for the baseline, a difference that comes from run-to-run variation alone.
2. Setup
Model. Learner 1.0, the same model as in the ten-skill experiment, with a fixed number of trainable parameters set before the first year. The mechanism is proprietary and is not described here.
What the learner sees. At each year the learner receives the benchmark's own prompt for that year and nothing else: no notes from earlier years and no retrieved examples. Its workspace is wiped before every year. Anything it carries from one year to the next is in its weights.
The loop. The learner plays a year. The benchmark scores it. A teacher LLM reads the learner's commands and their results for that year and writes lessons: what the year needed and which commands get there. The recipe turns each lesson into training examples built on states the learner actually reached, and the learner is trained on them before the next year. The teacher never forecasts and never acts in the benchmark.
Training. One pass, batch size one, no replay of earlier examples, no optimizer resets between years. Across the run the learner trained on 73 examples and 32,729 supervised tokens.
Runs. The task has one fixed year order, so runs differ only by seed. The stateless baseline plays every year on its own, with nothing carried over.
3. Every step of every year
The figure draws every step of every year, in order, for the baseline and for the learner. Each row is a year and each cell is one step, colored by what the step does.
Over the twelve years the baseline takes 169 steps and the learner 177. The learner's better scores do not come from doing more or less work. They come from what the work produces: in year 9, for example, the learner reaches a score of 0.94 in 9 steps, where the baseline reaches 0.37 in 13.
4. What it learned
After each year the teacher writes a short rationale for every lesson, and each training example carries one. A few of them, verbatim, with the first command of the example:
| trained after year | lesson (verbatim) | first line of the target command |
|---|---|---|
| 2 | catalog keys known from prior experience -> look up target IDs by the real name key directly | python3 -c " |
| 3 | Pull target-product history from each store file using that store's own id/year/quantity column names; filter on catalog IDs, not names. | python3 << 'EOF' |
| 9 | 2035: only Chicago changed -> inspect only its header | head -20 data/sales_chicago.csv |
| 11 | stdlib-shadowing import error -> remove shadowing file, then do the needed inspection | rm /app/inspect.py && head -20 data/sales_new_york.csv |
The three store files of this task name the same columns differently, and the catalog lists products by name with a numeric ID. Early lessons teach the learner to read each store's own column names and to look products up by ID. Later lessons teach it to read only the store file that changed, and to recover from errors it caused itself, such as a script that shadowed a standard library module.
5. Learner 1.0 learns efficiently online from very few examples
The forecast is never a training target. Every training example is a step of the learner's own attempt at a year: the prompt, the commands it ran so far and their real results, with the next command as the target. None of the 73 examples has the forecast itself as its target: every target is a command. Training happens only after a year has been played, and each year is played once.
6. Results
| lane | years | reward (of 12) | mean reward per year | gain over baseline | mean gain, years 2–12 | years ahead (2–12) |
|---|---|---|---|---|---|---|
| Learner 1.0 | 12 | 8.44 | 0.703 | +2.68 | +0.250 | 10 of 11 |
| Stateless baseline | 12 | 5.76 | 0.480 | – | – | – |
The gain grows over the run: +0.150 per year on average over years 2–6 and +0.334 over years 7–12.
| year | years forecast | Learner 1.0 | stateless baseline | gain |
|---|---|---|---|---|
| 1 | 2027–2031 | 0.469 | 0.539 | -0.070 |
| 2 | 2028–2032 | 0.683 | 0.573 | +0.111 |
| 3 | 2029–2033 | 0.663 | 0.514 | +0.149 |
| 4 | 2030–2034 | 0.579 | 0.600 | -0.020 |
| 5 | 2031–2035 | 0.704 | 0.487 | +0.217 |
| 6 | 2032–2036 | 0.742 | 0.448 | +0.295 |
| 7 | 2033–2037 | 0.768 | 0.454 | +0.314 |
| 8 | 2034–2038 | 0.894 | 0.442 | +0.453 |
| 9 | 2035–2039 | 0.941 | 0.367 | +0.574 |
| 10 | 2036–2040 | 0.659 | 0.482 | +0.176 |
| 11 | 2037–2041 | 0.801 | 0.485 | +0.316 |
| 12 | 2038–2042 | 0.537 | 0.367 | +0.169 |
7. Against the published leaderboard
Learner 1.0 is the only system on this leaderboard that learns strictly through its weights. Its context is reset before every year, so no in-context learning is allowed between years. Every other system carries its experience forward in its context, a notepad, a memory system or a coding agent's workspace.
The benchmark publishes results for twelve systems on this task. It ranks them by normalized reward, the total reward measured against GPT-5.4's stateless baseline, and by normalized gain, each system's gain over its own stateless baseline divided by its headroom up to the reference score.
| system | runs | normalized reward | reward rank | normalized gain | gain rank |
|---|---|---|---|---|---|
| ICL Notepad · Claude Sonnet 4.6 | 5 | 0.699 | 1 | 0.750 | 1 |
| ICL · Claude Sonnet 4.6 | 5 | 0.690 | 2 | 0.709 | 2 |
| Claude Code · Sonnet 4.6 | 5 | 0.621 | 3 | 0.651 | 3 |
| ICL · GPT-5.4 | 5 | 0.567 | 4 | 0.567 | 5 |
| ICL · Claude Opus 4.7 | 5 | 0.538 | 5 | 0.611 | 4 |
| ICL Notepad · GPT-5.4 | 5 | 0.451 | 6 | 0.532 | 7 |
| ICL · Gemini 3 Flash | 5 | 0.450 | 7 | 0.477 | 8 |
| Learner 1.0 (ours) | 1 | 0.444 | 8 | 0.430 | 10 |
| Codex · GPT-5.4 | 5 | 0.425 | 9 | 0.461 | 9 |
| ICL Notepad · Gemini 3.1 Pro Preview | 5 | 0.392 | 10 | 0.560 | 6 |
| Mem0 · GPT-5.4 | 5 | 0.346 | 11 | 0.423 | 11 |
| ICL · Gemini 3.1 Pro Preview | 5 | 0.136 | 12 | 0.316 | 12 |
| ACE · GPT-5.4 | 5 | 0.081 | 13 | 0.134 | 13 |
On normalized reward we rank 8 of 13, and on normalized gain 10 of 13, from one run against five for each published system.
8. The recipe
After each year. The teacher receives the learner's commands and results for the year just played, the benchmark's released feedback for it, the learner's trajectories on earlier years as evidence, and the lessons it wrote before. It returns lessons in a fixed format: what the year needed, the states that lead to a forecast, and the exact next command at each state. It never writes the forecast.
From lessons to examples. The recipe builds each example from a state the learner actually reached: the conversation up to that point, with the recorded command results, and the next command as the target. It keeps only examples grounded in recorded results and adds a few fixed examples about reading the data room. Every example is trained once.
Acting. Thinking when stuck, used on the database exploration task, was not used here.
9. Interpretation, and how to check it
Training after each year changed how the learner works through this task's data, and its forecasts scored higher than the same model's with learning off in 10 of the 11 years after the first.
To check: every trajectory of both lanes, the per-year scores and every training example are in the replication repository, and the figures regenerate from them.