Demonstrations · Continual Learning Bench
Twenty clinical studies, one model learning between them
On the cohort studies task of Continual Learning Bench, our Learner 1.0 model estimates patient survival from twenty clinical study databases in sequence and is trained after each study on what it did there. Over three runs it scores -0.020, -0.149 and -0.099 (information gain in bits per cohort, summed over the twenty studies) where the same model with learning off scores -0.241. Averaged over the three runs that places it 2nd of 13 on reward and 5th of 13 on gain among the systems on the published leaderboard. The gain is small: the runs beat the baseline on 12 of 20 studies on average.
Model Learner 1.0
Benchmark Continual Learning Bench, cohort studies, default schedule
Studies 20 per run, in 5 study families of 4
Runs 3 study orders (0 canonical, 1 and 2 permuted inside each family) and one stateless baseline
Learning after every finished study. One pass, one example per update, no replay
What the learner sees only the current study, exactly as the benchmark presents it
Teacher a teacher LLM reads only the study just finished and writes the training examples
Compute one H200 per run
Learner 1.0 is the only system on this benchmark that learns strictly through its weights. Its context is reset before every study, so there is no in-context learning between studies: anything it carries from one study to the next is in its weights. Every other system on the public leaderboard carries its experience forward in its context, a notepad, a memory store or a coding agent’s workspace.
1. The question and what was measured
The question: if a model is trained on its own experience after every study it analyses, does it analyse later studies better than the same model that never learns? Each study is a database of patients from one clinical study. The agent has twenty tool actions: read the study's metadata and a data summary, run SQL, estimate survival for any grouping of patients, and fit a cohort predictor. Then it submits survival estimates at 12, 24 and 36 months for 36 fixed cohorts, 108 numbers. Each study records different variables, so only 3 to 10 of the 36 cohorts can be measured in any one study. The rest have to be estimated from what similar studies showed. The twenty studies come in five families of four (HERALD, MERIDIAN, MOSAIC, FORGE, CADENCE) that share a design.
Reward for a study is the information gain of the submitted estimates over the benchmark's reference predictor, in bits per cohort: 0 means as good as the reference, and the best possible is 0.162. Reward is the sum over the twenty studies. Gain is reward minus the reward of the same model analysing the same studies with learning off, the benchmark's stateless baseline. It is the benchmark's learning metric.
The three runs end at rewards -0.0200, -0.1493 and -0.0988 against -0.2408 for the baseline. The gain over the baseline is +0.2208, +0.0915 and +0.1420. Normalized as the leaderboard does, +6.34%, +2.63% and +4.07%, mean +4.35% (standard error 1.08%).
2. Setup
Model. Learner 1.0, the same model as in the ten-skill experiment, with a fixed number of trainable parameters set before the first study. The mechanism is proprietary and is not described here.
What the learner sees. At each study the learner receives the benchmark's own task brief and the current study's tools, and its own steps in that study, nothing else: no earlier studies, no notes, no scores, no retrieved examples. Its prompt is the same as the baseline's. Anything it carries from one study to the next is in its weights.
The loop. The learner analyses a study and submits. A teacher LLM reads the learner's steps in that study and the tool results, and writes procedures: which next action is supported at which state. The teacher sees only the study just finished: no earlier studies, none of its own earlier procedures, and no scores. The recipe turns each procedure into training examples built on states the learner actually reached, the learner writes the reasoning for each example's action in its own words, and it is trained on them before the next study. The teacher never analyses a study and never acts in the benchmark. Section 9 gives the full recipe.
Training. One pass, batch size one, no replay of earlier examples, no optimizer resets between studies. Across a run the learner trained on 80 to 111 examples, 12,901 to 18,835 supervised tokens.
Runs. The benchmark's default schedule plays the twenty studies family by family, in a fixed canonical order (run 0) and in orders that permute the four studies inside each family (runs 1 and 2). The stateless baseline is run once. Because nothing carries over between its studies, its order does not matter. Each run is compared with the baseline study by study. All four lanes used the same model and the same decoding.
3. One study, before and after learning
Here is the same study analysed by the baseline and by run 0, which reached it ninth, after training on eight earlier studies. MOSAIC is a health survey that records a PRB1 category and does not record the comorbidity count the earlier families had.
Stateless baseline
- metadata
get_database_metadata()=== Study Info === study_name: MOSAIC study_label: MOSAIC Survey description: Government health survey administered in Urban and Suburban centres. Uses convenience sampling … … 17 more lines
- summary
get_data_summary()Patient count: 129 patient_id (numeric, n=129): mean=12510.155, std=6883.348, min=375, p25=6262, median=13487, p75=18768, max=232… region (categorical, 1 levels): Urban=129 (100.0%) time_to_death_months (numeric, n=129): mean=24.149, std=14.428, min=1.25, p25=12.92, median=21.91, p75=33.46, … … 11 more lines
- SQLI now have a good overview of the data. Key observations: 1. 129 patients, all from Urban region 2. High mortality rate: 95.3% died (mean of 'died' column) 3. Median time to death is ~22 months, with max 72 months 4. Variables available for cohort definition: age, sex, prb1_category, genotype_group, ambient_voc_ppb, years_of_education, ad8_score, hhbp_serum_level I need to estimate survival at 12, 24, and 36 months for 36 cohorts. The cohorts will likely be defined by combinations of key variables. Let me first understand the survival patterns by looking at some key variables. I should start by examining survival by age groups, since age is typically a strong predictor of survival. Let me query survival rates by age groups to understand the baseline patterns.
query_sql(SELECT CASE WHEN age < 40 THEN 'young' WHEN age < 60 THEN 'middle' ELSE 'old' END as age_group, COUNT(*) as n, SUM(died) as deaths, AVG(time_to_death_months) as mean_time FROM patients GROUP BY age_group ORDER BY age_group)age_group | n | deaths | mean_time middle | 61 | 57 | 20.98 old | 57 | 55 | 25.97 young | 11 | 11 | 32.26 … 1 more lines
- estimate
estimate_survival_by_group(CASE WHEN age < 40 THEN 'young' WHEN age < 60 THEN 'middle' ELSE 'old' END)=== Per-Group Survival === middle: n=61 (47.3%) S(12m)=0.7163 S(24m)=0.3533 S(36m)=0.1413 old: n=57 (44.2%) S(12m)=0.8404 S(24m)=0.5722 S(36m)=0.2030 young: n=11 (8.5%) S(12m)=0.8182 S(24m)=0.6364 S(36m)=0.4545 … 35 more lines
- estimate
estimate_survival_by_group(CASE WHEN sex = 'M' THEN 'male' ELSE 'female' END)=== Per-Group Survival === female: n=64 (49.6%) S(12m)=0.7462 S(24m)=0.4281 S(36m)=0.1647 male: n=65 (50.4%) S(12m)=0.8133 S(24m)=0.5234 S(36m)=0.2290 === Cohort Group Decomposition (10 observable, 26 unobservable) === … 34 more lines
- steps 6–21 not shown
Learner, run 0
- metadata
get_database_metadata()=== Study Info === study_name: MOSAIC study_label: MOSAIC Survey description: Government health survey administered in Urban and Suburban centres. Uses convenience sampling … … 17 more lines
- summary
get_data_summary()Patient count: 129 patient_id (numeric, n=129): mean=12510.155, std=6883.348, min=375, p25=6262, median=13487, p75=18768, max=232… region (categorical, 1 levels): Urban=129 (100.0%) time_to_death_months (numeric, n=129): mean=24.149, std=14.428, min=1.25, p25=12.92, median=21.91, p75=33.46, … … 11 more lines
- estimateI have reviewed the database metadata and univariate statistics. The study has 129 patients, all from the Urban region, with a high mortality rate (95.3%). Key variables include age, sex, prb1_category, genotype_group, ambient_voc_ppb, and ad8_score. To estimate survival for the 36 cohorts, I need to understand how these variables interact with survival outcomes. I will start by estimating survival for a broad risk group based on age and PRB1 ratio, as these are likely significant predictors. This will help me gauge the baseline survival patterns before refining the cohort definitions.
estimate_survival_by_group(CASE WHEN age >= 65 AND prb1_category IN ('high', 'elevated') THEN 'high_risk' WHEN age < 50 THEN 'low_risk' ELSE 'medium_risk' END)=== Per-Group Survival === high_risk: n=29 (22.5%) S(12m)=0.9298 S(24m)=0.7152 S(36m)=0.3011 low_risk: n=34 (26.4%) S(12m)=0.7647 S(24m)=0.5294 S(36m)=0.3529 medium_risk: n=66 (51.2%) S(12m)=0.7221 S(24m)=0.3414 S(36m)=0.0650 … 35 more lines
- estimate
estimate_survival_by_group(CASE WHEN genotype_group = 'GG-1' THEN 'GG-1' WHEN genotype_group = 'GG-2' THEN 'GG-2' ELSE 'GG-3' END)=== Per-Group Survival === GG-1: n=53 (41.1%) S(12m)=0.7690 S(24m)=0.4141 S(36m)=0.2169 GG-2: n=41 (31.8%) S(12m)=0.8049 S(24m)=0.5122 S(36m)=0.1707 GG-3: n=35 (27.1%) S(12m)=0.7694 S(24m)=0.5295 S(36m)=0.1986 … 35 more lines
- estimate
estimate_survival_by_group(CASE WHEN age >= 60 AND ad8_score >= 3 THEN 'elderly_cognitive_impairment' WHEN age < 60 AND ad8_score < 3 THEN 'younger_cognitive_intact' ELSE 'mixed' END)=== Per-Group Survival === elderly_cognitive_impairment: n=11 (8.5%) S(12m)=1.0000 S(24m)=0.6364 S(36m)=0.0909 mixed: n=87 (67.4%) S(12m)=0.7802 S(24m)=0.5005 S(36m)=0.2622 younger_cognitive_intact: n=31 (24.0%) S(12m)=0.6984 S(24m)=0.3385 S(36m)=0.0376 … 35 more lines
- steps 6–21 not shown
The baseline writes a summary of the data at step 3 and starts from age bands. Run 0 splits the patients into high, medium and low risk tiers at step 3, swapping the comorbidity count it learned on for the PRB1 category this study records. Its reasoning follows the shape of the example it was trained on after study 7 ("understand how survival varies across different risk groups" before refining the cohort definitions). Run 0 scores -0.0239 bits per cohort on this study, the baseline -0.0675.
The survival estimator returns, for each group, the share of patients alive at 12, 24 and 36 months, and lists which of the 36 cohorts can be measured in this study and how they split across the groups. Everything the agent submits for the unmeasured cohorts has to come from what it carries from earlier studies.
4. What it learned
After study 7 the teacher kept one step of run 0's analysis as a training example: open the analysis with a three-tier risk split. Run 0 trained on it once. From then on it opened 7 of its next 13 studies the same way, at step 3, each time choosing variables the study records. On those 7 studies it scored above the baseline on 6.
The example, as trained after study 7
The action was chosen by the teacher. The reasoning was written by the learner.
{
"thought": "I have reviewed the database metadata and data summary. The study has a high mortality rate (96.6% died), with a mean time to death of ~21 months. Key variables include age, comorbidity_count, prb1_ratio, and others. To begin estimating survival for the 36 cohorts, I need to understand how survival varies across different risk groups. I will start by estimating survival for a simple risk stratification based on age and comorbidity count, which are likely strong predictors. This will help me calibrate my understanding of survival patterns in this population before moving to more complex cohort definitions.",
"tool_call": {
"tool": "estimate_survival_by_group",
"group_expression": "CASE WHEN age >= 65 AND comorbidity_count >= 4 THEN 'high_risk' WHEN age < 50 AND comorbidity_count <= 1 THEN 'low_risk' ELSE 'medium_risk' END"
}
}
Two things qualify this. The benchmark's own tool description gives an example grouping with tier names like these, so every lane, the baseline included, has the idea in its prompt. The baseline uses it, but late, at steps 12 to 15. And run 1 opened studies 5 to 7 with the same split before any of its training examples targeted it. Run 2 never used it. So the opening move is reachable without training. What the training example changed in run 0 is that the move became its default opening for eight studies.
The table below counts what each lane does, family by family. It is the place to look for what changed in the agents' analyses beyond the opening move.
| per study family (HERALD · MERIDIAN · MOSAIC · FORGE · CADENCE) | baseline | Run 0 | Run 1 | Run 2 |
|---|---|---|---|---|
| studies opened with a risk-tier split (by step 3) | 0 · 0 · 0 · 0 · 0 | 0 · 2 · 3 · 3 · 0 | 0 · 3 · 0 · 0 · 2 | 0 · 0 · 0 · 0 · 0 |
| survival-estimator calls | 33 · 50 · 55 · 53 · 32 | 37 · 53 · 63 · 66 · 66 | 52 · 60 · 51 · 66 · 36 | 40 · 61 · 67 · 67 · 65 |
| SQL queries | 16 · 0 · 1 · 1 · 21 | 14 · 6 · 0 · 0 · 0 | 3 · 0 · 16 · 0 · 0 | 3 · 0 · 0 · 0 · 1 |
| cohort predictions | 12 · 3 · 6 · 5 · 10 | 2 · 2 · 0 · 0 · 5 | 1 · 0 · 0 · 1 · 32 | 1 · 6 · 0 · 0 · 0 |
5. Learner 1.0 learns efficiently online from very few examples
The whole effect in this article comes from 80 to 111 training examples per run, 12,901 to 18,835 supervised tokens, about five examples after each study. Each example was trained once.
The loss on each fresh example, measured in the step that trains on it and before the update, falls in every run: from 0.26, 0.22 and 0.26 on the examples written after studies 1 to 4 to 0.09, 0.10 and 0.08 after studies 17 to 19. The examples written late in the run are closer to what the model already does.
6. By study family
Learner minus baseline, summed over each family's four studies:
| family | run 0 | run 1 | run 2 | mean |
|---|---|---|---|---|
| HERALD | -0.030 | -0.089 | -0.017 | -0.045 |
| MERIDIAN | +0.044 | +0.051 | +0.061 | +0.052 |
| MOSAIC | +0.024 | +0.047 | -0.040 | +0.010 |
| FORGE | +0.287 | +0.027 | +0.208 | +0.174 |
| CADENCE | -0.105 | +0.056 | -0.070 | -0.040 |
MERIDIAN is ahead in every run by a small margin, and FORGE in every run, by a large margin in runs 0 and 2 and a small one in run 1. HERALD, the first family, is behind in every run. MOSAIC and CADENCE change sign between runs. The benchmark splits gain into the first study of each family, where nothing from that family has been seen yet, and the later ones: the runs gain -0.33%, +0.91% and +4.53% on first studies and +6.67%, +1.71% and -0.46% on later ones. Neither part is consistent across runs.
7. Results
| lane | reward, 20 studies | gain over baseline | normalized gain | studies above baseline | first in family / later | examples (supervised tokens) trained |
|---|---|---|---|---|---|---|
| Run 0 | -0.0200 | +0.2208 | +6.34% | 13 / 20 | -0.33% / +6.67% | 108 (16,854) |
| Run 1 | -0.1493 | +0.0915 | +2.63% | 9 / 20 | +0.91% / +1.71% | 111 (18,835) |
| Run 2 | -0.0988 | +0.1420 | +4.07% | 12 / 20 | +4.53% / -0.46% | 80 (12,901) |
| mean of the three runs | -0.0894 ± 0.0376 | +0.1514 | +4.35% ± 1.08% | 12 / 20 | 100 (16,197) | |
| Stateless baseline | -0.2408 | 0 | 0 | 0 |
Each run is compared with the baseline study by study. All three runs score above the baseline in total. The gain is not spread evenly: averaged over the three runs, the learner is above the baseline on 12 of 20 studies, which a one-sided sign test puts at p = 0.25. Most of the gain comes from a few studies with large differences.
| # | study | baseline | run 0 | run 1 | run 2 | mean − baseline |
|---|---|---|---|---|---|---|
| 1 | herald_suburban | -0.0036 | -0.0021 | -0.0485 | -0.0055 | -0.0151 |
| 2 | herald_rural | +0.0281 | +0.0031 | -0.0070 | +0.0260 | -0.0207 |
| 3 | herald_suburban_s2 | -0.0043 | -0.0022 | -0.0070 | +0.0000 | +0.0012 |
| 4 | herald_rural_s2 | +0.0070 | -0.0015 | +0.0004 | -0.0103 | -0.0108 |
| 5 | meridian_suburban | -0.0364 | -0.0032 | -0.0064 | +0.0001 | +0.0332 |
| 6 | meridian_rural | -0.0021 | +0.0137 | +0.0050 | +0.0146 | +0.0132 |
| 7 | meridian_suburban_s2 | +0.0054 | -0.0071 | +0.0038 | +0.0031 | -0.0055 |
| 8 | meridian_rural_s2 | -0.0400 | -0.0325 | -0.0246 | -0.0298 | +0.0110 |
| 9 | mosaic_urban | -0.0675 | -0.0239 | -0.0023 | -0.0122 | +0.0547 |
| 10 | mosaic_suburban | +0.0097 | -0.0036 | +0.0245 | -0.0616 | -0.0233 |
| 11 | mosaic_urban_s2 | -0.0052 | -0.0221 | -0.0312 | -0.0379 | -0.0252 |
| 12 | mosaic_suburban_s2 | -0.0152 | -0.0044 | -0.0226 | -0.0063 | +0.0041 |
| 13 | forge_industrial | -0.0843 | -0.0287 | -0.0315 | -0.0807 | +0.0373 |
| 14 | forge_urban | -0.1253 | -0.0153 | -0.0229 | -0.0081 | +0.1099 |
| 15 | forge_industrial_s2 | -0.0641 | +0.0245 | -0.1715 | -0.0136 | +0.0106 |
| 16 | forge_urban_s2 | -0.0396 | -0.0065 | -0.0601 | -0.0032 | +0.0163 |
| 17 | cadence_urban | +0.1403 | -0.0052 | +0.1317 | -0.0325 | -0.1090 |
| 18 | cadence_industrial | +0.0609 | +0.0729 | +0.0176 | +0.0390 | -0.0177 |
| 19 | cadence_urban_s2 | -0.0266 | +0.0080 | +0.0724 | +0.0899 | +0.0834 |
| 20 | cadence_industrial_s2 | +0.0220 | +0.0161 | +0.0309 | +0.0302 | +0.0037 |
8. Against the published leaderboard
Learner 1.0 is the only system on this leaderboard that learns strictly through its weights. Its context is reset before every study, so no in-context learning is allowed between studies. Every other system carries its experience forward in its context, a notepad, a memory system or a coding agent's workspace.
The benchmark publishes results for twelve systems on this task, each pairing a frontier model with a way of carrying experience forward: in context, through a notepad, through a memory system, or through a coding agent's workspace. The leaderboard ranks them on normalized reward (reward relative to GPT-5.4's stateless baseline) and normalized gain (gain relative to each system's own baseline).
| reward | reward rank of 13 | gain over own baseline | gain rank of 13 | |
|---|---|---|---|---|
| Learner 1.0, mean of three runs | -0.0894 | 2 | +0.1514 | 5 |
| Run 0 alone | -0.0200 | 1 | +0.2208 | 5 |
| Run 1 alone | -0.1493 | 3 | +0.0915 | 5 |
| Run 2 alone | -0.0988 | 2 | +0.1420 | 5 |
| best published on reward: ACE · GPT-5.4 | -0.0518 | 1 | +0.2800 | 4 |
| best published on gain: ICL Notepad · Claude Sonnet 4.6 | -0.7839 | 9 | +0.7953 | 1 |
| system | runs | reward | normalized reward | reward rank | gain | normalized gain | gain rank |
|---|---|---|---|---|---|---|---|
| ACE · GPT-5.4 | 5 | -0.052 | +0.0137 | 1 | +0.280 | +0.0783 | 4 |
| Learner 1.0 (ours) | 3 | -0.089 | +0.0025 | 2 | +0.151 | +0.0435 | 5 |
| ICL · Claude Opus 4.7 | 5 | -0.121 | -0.0069 | 3 | -0.091 | -0.0278 | 8 |
| ICL · GPT-5.4 | 5 | -0.388 | -0.0868 | 4 | -0.290 | -0.0868 | 10 |
| ICL Notepad · GPT-5.4 | 5 | -0.388 | -0.0870 | 5 | -0.059 | -0.0166 | 7 |
| ICL · Gemini 3 Flash | 5 | -0.467 | -0.1106 | 6 | +0.352 | +0.0866 | 3 |
| Mem0 · GPT-5.4 | 5 | -0.468 | -0.1108 | 7 | +0.411 | +0.0998 | 2 |
| Codex · GPT-5.4 | 1 | -0.521 | -0.1266 | 8 | -0.147 | -0.0406 | 9 |
| ICL Notepad · Claude Sonnet 4.6 | 5 | -0.784 | -0.2053 | 9 | +0.795 | +0.1649 | 1 |
| Claude Code · Sonnet 4.6 | 5 | -0.932 | -0.2497 | 10 | -0.421 | -0.1121 | 11 |
| ICL · Claude Sonnet 4.6 | 5 | -0.987 | -0.2662 | 11 | -0.056 | -0.0133 | 6 |
| ICL Notepad · Gemini 3.1 Pro Preview | 5 | -0.991 | -0.2674 | 12 | -0.761 | -0.2190 | 12 |
| ICL · Gemini 3.1 Pro Preview | 5 | -1.597 | -0.4487 | 13 | -1.347 | -0.3857 | 13 |
On reward our three runs average -0.0894, 2nd of 13, behind ACE · GPT-5.4 (-0.0518). Run 0 alone would rank first. On gain they average +0.1514, 5th of 13. The four systems above range from +7.8% to +16.5% normalized gain against our +4.35%. Every system on this task, ours included, scores below zero on reward: none beats the benchmark's reference predictor. Two cautions: the published figures average five runs, ours three, and this is one of the leaderboard's six tasks, so it is a placement on the cohort studies tab, not an overall entry.
9. The recipe
After each study. The teacher receives the learner's steps in the study just finished, with the tool results, and the task brief. It returns procedures in a fixed JSON format: which next action is supported at which state, with the evidence for it. Its full prompt is in the replication repository (recipe/teacher_prompt.txt).
From procedures to examples. Each example is a state the learner actually reached, the conversation so far with its recorded tool results, and the next action as the target. Only states the learner reached are used and no tool result is invented. Two fixed examples are added at every study: read the database metadata first, and after a tool error read the metadata if not yet read.
Reasoning. The teacher chooses the action. The learner writes the reasoning for it, given the state and the action (the prompt is in recipe/reasoning_prompt.txt).
Acting. Thinking when stuck, used on the database exploration task, was not used here.
Training and limits. One pass, batch one, no replay, at most 50,000 supervised tokens per study. The learner's context is 32,000 tokens and each step's output is limited to 4,096 tokens. The teacher call is limited to 15 minutes.
10. Interpretation, and how to check it
Training after each study on about five examples written from that study gives a small gain over the same model with learning off in all three runs, and in run 0 a procedure learned in one study carried to the next ones with different variables. Because neither the learner's prompt nor the teacher's input holds anything from earlier studies, the only path from one study to the next is the weights. The size of the effect varies by run (+2.63% to +6.34% normalized gain).
To check: every step of every study in every lane, every training example, the per-example training losses and the scores are in the replication repository, and the figures regenerate from them.