Learner LabsLearner 1.0Foundation models that continually learn

Demonstrations · Continual Learning Bench

Twenty clinical studies, one model learning between them

On the cohort studies task of Continual Learning Bench, our Learner 1.0 model estimates patient survival from twenty clinical study databases in sequence and is trained after each study on what it did there. Over three runs it scores -0.020, -0.149 and -0.099 (information gain in bits per cohort, summed over the twenty studies) where the same model with learning off scores -0.241. Averaged over the three runs that places it 2nd of 13 on reward and 5th of 13 on gain among the systems on the published leaderboard. The gain is small: the runs beat the baseline on 12 of 20 studies on average.

Model Learner 1.0

Benchmark Continual Learning Bench, cohort studies, default schedule

Studies 20 per run, in 5 study families of 4

Runs 3 study orders (0 canonical, 1 and 2 permuted inside each family) and one stateless baseline

Learning after every finished study. One pass, one example per update, no replay

What the learner sees only the current study, exactly as the benchmark presents it

Teacher a teacher LLM reads only the study just finished and writes the training examples

Compute one H200 per run

Learner 1.0 is the only system on this benchmark that learns strictly through its weights. Its context is reset before every study, so there is no in-context learning between studies: anything it carries from one study to the next is in its weights. Every other system on the public leaderboard carries its experience forward in its context, a notepad, a memory store or a coding agent’s workspace.

1. The question and what was measured

The question: if a model is trained on its own experience after every study it analyses, does it analyse later studies better than the same model that never learns? Each study is a database of patients from one clinical study. The agent has twenty tool actions: read the study's metadata and a data summary, run SQL, estimate survival for any grouping of patients, and fit a cohort predictor. Then it submits survival estimates at 12, 24 and 36 months for 36 fixed cohorts, 108 numbers. Each study records different variables, so only 3 to 10 of the 36 cohorts can be measured in any one study. The rest have to be estimated from what similar studies showed. The twenty studies come in five families of four (HERALD, MERIDIAN, MOSAIC, FORGE, CADENCE) that share a design.

Reward for a study is the information gain of the submitted estimates over the benchmark's reference predictor, in bits per cohort: 0 means as good as the reference, and the best possible is 0.162. Reward is the sum over the twenty studies. Gain is reward minus the reward of the same model analysing the same studies with learning off, the benchmark's stateless baseline. It is the benchmark's learning metric.

Reward as the learner trains on each finished study measured Running average of reward per study, 20 studies in order reward per study = information gain in bits per cohort against the benchmark reference predictor. 0 = as good as the reference, at most 0.162 -0.03 -0.02 -0.01 0 +0.01 +0.02 +0.03 1 4 8 12 16 20 study (position in the run) running average reward HERALD MERIDIAN MOSAIC FORGE CADENCE run 0 -0.0010 run 1 -0.0075 run 2 -0.0049 stateless baseline -0.0120 Three runs of the same recipe in the benchmark default schedule: order 0 canonical, orders 1 and 2 permute the four studies inside each study family. The stateless baseline is the same model with learning off, scored once per study, drawn in order 0. Source: data/trajectories_*.json.
How to read it: reward is measured against the benchmark's reference predictor, so 0 means as good as the reference and higher is better. Every lane stays below 0, and the closer a curve is to 0, the better. All three runs end above the stateless baseline. Final running averages: -0.0010, -0.0075 and -0.0049 bits per cohort for the three runs against -0.0120 for the baseline.

The three runs end at rewards -0.0200, -0.1493 and -0.0988 against -0.2408 for the baseline. The gain over the baseline is +0.2208, +0.0915 and +0.1420. Normalized as the leaderboard does, +6.34%, +2.63% and +4.07%, mean +4.35% (standard error 1.08%).

Gain over the stateless baseline, study by study measured Running average gain per study gain per study = reward of the training run − reward of the stateless baseline on the same study -0.03 -0.02 -0.01 0 +0.01 +0.02 +0.03 1 4 8 12 16 20 study (position in the run) running average gain HERALD MERIDIAN MOSAIC FORGE CADENCE run 0 +0.0110 run 1 +0.0046 run 2 +0.0071 Gain is the CL-Bench learning metric: each study is paired with the same study answered by the same model with learning off, so the difficulty of the study cancels. Source: data/trajectories_*.json.
Run 0 is ahead from study 5 on and run 2 from study 11. Run 1 moves around zero from study 9 and ends just above it. Most of the gain comes in the FORGE family (studies 13–16). Total gain over 20 studies: +0.2208, +0.0915 and +0.1420.

2. Setup

Model. Learner 1.0, the same model as in the ten-skill experiment, with a fixed number of trainable parameters set before the first study. The mechanism is proprietary and is not described here.

What the learner sees. At each study the learner receives the benchmark's own task brief and the current study's tools, and its own steps in that study, nothing else: no earlier studies, no notes, no scores, no retrieved examples. Its prompt is the same as the baseline's. Anything it carries from one study to the next is in its weights.

The loop. The learner analyses a study and submits. A teacher LLM reads the learner's steps in that study and the tool results, and writes procedures: which next action is supported at which state. The teacher sees only the study just finished: no earlier studies, none of its own earlier procedures, and no scores. The recipe turns each procedure into training examples built on states the learner actually reached, the learner writes the reasoning for each example's action in its own words, and it is trained on them before the next study. The teacher never analyses a study and never acts in the benchmark. Section 9 gives the full recipe.

Training. One pass, batch size one, no replay of earlier examples, no optimizer resets between studies. Across a run the learner trained on 80 to 111 examples, 12,901 to 18,835 supervised tokens.

Runs. The benchmark's default schedule plays the twenty studies family by family, in a fixed canonical order (run 0) and in orders that permute the four studies inside each family (runs 1 and 2). The stateless baseline is run once. Because nothing carries over between its studies, its order does not matter. Each run is compared with the baseline study by study. All four lanes used the same model and the same decoding.

3. One study, before and after learning

Here is the same study analysed by the baseline and by run 0, which reached it ninth, after training on eight earlier studies. MOSAIC is a health survey that records a PRB1 category and does not record the comorbidity count the earlier families had.

Stateless baseline

study 9 of 20 in this run · 20 tool calls · reward -0.0675 bits per cohort
  1. metadataget_database_metadata()
    === Study Info ===
      study_name: MOSAIC
      study_label: MOSAIC Survey
      description: Government health survey administered in Urban and Suburban centres. Uses convenience sampling …
    … 17 more lines
  2. summaryget_data_summary()
    Patient count: 129
    patient_id (numeric, n=129): mean=12510.155, std=6883.348, min=375, p25=6262, median=13487, p75=18768, max=232…
    region (categorical, 1 levels): Urban=129 (100.0%)
    time_to_death_months (numeric, n=129): mean=24.149, std=14.428, min=1.25, p25=12.92, median=21.91, p75=33.46, …
    … 11 more lines
  3. SQL
    I now have a good overview of the data. Key observations: 1. 129 patients, all from Urban region 2. High mortality rate: 95.3% died (mean of 'died' column) 3. Median time to death is ~22 months, with max 72 months 4. Variables available for cohort definition: age, sex, prb1_category, genotype_group, ambient_voc_ppb, years_of_education, ad8_score, hhbp_serum_level I need to estimate survival at 12, 24, and 36 months for 36 cohorts. The cohorts will likely be defined by combinations of key variables. Let me first understand the survival patterns by looking at some key variables. I should start by examining survival by age groups, since age is typically a strong predictor of survival. Let me query survival rates by age groups to understand the baseline patterns.
    query_sql(SELECT CASE WHEN age < 40 THEN 'young' WHEN age < 60 THEN 'middle' ELSE 'old' END as age_group, COUNT(*) as n, SUM(died) as deaths, AVG(time_to_death_months) as mean_time FROM patients GROUP BY age_group ORDER BY age_group)
    age_group | n | deaths | mean_time
    middle | 61 | 57 | 20.98
    old | 57 | 55 | 25.97
    young | 11 | 11 | 32.26
    … 1 more lines
  4. estimateestimate_survival_by_group(CASE WHEN age < 40 THEN 'young' WHEN age < 60 THEN 'middle' ELSE 'old' END)
    === Per-Group Survival ===
      middle: n=61 (47.3%)  S(12m)=0.7163  S(24m)=0.3533  S(36m)=0.1413
      old: n=57 (44.2%)  S(12m)=0.8404  S(24m)=0.5722  S(36m)=0.2030
      young: n=11 (8.5%)  S(12m)=0.8182  S(24m)=0.6364  S(36m)=0.4545
    … 35 more lines
  5. estimateestimate_survival_by_group(CASE WHEN sex = 'M' THEN 'male' ELSE 'female' END)
    === Per-Group Survival ===
      female: n=64 (49.6%)  S(12m)=0.7462  S(24m)=0.4281  S(36m)=0.1647
      male: n=65 (50.4%)  S(12m)=0.8133  S(24m)=0.5234  S(36m)=0.2290
    === Cohort Group Decomposition (10 observable, 26 unobservable) ===
    … 34 more lines
  6. steps 6–21 not shown

Learner, run 0

study 9 of 20 in this run · 20 tool calls · reward -0.0239 bits per cohort
  1. metadataget_database_metadata()
    === Study Info ===
      study_name: MOSAIC
      study_label: MOSAIC Survey
      description: Government health survey administered in Urban and Suburban centres. Uses convenience sampling …
    … 17 more lines
  2. summaryget_data_summary()
    Patient count: 129
    patient_id (numeric, n=129): mean=12510.155, std=6883.348, min=375, p25=6262, median=13487, p75=18768, max=232…
    region (categorical, 1 levels): Urban=129 (100.0%)
    time_to_death_months (numeric, n=129): mean=24.149, std=14.428, min=1.25, p25=12.92, median=21.91, p75=33.46, …
    … 11 more lines
  3. estimate
    I have reviewed the database metadata and univariate statistics. The study has 129 patients, all from the Urban region, with a high mortality rate (95.3%). Key variables include age, sex, prb1_category, genotype_group, ambient_voc_ppb, and ad8_score. To estimate survival for the 36 cohorts, I need to understand how these variables interact with survival outcomes. I will start by estimating survival for a broad risk group based on age and PRB1 ratio, as these are likely significant predictors. This will help me gauge the baseline survival patterns before refining the cohort definitions.
    estimate_survival_by_group(CASE WHEN age >= 65 AND prb1_category IN ('high', 'elevated') THEN 'high_risk' WHEN age < 50 THEN 'low_risk' ELSE 'medium_risk' END)
    === Per-Group Survival ===
      high_risk: n=29 (22.5%)  S(12m)=0.9298  S(24m)=0.7152  S(36m)=0.3011
      low_risk: n=34 (26.4%)  S(12m)=0.7647  S(24m)=0.5294  S(36m)=0.3529
      medium_risk: n=66 (51.2%)  S(12m)=0.7221  S(24m)=0.3414  S(36m)=0.0650
    … 35 more lines
  4. estimateestimate_survival_by_group(CASE WHEN genotype_group = 'GG-1' THEN 'GG-1' WHEN genotype_group = 'GG-2' THEN 'GG-2' ELSE 'GG-3' END)
    === Per-Group Survival ===
      GG-1: n=53 (41.1%)  S(12m)=0.7690  S(24m)=0.4141  S(36m)=0.2169
      GG-2: n=41 (31.8%)  S(12m)=0.8049  S(24m)=0.5122  S(36m)=0.1707
      GG-3: n=35 (27.1%)  S(12m)=0.7694  S(24m)=0.5295  S(36m)=0.1986
    … 35 more lines
  5. estimateestimate_survival_by_group(CASE WHEN age >= 60 AND ad8_score >= 3 THEN 'elderly_cognitive_impairment' WHEN age < 60 AND ad8_score < 3 THEN 'younger_cognitive_intact' ELSE 'mixed' END)
    === Per-Group Survival ===
      elderly_cognitive_impairment: n=11 (8.5%)  S(12m)=1.0000  S(24m)=0.6364  S(36m)=0.0909
      mixed: n=87 (67.4%)  S(12m)=0.7802  S(24m)=0.5005  S(36m)=0.2622
      younger_cognitive_intact: n=31 (24.0%)  S(12m)=0.6984  S(24m)=0.3385  S(36m)=0.0376
    … 35 more lines
  6. steps 6–21 not shown

The baseline writes a summary of the data at step 3 and starts from age bands. Run 0 splits the patients into high, medium and low risk tiers at step 3, swapping the comorbidity count it learned on for the PRB1 category this study records. Its reasoning follows the shape of the example it was trained on after study 7 ("understand how survival varies across different risk groups" before refining the cohort definitions). Run 0 scores -0.0239 bits per cohort on this study, the baseline -0.0675.

The survival estimator returns, for each group, the share of patients alive at 12, 24 and 36 months, and lists which of the 36 cohorts can be measured in this study and how they split across the groups. Everything the agent submits for the unmeasured cohorts has to come from what it carries from earlier studies.

4. What it learned

An opening move learned in one study, reused in the next ones measured Step at which the agent first splits patients into high, medium and low risk tiers one dot per study where it happens. No dot = never in that study. Diamonds mark the studies after which a training example targeted this split run 0 run 1 stateless baseline training example on the split 0 5 10 15 20 1 4 8 12 16 20 study (position in the run) step (of up to 20) Run 2 never makes this split. The benchmark's tool description shows an example expression with these tier names, so the move is available to every lane from its prompt. Source: data/trajectories_*.json. Lesson files in the replication repository.
After one training example on this split (after study 7), run 0 opens studies 8, 9 and 11–15 with it at step 3. Run 1 opens studies 5–7 and 18–19 the same way. Its only example on the split came after study 18. The baseline uses the split only late, at steps 12–15, in three studies.

After study 7 the teacher kept one step of run 0's analysis as a training example: open the analysis with a three-tier risk split. Run 0 trained on it once. From then on it opened 7 of its next 13 studies the same way, at step 3, each time choosing variables the study records. On those 7 studies it scored above the baseline on 6.

The example, as trained after study 7

The action was chosen by the teacher. The reasoning was written by the learner.

{
 "thought": "I have reviewed the database metadata and data summary. The study has a high mortality rate (96.6% died), with a mean time to death of ~21 months. Key variables include age, comorbidity_count, prb1_ratio, and others. To begin estimating survival for the 36 cohorts, I need to understand how survival varies across different risk groups. I will start by estimating survival for a simple risk stratification based on age and comorbidity count, which are likely strong predictors. This will help me calibrate my understanding of survival patterns in this population before moving to more complex cohort definitions.",
 "tool_call": {
  "tool": "estimate_survival_by_group",
  "group_expression": "CASE WHEN age >= 65 AND comorbidity_count >= 4 THEN 'high_risk' WHEN age < 50 AND comorbidity_count <= 1 THEN 'low_risk' ELSE 'medium_risk' END"
 }
}

Two things qualify this. The benchmark's own tool description gives an example grouping with tier names like these, so every lane, the baseline included, has the idea in its prompt. The baseline uses it, but late, at steps 12 to 15. And run 1 opened studies 5 to 7 with the same split before any of its training examples targeted it. Run 2 never used it. So the opening move is reachable without training. What the training example changed in run 0 is that the move became its default opening for eight studies.

The table below counts what each lane does, family by family. It is the place to look for what changed in the agents' analyses beyond the opening move.

per study family (HERALD · MERIDIAN · MOSAIC · FORGE · CADENCE)baselineRun 0Run 1Run 2
studies opened with a risk-tier split (by step 3)0 · 0 · 0 · 0 · 00 · 2 · 3 · 3 · 00 · 3 · 0 · 0 · 20 · 0 · 0 · 0 · 0
survival-estimator calls33 · 50 · 55 · 53 · 3237 · 53 · 63 · 66 · 6652 · 60 · 51 · 66 · 3640 · 61 · 67 · 67 · 65
SQL queries16 · 0 · 1 · 1 · 2114 · 6 · 0 · 0 · 03 · 0 · 16 · 0 · 03 · 0 · 0 · 0 · 1
cohort predictions12 · 3 · 6 · 5 · 102 · 2 · 0 · 0 · 51 · 0 · 0 · 1 · 321 · 6 · 0 · 0 · 0

5. Learner 1.0 learns efficiently online from very few examples

Training examples written after each study measured Cumulative supervised tokens trained, per run one training step after each study 1–19. One pass, one example per update, no replay 0 5k 10k 15k 20k 1 4 8 12 16 19 after study n supervised tokens run 0 16,854 total run 1 18,835 total run 2 12,901 total Supervised tokens are the answer tokens the loss is computed on. Run 0 trained nothing after study 10 (an outage, it resumed from the step after study 9). Source: f5_cohort/data/dose_losses.json.
Each run trains on a few hundred to about two thousand supervised tokens after each study: 108, 111 and 80 examples, 16,854, 18,835 and 12,901 tokens over the whole run.

The whole effect in this article comes from 80 to 111 training examples per run, 12,901 to 18,835 supervised tokens, about five examples after each study. Each example was trained once.

Each new lesson surprises the model less as the run goes on measured Training loss on each new example, measured before the model trains on it mean over the examples written after the studies of each family. One pass, so every example is new to the model 0 0.05 0.1 0.15 0.2 0.25 0.3 loss per supervised token studies 1–4 5–8 9–12 13–16 17–19 run 0 0.26 → 0.09 run 1 0.22 → 0.10 run 2 0.26 → 0.08 Loss of each example in the training step that trains on it, before the update (batch 1). Source: f5_cohort/data/dose_losses.json.
In all three runs the loss on a fresh lesson falls by roughly a factor of two to three between the first and the last family: the lessons written late in the run are closer to what the model already does.

The loss on each fresh example, measured in the step that trains on it and before the update, falls in every run: from 0.26, 0.22 and 0.26 on the examples written after studies 1 to 4 to 0.09, 0.10 and 0.08 after studies 17 to 19. The examples written late in the run are closer to what the model already does.

6. By study family

Learner minus baseline, summed over each family's four studies:

familyrun 0run 1run 2mean
HERALD-0.030-0.089-0.017-0.045
MERIDIAN+0.044+0.051+0.061+0.052
MOSAIC+0.024+0.047-0.040+0.010
FORGE+0.287+0.027+0.208+0.174
CADENCE-0.105+0.056-0.070-0.040

MERIDIAN is ahead in every run by a small margin, and FORGE in every run, by a large margin in runs 0 and 2 and a small one in run 1. HERALD, the first family, is behind in every run. MOSAIC and CADENCE change sign between runs. The benchmark splits gain into the first study of each family, where nothing from that family has been seen yet, and the later ones: the runs gain -0.33%, +0.91% and +4.53% on first studies and +6.67%, +1.71% and -0.46% on later ones. Neither part is consistent across runs.

7. Results

lanereward, 20 studiesgain over baselinenormalized gainstudies above baselinefirst in family / laterexamples (supervised tokens) trained
Run 0-0.0200+0.2208+6.34%13 / 20-0.33% / +6.67%108 (16,854)
Run 1-0.1493+0.0915+2.63%9 / 20+0.91% / +1.71%111 (18,835)
Run 2-0.0988+0.1420+4.07%12 / 20+4.53% / -0.46%80 (12,901)
mean of the three runs-0.0894 ± 0.0376+0.1514+4.35% ± 1.08%12 / 20100 (16,197)
Stateless baseline-0.2408000

Each run is compared with the baseline study by study. All three runs score above the baseline in total. The gain is not spread evenly: averaged over the three runs, the learner is above the baseline on 12 of 20 studies, which a one-sided sign test puts at p = 0.25. Most of the gain comes from a few studies with large differences.

Every study, every run measured Learner minus baseline on each of the 20 studies, bits per cohort filled teal = learner above the baseline, open ochre = below. Studies in order 0, paired by study id (runs 1 and 2 met them in their own order) HERALD MERIDIAN MOSAIC FORGE CADENCE run 0 0 -2 0 -1 +3 +2 -1 +1 +4 -1 -2 +1 +6 +11 +9 +3 -15 +1 +3 -1 13/20 run 1 -4 -4 0 -1 +3 +1 0 +2 +7 +1 -3 -1 +5 +10 -11 -2 -1 -4 +10 +1 9/20 run 2 0 0 0 -2 +4 +2 0 +1 +6 -7 -3 +1 0 +12 +5 +4 -17 -2 +12 +1 12/20 mean of 3 -2 -2 0 -1 +3 +1 -1 +1 +5 -2 -3 0 +4 +11 +1 +2 -11 -2 +8 0 12/20 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 Cell value = difference × 100 (hundredths of a bit per cohort). Source: data/trajectories_*.json.
The runs are above the baseline on 13, 9 and 12 of 20 studies. Averaged over the three runs, on 12 of 20 (one-sided sign test p = 0.25). The largest differences sit in a few FORGE and CADENCE studies.
#studybaselinerun 0run 1run 2mean − baseline
1herald_suburban-0.0036-0.0021-0.0485-0.0055-0.0151
2herald_rural+0.0281+0.0031-0.0070+0.0260-0.0207
3herald_suburban_s2-0.0043-0.0022-0.0070+0.0000+0.0012
4herald_rural_s2+0.0070-0.0015+0.0004-0.0103-0.0108
5meridian_suburban-0.0364-0.0032-0.0064+0.0001+0.0332
6meridian_rural-0.0021+0.0137+0.0050+0.0146+0.0132
7meridian_suburban_s2+0.0054-0.0071+0.0038+0.0031-0.0055
8meridian_rural_s2-0.0400-0.0325-0.0246-0.0298+0.0110
9mosaic_urban-0.0675-0.0239-0.0023-0.0122+0.0547
10mosaic_suburban+0.0097-0.0036+0.0245-0.0616-0.0233
11mosaic_urban_s2-0.0052-0.0221-0.0312-0.0379-0.0252
12mosaic_suburban_s2-0.0152-0.0044-0.0226-0.0063+0.0041
13forge_industrial-0.0843-0.0287-0.0315-0.0807+0.0373
14forge_urban-0.1253-0.0153-0.0229-0.0081+0.1099
15forge_industrial_s2-0.0641+0.0245-0.1715-0.0136+0.0106
16forge_urban_s2-0.0396-0.0065-0.0601-0.0032+0.0163
17cadence_urban+0.1403-0.0052+0.1317-0.0325-0.1090
18cadence_industrial+0.0609+0.0729+0.0176+0.0390-0.0177
19cadence_urban_s2-0.0266+0.0080+0.0724+0.0899+0.0834
20cadence_industrial_s2+0.0220+0.0161+0.0309+0.0302+0.0037

8. Against the published leaderboard

Learner 1.0 is the only system on this leaderboard that learns strictly through its weights. Its context is reset before every study, so no in-context learning is allowed between studies. Every other system carries its experience forward in its context, a notepad, a memory system or a coding agent's workspace.

The benchmark publishes results for twelve systems on this task, each pairing a frontier model with a way of carrying experience forward: in context, through a notepad, through a memory system, or through a coding agent's workspace. The leaderboard ranks them on normalized reward (reward relative to GPT-5.4's stateless baseline) and normalized gain (gain relative to each system's own baseline).

rewardreward rank of 13gain over own baselinegain rank of 13
Learner 1.0, mean of three runs-0.08942+0.15145
Run 0 alone-0.02001+0.22085
Run 1 alone-0.14933+0.09155
Run 2 alone-0.09882+0.14205
best published on reward: ACE · GPT-5.4-0.05181+0.28004
best published on gain: ICL Notepad · Claude Sonnet 4.6-0.78399+0.79531
Cohort studies: our runs against the published leaderboard measured Reward over 20 studies (mean ± standard error) -1.5 -1 -0.5 0 ACE · GPT-5.4 -0.05 Learner 1.0 (ours, 3 runs) -0.09 ICL · Claude Opus 4.7 -0.12 ICL · GPT-5.4 -0.39 ICL Notepad · GPT-5.4 -0.39 ICL · Gemini 3 Flash -0.47 Mem0 · GPT-5.4 -0.47 Codex · GPT-5.4 -0.52 ICL Notepad · Claude Sonnet 4.6 -0.78 Claude Code · Sonnet 4.6 -0.93 ICL · Claude Sonnet 4.6 -0.99 ICL Notepad · Gemini 3.1 Pro Preview -0.99 ICL · Gemini 3.1 Pro Preview -1.60 Gain over the system's own stateless baseline -1 -0.5 0 0.5 ICL Notepad · Claude Sonnet 4.6 +0.80 Mem0 · GPT-5.4 +0.41 ICL · Gemini 3 Flash +0.35 ACE · GPT-5.4 +0.28 Learner 1.0 (ours, 3 runs) +0.15 ICL · Claude Sonnet 4.6 -0.06 ICL Notepad · GPT-5.4 -0.06 ICL · Claude Opus 4.7 -0.09 Codex · GPT-5.4 -0.15 ICL · GPT-5.4 -0.29 Claude Code · Sonnet 4.6 -0.42 ICL Notepad · Gemini 3.1 Pro Preview -0.76 ICL · Gemini 3.1 Pro Preview -1.35 Published systems: CL-Bench leaderboard data of 2026-07-18, cohort_studies task, 5 runs each (Codex · GPT-5.4: 1 run). Ours: 3 runs (orders 0, 1, 2) and one stateless baseline. Gain uses each system's own stateless baseline. One published standard error (ICL · Gemini 3.1 Pro Preview, ±1.53) is clipped at the axis. Sources: data/leaderboard_data_2026-07-18.json, f5_cohort/data/three_run_analysis.json.
On reward our three runs average -0.0894, 2nd of 13 behind ACE · GPT-5.4 (-0.0518). On gain, the improvement over the same model with learning off, they average +0.1514, 5th of 13.
systemrunsrewardnormalized rewardreward rankgainnormalized gaingain rank
ACE · GPT-5.45-0.052+0.01371+0.280+0.07834
Learner 1.0 (ours)3-0.089+0.00252+0.151+0.04355
ICL · Claude Opus 4.75-0.121-0.00693-0.091-0.02788
ICL · GPT-5.45-0.388-0.08684-0.290-0.086810
ICL Notepad · GPT-5.45-0.388-0.08705-0.059-0.01667
ICL · Gemini 3 Flash5-0.467-0.11066+0.352+0.08663
Mem0 · GPT-5.45-0.468-0.11087+0.411+0.09982
Codex · GPT-5.41-0.521-0.12668-0.147-0.04069
ICL Notepad · Claude Sonnet 4.65-0.784-0.20539+0.795+0.16491
Claude Code · Sonnet 4.65-0.932-0.249710-0.421-0.112111
ICL · Claude Sonnet 4.65-0.987-0.266211-0.056-0.01336
ICL Notepad · Gemini 3.1 Pro Preview5-0.991-0.267412-0.761-0.219012
ICL · Gemini 3.1 Pro Preview5-1.597-0.448713-1.347-0.385713

On reward our three runs average -0.0894, 2nd of 13, behind ACE · GPT-5.4 (-0.0518). Run 0 alone would rank first. On gain they average +0.1514, 5th of 13. The four systems above range from +7.8% to +16.5% normalized gain against our +4.35%. Every system on this task, ours included, scores below zero on reward: none beats the benchmark's reference predictor. Two cautions: the published figures average five runs, ours three, and this is one of the leaderboard's six tasks, so it is a placement on the cohort studies tab, not an overall entry.

9. The recipe

After each study. The teacher receives the learner's steps in the study just finished, with the tool results, and the task brief. It returns procedures in a fixed JSON format: which next action is supported at which state, with the evidence for it. Its full prompt is in the replication repository (recipe/teacher_prompt.txt).

From procedures to examples. Each example is a state the learner actually reached, the conversation so far with its recorded tool results, and the next action as the target. Only states the learner reached are used and no tool result is invented. Two fixed examples are added at every study: read the database metadata first, and after a tool error read the metadata if not yet read.

Reasoning. The teacher chooses the action. The learner writes the reasoning for it, given the state and the action (the prompt is in recipe/reasoning_prompt.txt).

Acting. Thinking when stuck, used on the database exploration task, was not used here.

Training and limits. One pass, batch one, no replay, at most 50,000 supervised tokens per study. The learner's context is 32,000 tokens and each step's output is limited to 4,096 tokens. The teacher call is limited to 15 minutes.

10. Interpretation, and how to check it

Training after each study on about five examples written from that study gives a small gain over the same model with learning off in all three runs, and in run 0 a procedure learned in one study carried to the next ones with different variables. Because neither the learner's prompt nor the teacher's input holds anything from earlier studies, the only path from one study to the next is the weights. The size of the effect varies by run (+2.63% to +6.34% normalized gain).

To check: every step of every study in every lane, every training example, the per-example training losses and the scores are in the replication repository, and the figures regenerate from them.

Cohort studies · Learner Labs