Learner LabsLearner 1.0Foundation models that continually learn

Demonstrations · Continual Learning Bench

Forty database questions, one model learning as it answers

On the database exploration task of Continual Learning Bench, our Learner 1.0 model answers forty questions about one product database in sequence and is trained on each question after it answers it. Over three runs it answers 21, 15 and 16 questions correctly where the same model with learning off answers 7.

Model Learner 1.0

Benchmark Continual Learning Bench, database exploration, default schedule

Questions 40 per run, 20 before and 20 after a schema migration

Runs 3 question orders (0 canonical, 1 and 2 permuted) and one stateless baseline

Learning after every answered question. One pass, one example per update, no replay

What the learner sees only the current question, exactly as the benchmark presents it

Teacher a teacher LLM writes the training examples after each question

Compute one H200 per run

Learner 1.0 is the only system on this benchmark that learns strictly through its weights. Its context is reset before every question, so there is no in-context learning between questions: anything it carries from one question to the next is in its weights. Every other system on the public leaderboard carries its experience forward in its context, a notepad, a memory store or a coding agent’s workspace.

1. The question and what was measured

The question: can our continual learning agent learn from its own experience by updating its weights? There are 40 questions in the benchmark, and the agent must learn from its own experience incrementally and sequentially, and improve its performance. We use only weight-based learning and reset the context between questions, so there is no in-context learning between questions. If the agent improves over the baseline, it must be due to online parametric continual learning.

There are two challenges. 1) Data efficiency: each question of the benchmark generates very little data, so the agent must learn efficiently from just a handful of examples. 2) It must retain its learnings in the weights as it is incrementally trained. As per our understanding, no other technique today can achieve these two traits when it comes to parametric continual learning.

Each question is a natural-language question about a SQLite database of three product datasets (office products, electronics, musical instruments). The agent may run up to 15 exploratory queries, then submits one answer, and is told whether it was right and, if not, the correct answer. Reward for a question is 1 − queries/15 if the answer is correct and 0 if it is wrong. After question 20 the database is migrated to a new schema, and questions 21–40 are asked against it.

Reward is the sum over the forty questions (at most 40). Gain is reward minus the reward of the same model answering the same questions with learning off, which the benchmark calls the stateless baseline. Gain is the benchmark's learning metric.

Reward rises as the learner trains on each answered question measured Running average of reward per question, 40 questions in order reward per question = 1 − queries/15 if correct, else 0. Mean over the questions answered so far 0 0.1 0.2 0.3 0.4 0.5 1 5 10 15 20 25 30 35 40 question (position in the run) running average reward database migration run 0 0.303 run 1 0.182 run 2 0.267 stateless baseline 0.047 Three runs of the same recipe in three question orders (the benchmark default schedule: order 0 canonical, orders 1 and 2 permuted within each half). The stateless baseline is the same model with learning off and the same decoding, scored once per question. Questions 21–40 follow a schema migration of the database. Source: data/de_lanes.json.
All three runs end above the stateless baseline. Final running averages: 0.303, 0.182 and 0.267 for runs 0, 1 and 2 against 0.047 for the baseline.

The three runs end at rewards 12.13, 7.27 and 10.67 against 1.87 for the baseline.

Gain over the stateless baseline measured Running average gain per question gain per question = reward of the training run − reward of the stateless baseline on the same question -0.2 -0.1 0 0.1 0.2 0.3 0.4 0.5 1 5 10 15 20 25 30 35 40 question (position in the run) running average gain database migration run 0 +0.257 run 1 +0.135 run 2 +0.220 Gain is the CL-Bench learning metric: each question is paired with the same question answered by the same model with learning off, so the difficulty of the question cancels. Source: data/de_lanes.json.
The total gain over 40 questions is +10.27, +5.40 and +8.80 for runs 0, 1 and 2.

2. Setup

Model. Learner 1.0, the same model as in the ten-skill experiment, with a fixed number of trainable parameters set before the first question. The mechanism is proprietary and is not described here.

What the learner sees. At each question the learner receives the benchmark's own system prompt and the current question, and nothing else: no earlier questions, no notes, no retrieved examples. Its prompt is the same as the baseline's. Anything it carries from one question to the next is in its weights.

Thinking when stuck. Before each query the learner normally writes its next action directly. When it is stuck, it first writes up to 512 tokens of reasoning that the benchmark does not see, and then the action. It counts as stuck when its last query returned an error or no rows, or when it has used 8 or more of its 15 queries on the question. The rule reads only the text the benchmark returns, and the baseline uses the same rule. The learner thought before 53, 66 and 39 of its 299, 291 and 228 actions in the three runs. The baseline before 198 of 522.

The loop. The learner answers a question. The benchmark scores it and releases its feedback. A teacher LLM reads the learner's trajectory on that question and the feedback, and writes procedures: what the question needed and which queries answer it. The recipe turns each procedure into training examples built on states the learner actually reached, and the learner is trained on them before the next question. The teacher never answers a question and never acts in the benchmark.

Training. One pass, batch size one, no replay of earlier examples, no optimizer resets between questions. Across a run the learner trained on 240 to 539 examples and 20,547 to 45,096 supervised tokens.

Runs. The benchmark's default schedule asks the same forty questions in a fixed canonical order (run 0) and in permuted orders that keep the twenty pre-migration questions first (runs 1 and 2). The stateless baseline is run once. Nothing carries over between its questions, so its order does not matter.

3. A few example questions compared, before and after learning

Three questions asked before the migration, so both agents faced the same database. Each is one the baseline gets wrong and run 0 gets right.

For each question the two columns show the agent's actions in order. Each numbered step is one SQL query the agent sent, with the start of the result the database returned below it. The last line is the answer it submitted and whether it was right. The left column is the stateless baseline, the same model with learning off. The right column is the learner. The full trajectories of every run, all forty questions each, are in the replication repository.

Question. What is the average review rating given to office products that cost more than $50?

Stateless baseline

question 19 of 40 in this run · 4 queries
  1. list.tables
    attrs_g1
    attrs_g3
    fdbk_g1
    … 9 more lines
  2. listPRAGMA table_info(items_g1);
    cid | name        | type    | notnull | dflt_value | pk
    0   | ref_id      | TEXT    | 0       | NULL       | 1
    1   | ttl         | TEXT    | 0       | NULL       | 0
    … 8 more lines
  3. sampleSELECT DISTINCT main_cat FROM items_g1 LIMIT 10;
    main_cat
    All Electronics
    Amazon Home
    … 8 more lines
  4. computeSELECT AVG(avg_rtg) FROM items_g1 WHERE main_cat = 'Office Products' AND prc > 50;
    AVG(avg_rtg)
    4.16020942408377
answered 4.16020942408377 · wrong (correct: 4.11)

Learner, run 0

question 19 of 40 in this run · 3 queries
  1. list.tables
    attrs_g1
    attrs_g3
    fdbk_g1
    … 9 more lines
  2. read schemaSELECT name, sql FROM sqlite_master WHERE type='table' AND name LIKE '%g1%'
    name          | sql
    items_g1      | CREATE TABLE items_g1 (     ref_id TEXT PRIMARY KEY,     ttl TEXT,     avg_rtg REAL,     avg_rtg_30d REA…
    fdbk_g1       | CREATE TABLE fdbk_g1 (     id INTEGER PRIMARY KEY AUTOINCREMENT,     item_id TEXT,     ref_id TEXT,     …
    … 3 more lines
  3. computeSELECT AVG(rtg) FROM fdbk_g1 WHERE ref_id IN (SELECT ref_id FROM items_g1 WHERE prc > 50.0);
    AVG(rtg)
    4.105430924150588
answered 4.105430924150588 · correct

Question. Among electronics reviews with a review body longer than 500 characters, considering only products whose list price is at most $40, what percentage are 1-star? Round to 2 decimal places.

Stateless baseline

question 16 of 40 in this run · 15 queries
  1. list.tables
    attrs_g1
    attrs_g3
    fdbk_g1
    … 9 more lines
  2. listPRAGMA table_info(items_g1);
    cid | name        | type    | notnull | dflt_value | pk
    0   | ref_id      | TEXT    | 0       | NULL       | 1
    1   | ttl         | TEXT    | 0       | NULL       | 0
    … 8 more lines
  3. sampleSELECT main_cat FROM items_g1 LIMIT 5;
    main_cat
    All Electronics
    Amazon Home
    … 3 more lines
  4. sampleSELECT main_cat FROM items_g2 LIMIT 5;
    main_cat
    Computers
    Computers
    … 3 more lines
  5. sampleSELECT main_cat FROM items_g3 LIMIT 5;
    main_cat
    Musical Instruments
    Musical Instruments
    … 3 more lines
  6. listPRAGMA table_info(fdbk_g1);
    cid | name    | type    | notnull | dflt_value | pk
    0   | id      | INTEGER | 0       | NULL       | 1
    1   | item_id | TEXT    | 0       | NULL       | 0
    … 8 more lines
  7. listPRAGMA table_info(fdbk_g2);
    cid | name    | type    | notnull | dflt_value | pk
    0   | id      | INTEGER | 0       | NULL       | 1
    1   | item_id | TEXT    | 0       | NULL       | 0
    … 8 more lines
  8. listPRAGMA table_info(fdbk_g3);
    cid | name   | type    | notnull | dflt_value | pk
    0   | id     | INTEGER | 0       | NULL       | 1
    1   | ref_id | TEXT    | 0       | NULL       | 0
    … 7 more lines
  9. sampleSELECT DISTINCT main_cat FROM items_g1;
    main_cat
    All Electronics
    Amazon Home
    … 31 more lines
  10. sampleSELECT DISTINCT main_cat FROM items_g2;
    main_cat
    Computers
    Cell Phones & Accessories
    … 31 more lines
  11. computeSELECT COUNT(*) FROM items_g2 WHERE main_cat = 'All Electronics';
    COUNT(*)
    4562
  12. computeSELECT COUNT(*) FROM items_g1 WHERE main_cat = 'Office Products';
    COUNT(*)
    13933
  13. computeSELECT COUNT(*) FROM items_g3 WHERE main_cat = 'Musical Instruments';
    COUNT(*)
    17162
  14. computeSELECT COUNT(*) FROM fdbk_g2 WHERE LENGTH(body) > 500;
    COUNT(*)
    9875
  15. computeSELECT ROUND(100.0 * SUM(CASE WHEN f.rtg = 1 THEN 1 ELSE 0 END) / COUNT(*), 2) FROM fdbk_g2 f JOIN items_g2 i ON f.ref_id = i.ref_id WHERE i.main_cat = 'All Electronics' AND i.prc <= 40 AND LENGTH(f.body) > 500;
    ROUND(100.0 * SUM(CASE WHEN f.rtg = 1 THEN 1 ELSE 0 END) / COUNT(*), 2)
    NULL
answered 0.00 · wrong (correct: 10.22)

Learner, run 0

question 16 of 40 in this run · 3 queries
  1. read schemaSELECT name, sql FROM sqlite_master WHERE type='table'
    name            | sql
    items_g1        | CREATE TABLE items_g1 (     ref_id TEXT PRIMARY KEY,     ttl TEXT,     avg_rtg REAL,     avg_rtg_30d R…
    fdbk_g1         | CREATE TABLE fdbk_g1 (     id INTEGER PRIMARY KEY AUTOINCREMENT,     item_id TEXT,     ref_id TEXT,   …
    … 10 more lines
  2. computeSELECT COUNT(*) FROM fdbk_g2 WHERE LENGTH(body) > 500 AND rtg = 1.0 AND ref_id IN (SELECT ref_id FROM items_g2 WHERE prc_usd <= 40.0);
    COUNT(*)
    222
  3. computeSELECT COUNT(*) FROM fdbk_g2 WHERE LENGTH(body) > 500 AND ref_id IN (SELECT ref_id FROM items_g2 WHERE prc_usd <= 40.0);
    COUNT(*)
    2172
answered 10.22 · correct

Question. Among electronics products that have zero reviews, what percentage have a price listed (non-NULL, non-zero)? Round to 2 decimal places.

Stateless baseline

question 13 of 40 in this run · 15 queries
  1. list.tables
    attrs_g1
    attrs_g3
    fdbk_g1
    … 9 more lines
  2. listPRAGMA table_info(items_g1);
    cid | name        | type    | notnull | dflt_value | pk
    0   | ref_id      | TEXT    | 0       | NULL       | 1
    1   | ttl         | TEXT    | 0       | NULL       | 0
    … 8 more lines
  3. computeSELECT main_cat, COUNT(*) FROM items_g1 GROUP BY main_cat LIMIT 10;
    main_cat              | COUNT(*)
    NULL                  | 741
    AMAZON FASHION        | 224
    All Beauty            | 25
    All Electronics       | 1308
    Amazon Fire TV        | 1
    … 5 more lines
  4. computeSELECT main_cat, COUNT(*) FROM items_g2 GROUP BY main_cat LIMIT 10;
    main_cat              | COUNT(*)
    NULL                  | 1304
    AMAZON FASHION        | 232
    All Beauty            | 25
    All Electronics       | 4562
    Amazon Devices        | 106
    … 5 more lines
  5. computeSELECT main_cat, COUNT(*) FROM items_g3 GROUP BY main_cat LIMIT 10;
    main_cat              | COUNT(*)
    NULL                  | 321
    AMAZON FASHION        | 177
    All Beauty            | 16
    All Electronics       | 534
    Amazon Home           | 253
    … 5 more lines
  6. computeSELECT COUNT(*) FROM items_g1 WHERE main_cat = 'All Electronics' AND rtg_ct = 0;
    COUNT(*)
    0
  7. computeSELECT COUNT(*) FROM items_g2 WHERE main_cat = 'All Electronics' AND rtg_ct = 0;
    ERROR: no such column: rtg_ct
  8. sampleSELECT * FROM taxn_g1 LIMIT 10;
    ref_id     | cat_lvl | cat_nm
    B001DZWCDO | 0       | Office Products
    B001DZWCDO | 1       | Office Electronics
    … 8 more lines
  9. listPRAGMA table_info(items_g2);
    cid | name     | type    | notnull | dflt_value | pk
    0   | ref_id   | TEXT    | 0       | NULL       | 1
    1   | ttl      | TEXT    | 0       | NULL       | 0
    … 7 more lines
  10. listPRAGMA table_info(items_g3);
    cid | name     | type    | notnull | dflt_value | pk
    0   | ref_id   | TEXT    | 0       | NULL       | 1
    1   | ttl      | TEXT    | 0       | NULL       | 0
    … 6 more lines
  11. sampleSELECT cat_nm FROM taxn_g2 WHERE cat_lvl = 0 LIMIT 5;
    cat_nm
    Electronics
    Electronics
    … 3 more lines
  12. listPRAGMA table_info(fdbk_g2);
    cid | name    | type    | notnull | dflt_value | pk
    0   | id      | INTEGER | 0       | NULL       | 1
    1   | item_id | TEXT    | 0       | NULL       | 0
    … 8 more lines
  13. computeSELECT COUNT(*) FROM items_g2 WHERE main_cat = 'All Electronics';
    COUNT(*)
    4562
  14. computeSELECT COUNT(*) FROM items_g2 WHERE main_cat = 'All Electronics' AND avg_rtg IS NULL;
    COUNT(*)
    0
  15. repeatSELECT COUNT(*) FROM items_g2 WHERE main_cat = 'All Electronics' AND avg_rtg IS NULL;
    COUNT(*)
    0
answered 0.00 · wrong (correct: 29.93)

Learner, run 0

question 13 of 40 in this run · 3 queries
  1. read schemaSELECT name, sql FROM sqlite_master WHERE type='table'
    name            | sql
    items_g1        | CREATE TABLE items_g1 (     ref_id TEXT PRIMARY KEY,     ttl TEXT,     avg_rtg REAL,     avg_rtg_30d R…
    fdbk_g1         | CREATE TABLE fdbk_g1 (     id INTEGER PRIMARY KEY AUTOINCREMENT,     item_id TEXT,     ref_id TEXT,   …
    … 10 more lines
  2. computeSELECT COUNT(*) FROM items_g2 WHERE ref_id NOT IN (SELECT ref_id FROM fdbk_g2) AND prc IS NOT NULL AND prc > 0;
    COUNT(*)
    3395
  3. computeSELECT COUNT(*) FROM items_g2 WHERE ref_id NOT IN (SELECT ref_id FROM fdbk_g2);
    COUNT(*)
    11345
answered 29.93 · correct

4. What it learned

The figure draws every step of every question, in the order the run answered them, for the baseline and for run 0. Each row is a question. Each cell is one query, colored by what the query does.

Every step of every question: learning off and learning on measured read the full schema in one query list tables or columns look at sample rows compute toward the answer repeat an earlier identical query answer: correct (filled) or wrong (open) Stateless baseline (learning off) 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 migration 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 7 of 40 correct · 479 queries Learner, run 0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 migration 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 21 of 40 correct · 258 queries Each row is one question, in the order the run answered it (row labels are positions 1–40). Each cell is one action, left to right. Up to 15 queries are allowed before the answer. The dashed line marks the schema migration after question 20. Classification of each query is by its text (scripts/step_classes.py). Source: data/traces.json.
The baseline answers 7 of 40 questions with 479 queries. Run 0 answers 21 with 258.
The other two runs measured read the full schema in one query list tables or columns look at sample rows compute toward the answer repeat an earlier identical query answer: correct (filled) or wrong (open) Learner, run 1 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 migration 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 15 of 40 correct · 248 queries Learner, run 2 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 migration 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 16 of 40 correct · 187 queries Each row is one question, in the order the run answered it (row labels are positions 1–40). Each cell is one action, left to right. Up to 15 queries are allowed before the answer. The dashed line marks the schema migration after question 20. Classification of each query is by its text (scripts/step_classes.py). Source: data/traces.json.
Run 1 answers 15 of 40 with 248 queries. Run 2 answers 16 with 187.

Counted per block of ten questions. A sample-row query looks at a few raw rows of a table (for example SELECT … LIMIT 5) instead of computing an answer. A schema read asks the database for its table definitions. A repeat is a query identical to one already sent on the same question.

per block of ten questions (positions 1–10, 11–20, 21–30, 31–40)baselinerun 0run 1run 2
questions with a schema read in the first two queries0 · 0 · 0 · 00 · 6 · 5 · 51 · 6 · 6 · 102 · 10 · 10 · 10
sample-row queries36 · 40 · 30 · 4427 · 6 · 3 · 133 · 3 · 2 · 030 · 1 · 7 · 0
queries in total108 · 113 · 133 · 12586 · 63 · 52 · 57103 · 47 · 61 · 3783 · 40 · 42 · 22
repeats of an identical earlier query5 · 6 · 5 · 23 · 4 · 0 · 53 · 3 · 4 · 04 · 4 · 4 · 0

From the second block of ten questions on, every run sends about half as many queries as the baseline or fewer, and the baseline never improves. In the last block the three runs send 57, 37 and 22 queries against the baseline's 125. Looking at sample rows almost disappears after the first ten questions, from 27 to 33 per block to 0 to 7, and reading the schema at the start of a question goes from 0 to 2 of the first ten questions to 5 to 10 of each later ten.

A lesson and where it was used

Run 0's 24th question was about musical-instrument brands, and it got it wrong. It grouped products by the catalog's brand column, which the migration had left empty for instruments, found no brand with fifty products, tried the other product groups' tables, and answered Yamaha (correct: Fender).

Learner, run 0, question 24

question 24 of 40 in this run · 6 queries
  1. list.tables
    attrs_g1
    attrs_g3_legacy
    fdbk_g1
    … 10 more lines
  2. listPRAGMA table_info(items_g3);
    cid | name     | type    | notnull | dflt_value | pk
    0   | ref_id   | TEXT    | 0       | NULL       | 1
    1   | ttl      | TEXT    | 0       | NULL       | 0
    … 6 more lines
  3. computeSELECT str_nm, COUNT(DISTINCT ref_id) AS n_products, AVG(prc) AS avg_price FROM items_g3 GROUP BY str_nm HAVING COUNT(DISTINCT ref_id) >= 50 ORDER BY avg_price DESC, str_nm ASC LIMIT 10
    str_nm | n_products | avg_price
    NULL   | 20000      | 123.38623097244937
  4. computeSELECT str_nm, COUNT(DISTINCT ref_id) AS n_products, AVG(prc) AS avg_price FROM items_g3 WHERE str_nm IS NOT NULL GROUP BY str_nm HAVING COUNT(DISTINCT ref_id) >= 50 ORDER BY avg_price DESC, str_nm ASC LIMIT 10
    (no results)
  5. computeSELECT str_nm, COUNT(DISTINCT ref_id) AS n_products, AVG(prc) AS avg_price FROM items_g1 WHERE str_nm IS NOT NULL GROUP BY str_nm HAVING COUNT(DISTINCT ref_id) >= 50 ORDER BY avg_price DESC, str_nm ASC LIMIT 10
    str_nm         | n_products | avg_price
    HP             | 104        | 374.79033898305084
    Canon          | 66         | 227.96085714285715
    Amazon Renewed | 66         | 187.9235294117647
    Brother        | 66         | 137.94026315789472
    Epson          | 65         | 123.12
    … 5 more lines
  6. computeSELECT str_nm, COUNT(DISTINCT ref_id) AS n_products, AVG(prc) AS avg_price FROM items_g2 WHERE str_nm IS NOT NULL GROUP BY str_nm HAVING COUNT(DISTINCT ref_id) >= 50 ORDER BY avg_price DESC, str_nm ASC LIMIT 10
    str_nm         | n_products | avg_price
    Canon          | 54         | 67578.72222222222
    Lenovo         | 108        | 66982.97674418605
    ASUS           | 82         | 49604.0
    HP             | 206        | 44155.68292682927
    Dell           | 73         | 42321.31428571429
    … 5 more lines
answered Yamaha · wrong (correct: Fender)

After that question the teacher wrote this procedure (verbatim):

Musical-instrument brand ranking (brands with at least N distinct products, highest average list price, alphabetical tie-break). The items_g3.str_nm brand column gives no brand with N or more products, so find where brand is recorded (the g3 key-value attribute tables) before computing or answering. Do not switch to other categories' tables or answer with an uncomputed guess.

The recipe built 6 training examples that query the migration's new attribute table, 522 supervised tokens in total. Nine questions later run 0 was asked how many instrument products have brand information in the current attribute feed but not in the old export. After a table listing and one exploratory count it read the instrument schema, then counted the rows with the key Brand in the new attribute table: 1145, correct, in 4 queries.

Learner, run 0, question 33

question 33 of 40 in this run · 4 queries
  1. list.tables
    attrs_g1
    attrs_g3_legacy
    fdbk_g1
    … 10 more lines
  2. computeSELECT COUNT(*) AS n FROM items_g3 WHERE category LIKE '%Musical%' OR category LIKE '%Instrument%' LIMIT 1
    ERROR: no such column: category
  3. read schemaSELECT name, sql FROM sqlite_master WHERE type='table' AND name LIKE '%g3%';
    name                  | sql
    items_g3              | CREATE TABLE items_g3 (     ref_id TEXT PRIMARY KEY,     ttl TEXT,     avg_rtg REAL,     rtg_ct …
    fdbk_g3               | CREATE TABLE fdbk_g3 (     id INTEGER PRIMARY KEY AUTOINCREMENT,     ref_id TEXT,     uid TEXT, …
    … 2 more lines
  4. computeSELECT COUNT(*) AS n FROM product_attributes_g3 WHERE attr_key = 'Brand' AND ref_id NOT IN (SELECT ref_id FROM attrs_g3_legacy WHERE attr_key = 'Brand')
    n
    1145
answered 1145 · correct

Run 0 used that table in 5 later questions and answered all 5 correctly. Before question 24 it had not written the table's name in any query.

5. Learner 1.0 learns efficiently online from very few examples

A run trains once after each question, on examples built from that question alone. Over forty questions the three runs trained on 539, 428 and 240 examples, 45,096, 37,579 and 20,547 supervised tokens: on average 13.5, 10.7 and 6.0 examples and 1,127, 939 and 514 supervised tokens per question. The table lists names each run did not write in any query until a training example used them. Each name was already visible in the schema listings the learner had read. What the examples taught was to use it.

Training examples written after each question measured Cumulative supervised tokens trained, per run one boundary after each answered question 1–39. One pass, one example per update, no replay 0 10,000 20,000 30,000 40,000 1 5 10 15 20 25 30 35 39 boundary (after question n) supervised tokens run 0 45,096 total run 1 37,579 total run 2 20,547 total Supervised tokens are the target tokens the loss is computed on. Prompt tokens are not counted. Source: data/de_lanes.json (training).
Over the forty questions the three runs train on 45,096, 37,579 and 20,547 supervised tokens, a few hundred to a few thousand after each question. Run 0, which answered the most questions correctly, trained on 45,096.

We never train on a question paired with its answer. Every training example is a step of the learner's own attempt at a question: the question, the queries it sent so far and the database's real results, with the next action as the target. Most examples, 989 of the 1,207 across the three runs, teach the next query. The other 218 end in submitting an answer, and that value follows from the query results shown in the same example. Training happens only after a question has been answered, and each of the forty questions is asked once. Every training example is in the replication repository.

runnamewhat it istaught after questionexamplessupervised tokenslater questions using itcorrect
run 0review_yearyear column added to musical-instrument reviews by the migration22354322
run 0attrs_g3_legacythe old musical-instrument attribute table, renamed by the migration24652244
run 0product_attributes_g3musical-instrument attribute table added by the migration24652255
run 1cat_lvltaxonomy level column771,51921
run 1cat_nmtaxonomy category name column771,51921
run 1prc_usdelectronics price in dollars1251,30062
run 1attrs_g1office product attribute table14225421
run 1verified_statusoffice review verification label added by the migration23110422
run 1review_yearyear column added to musical-instrument reviews by the migration24244021
run 1attrs_g3_legacythe old musical-instrument attribute table, renamed by the migration25475211
run 1product_attributes_g3musical-instrument attribute table added by the migration25475232
run 2taxn_g1office product taxonomy table2518042
run 2fdbk_g2electronics review table2518096
run 2fdbk_g3musical-instrument review table2518073
run 2item_idsecond product identifier in office reviews451,32011
run 2cat_lvltaxonomy level column451,32020
run 2cat_nmtaxonomy category name column451,32020

Learning the rule, not memorizing the examples. Each training example is used for one update and never again. Its loss is recorded in that same step, before the update, so every point below is the model's error on an example it has not trained on yet. A falling line means the model gets better at examples it has never trained on: it is learning how to act on this database, not memorizing the examples it was shown. The examples written after question 30 come from question 30, which the model has not trained on, so a lower loss there is carried over from earlier questions. From the first ten questions to the last ten, the mean loss on a new example falls from 0.27 to 0.11, from 0.23 to 0.10 and from 0.30 to 0.13 nats per target token in runs 0, 1 and 2.

The loss on each new training example falls as the run goes on measured Training loss on each new example, measured just before the model trains on it mean over the examples written after each question. No example is trained twice, so the model has not trained on it when this loss is measured 0 0.2 0.4 0.6 0.8 1 5 10 15 20 25 30 35 39 question after which the examples were written mean loss (nats per target token) database migration run 0 0.27 → 0.11 run 1 0.23 → 0.10 run 2 0.30 → 0.13 Thin lines: mean loss of the examples written after each question. Thick lines and dots: mean over blocks of ten questions (1–10, 11–20, 21–30, 31–39). Each example is one optimizer step at batch size 1 in a single pass. The loss is the one computed for that step, before the update. Source: data/losses.json.
From the first ten questions to the last ten, the mean loss on a new example goes from 0.27 to 0.11 in run 0, from 0.23 to 0.10 in run 1, from 0.30 to 0.13 in run 2. It falls in all three runs.

Each lesson takes only a handful of examples. After a question the teacher usually writes several examples for each procedure it teaches. The first is often the step the learner already took, so its loss is low. The second is the first that teaches something new, with a median loss of 0.37 nats per target token. By the eighth example the median is 0.07, all within the training that follows that one question.

A lesson is absorbed within a handful of examples measured Loss on the k-th example of the same lesson, before training on it examples the teacher wrote for one lesson after one question. Median over all lessons in the three runs 0 0.1 0.2 0.3 0.4 1 2 3 4 5 6 7 8 9 10 example number within the lesson median loss (nats per target token) 0.07 0.37 0.24 0.22 0.17 0.23 0.12 0.07 0.04 0.06 Examples at each position, 1 to 10: 130, 126, 124, 100, 76, 55, 48, 39, 29, 24. Positions with fewer than 10 examples are not shown. Only examples whose target the teacher wrote for a state the learner actually reached. The first example of a lesson is usually the step the learner already took correctly, so its loss is near zero (grey). Each example is used once. Source: lane training events and schedules (data/lesson_th.json).
From the second example of a lesson to the tenth, the median loss falls from 0.37 to 0.06: the model has largely learned the lesson before the last of its examples arrives.

6. After the migration

At question 21 the database is migrated: the musical-instrument attribute table is renamed and replaced by a new one, musical-instrument timestamps are split into year and month columns, office reviews gain a partially filled verification label, and electronics gain a second price column. The paper that introduces the benchmark (Asawa et al., 2026, arXiv:2606.05661) illustrates the failure this is meant to catch: an agent that reuses what it learned before the migration and queries tables that no longer exist.

In the sixty post-migration questions of the three runs, 46 read the schema in their first two queries, and 0 queries used the renamed table's old name.

The three runs answered 9, 7 and 8 of the 20 questions before the migration and 12, 8 and 8 of the 20 after it. The stateless baseline answered 4 and 3. No run did worse after the database changed. By product group, office questions went from 1 to 3 of 8 correct before the migration to 3 to 5 of 6 after it, and musical instruments from 1 to 2 of 3 to 3 to 5 of 7. Electronics went the other way, from 4 to 5 of 8 to 1 to 2 of 7, against 2 of 7 for the baseline. Electronics is the group to which the migration added a second price column.

Before and after the database migration, by product group measured Share of questions answered correctly, before and after the database migration by product group. Each pair of bars is one run, before (light) and after (dark) the migration 0% 25% 50% 75% 100% correct r0 5/6 r1 4/6 r2 3/6 Office r0 5/7 r1 3/7 r2 4/7 Musical instruments r0 2/7 r1 1/7 r2 1/7 Electronics Before: questions 1–20 (Office 8, Musical instruments 3, Electronics 8 per run). After: questions 21–40 on the migrated database (Office 6, Musical instruments 7, Electronics 7). Questions that span groups are left out. Source: data/de_lanes.json, data/traces.json.
Correct answers per run before → after the migration: office 1–3 of 8 → 3–5 of 6. Musical instruments 1–2 of 3 → 3–5 of 7. Electronics 4–5 of 8 → 1–2 of 7.

7. Results

lanecorrect of 40reward (of 40)gain over baselinequeriessupervised tokens trainedvs baseline: gained / lostexact McNemar p
Run 02112.13+10.2725845,09615 / 10.001
Run 1157.27+5.4024837,57910 / 20.039
Run 21610.67+8.8018720,54712 / 30.035
mean of the three runs17.33 ± 1.8610.02+8.1623134,407
Stateless baseline71.8704790

What thinking adds without learning. The baseline uses the same thinking-when-stuck rule, so the gain above is what training adds on top of thinking. Thinking alone moves little: the baseline answers 7 questions with a reward of 1.87, against 6 and 1.40 for the same model with thinking off (the baseline of the earlier, no-thinking recipe). The baseline with thinking answers 2 questions that the no-thinking baseline misses, and misses 1 that the no-thinking baseline answers.

"Gained" counts questions a run answers correctly and the baseline gets wrong. "lost" counts the reverse. Pooled over the three runs, the learner gains 37 questions and loses 6.

Every question, every run measured Correct (filled) or wrong (open) on each of the 40 questions, by question id questions 1–20 before the migration, 101–120 after it. Each run answered them in its own order baseline 7/40 run 0 21/40 run 1 15/40 run 2 16/40 1 3 5 7 9 11 13 15 17 19 101 103 105 107 109 111 113 115 117 119 Source: data/de_lanes.json.
The baseline solves 7 questions and the three runs solve 21, 15 and 16. 12 questions are solved by at least two runs while the baseline misses them. 1 of the questions the baseline solves is missed by all three runs.
Exploratory queries per question measured Exploratory queries per question, moving average over five questions each question allows up to 15 queries before the answer 0 3 6 9 12 15 1 5 10 15 20 25 30 35 40 question (position in the run) queries per question database migration run 0 6.4 run 1 4.4 run 2 2.4 stateless baseline 11.0 Mean queries per question over all 40: 6.45, 6.20 and 4.67 for the three runs, 11.97 for the baseline. Source: data/de_lanes.json.
The baseline spends 12.0 queries per question on average. The three runs spend 6.5, 6.2 and 4.7.

8. Against the published leaderboard

Learner 1.0 is the only system on this leaderboard that learns strictly through its weights. Its context is reset before every question, so no in-context learning is allowed between questions. Every other system carries its experience forward in its context, a notepad, a memory system or a coding agent's workspace.

The benchmark publishes results for twelve systems on this task, each pairing a frontier model with a way of carrying experience forward: in context, through a notepad, through a memory system, or through a coding agent's workspace. Our three runs average a reward of 10.02 and a gain of +8.16.

Database exploration: our runs against the published leaderboard measured Reward over 40 questions (mean ± standard error) 0 5 10 15 20 25 Claude Code · Sonnet 4.6 22.1 Mem0 · GPT-5.4 17.2 ICL · Claude Opus 4.7 15.7 ICL · Gemini 3 Flash 15.0 ICL · Claude Sonnet 4.6 15.0 ICL · GPT-5.4 13.9 ICL Notepad · GPT-5.4 12.4 ICL · Gemini 3.1 Pro Preview 11.6 ICL Notepad · Claude Sonnet 4.6 11.0 Learner 1.0 (ours, 3 runs) 10.0 Codex · GPT-5.4 9.6 ICL Notepad · Gemini 3.1 Pro Preview 8.5 ACE · GPT-5.4 7.9 Normalized gain: gain ÷ (40 − own baseline) 0 0.1 0.2 0.3 0.4 0.5 Claude Code · Sonnet 4.6 0.436 Mem0 · GPT-5.4 0.362 ICL · Gemini 3 Flash 0.315 ICL · Claude Opus 4.7 0.283 ICL · Claude Sonnet 4.6 0.253 ICL · GPT-5.4 0.242 Learner 1.0 (ours, 3 runs) 0.214 ICL · Gemini 3.1 Pro Preview 0.194 ICL Notepad · GPT-5.4 0.187 Codex · GPT-5.4 0.168 ICL Notepad · Gemini 3.1 Pro Preview 0.121 ICL Notepad · Claude Sonnet 4.6 0.112 ACE · GPT-5.4 0.069 Published systems: CL-Bench leaderboard data of 2026-07-18, database_exploration task, 5 runs each (Codex · GPT-5.4: 1 run). Ours: 3 runs (orders 0, 1, 2) and one stateless baseline with the same decoding. Gain is normalized as the leaderboard does: gain over the system's own stateless baseline, divided by the headroom above that baseline (40 − baseline). Sources: data/leaderboard_data_2026-07-18.json, data/de_lanes.json.
Our three runs average a reward of 10.02 (10 of 13) and a normalized gain of 0.214 (7 of 13).
systemrunsrewardreward rankgainnormalized gaingain rank
Claude Code · Sonnet 4.6522.051+13.850.4361
Mem0 · GPT-5.4517.242+12.910.3622
ICL · Claude Opus 4.7515.653+9.590.2834
ICL · Gemini 3 Flash515.034+11.490.3153
ICL · Claude Sonnet 4.6515.015+8.480.2535
ICL · GPT-5.4513.886+8.350.2426
ICL Notepad · GPT-5.4512.377+6.370.1879
ICL · Gemini 3.1 Pro Preview511.568+6.830.1948
ICL Notepad · Claude Sonnet 4.6511.009+3.670.11212
Learner 1.0 (ours)310.0210+8.160.2147
Codex · GPT-5.419.6011+6.130.16810
ICL Notepad · Gemini 3.1 Pro Preview58.5212+4.320.12111
ACE · GPT-5.457.8513+2.390.06913

On reward we rank 10 of 13. Gain is ranked as the leaderboard ranks it, normalized by each system's headroom: gain divided by 40 minus the system's own baseline, because a system whose baseline is already high has less room to gain. On that measure we rank 7 of 13. The same convention is used for the sales task.

9. Against the earlier recipe

The earlier version of this recipe (full write-up) was identical except that the learner never thought before acting. It ran on the same three question orders. The final recipe answers 52 questions correctly over the three orders against 43, and repeats an identical query 34 times against 184. Its total reward is 30.07 against 30.00: the same.

Thinking when stuck against the earlier recipe measured correct answers (of 40) 13 21 run 0 16 15 run 1 14 16 run 2 reward (of 40) 8.9 12.1 run 0 12.0 7.3 run 1 9.1 10.7 run 2 repeated identical queries 36 12 run 0 55 10 run 1 93 12 run 2 earlier recipe final recipe: think when stuck Each pair of bars is one question order. Both recipes answered the same forty questions in that order with the same model, teacher and training settings. Reward = 1 − queries/15 for a correct answer, else 0. Source: data/de_lanes.json, data/traces.json.
Over the three orders the final recipe answers 52 questions correctly against 43, and repeats an identical query 34 times against 184. Total reward is 30.07 against 30.00: the extra correct answers took more queries each, so the benchmark score is the same.

The split by whether the learner got stuck shows why the two totals match. A question counts as stuck if at any step the stuck rule of section 2 would have fired. For the earlier recipe the rule is applied afterwards to its recorded trajectories.

questions, pooled over the three ordersfinal recipe: questionscorrectrewardearlier recipe: questionscorrectreward
the learner got stuck at some step55258.8058124.80
the learner never got stuck652721.27623125.20

Thinking happens only when the learner is stuck, and there it helps: the final recipe answers 25 of 55 such questions against 12 of 58. Those answers come late, after many queries, so each earns little reward. On questions where the learner never got stuck no thinking happened, yet the final recipe answers 27 of 65 against 31 of 62. That difference cannot come from thinking. It comes from what each run's training taught. Run 1 accounts for it: on its never-stuck questions it answers 5 of 20 against 11 of 25 for the earlier recipe on the same order, while runs 0 and 2 answer 12 of 21 and 10 of 24 against 9 of 18 and 11 of 19. Runs 0 and 1 open most later questions with a table listing before the schema read (24 and 26 of questions 12–40), where the earlier runs opened with the schema read directly (3 and 0). Run 2 opens with the schema read. The extra listing costs one query per question, but it does not explain run 1's shortfall by itself, because run 0 opens the same way and does better than before. We do not know what in run 1's training caused it.

10. The recipe

After each question. The teacher receives the learner's trajectory on the question just answered, the benchmark's released feedback for it, the learner's trajectories on earlier questions as evidence, and the procedures it wrote before. It returns procedures in a fixed format: what the question needed, the states that lead to an answer, and the exact next query or answer at each state.

From procedures to examples. The recipe builds each example from a state the learner actually reached: the conversation up to that point, with the recorded query results, and the next action as the target. It keeps only examples grounded in recorded results, adds three fixed examples about reading the database's state, and caps any single procedure at a quarter of a question's examples. Every example is trained once.

Acting. Greedy decoding. Thinking before an action only when stuck, as defined in section 2: at most 512 tokens of reasoning, then at most 384 tokens for the action. The learner's context is 32,000 tokens.

11. Interpretation, and how to check it

Training on each answered question changed how the learner explores this database, and thinking when stuck recovered questions the earlier recipe lost. The benchmark score does not show the second effect, because those answers took many queries and because one run regressed on questions where it never got stuck.

To check: every trajectory of every lane, the feedback, the training examples and the scores are in the replication repository, and the figures regenerate from them.

Database exploration · Learner Labs