Learner LabsLearner 1.0Foundation models that continually learn

Demonstrations · Continual Learning Bench

Database exploration: the earlier recipe

On the database exploration task of Continual Learning Bench, our Learner 1.0 model answers forty questions about one product database in sequence and is trained on each question after it answers it. Over three runs it answers 13, 16 and 14 questions correctly where the same model with learning off answers 6, and it uses 39% fewer queries to get there.

Model Learner 1.0

Benchmark Continual Learning Bench, database exploration, default schedule

Questions 40 per run, 20 before and 20 after a schema migration

Runs 3 question orders (0 canonical, 1 and 2 permuted) and one stateless baseline

Learning after every answered question. One pass, one example per update, no replay

What the learner sees only the current question, exactly as the benchmark presents it

Teacher a teacher LLM writes the training examples after each question

Compute one H200 per run

Learner 1.0 is the only system on this benchmark that learns strictly through its weights. Its context is reset before every question, so there is no in-context learning between questions: anything it carries from one question to the next is in its weights. Every other system on the public leaderboard carries its experience forward in its context, a notepad, a memory store or a coding agent’s workspace.

1. The question and what was measured

The question: if a model is trained on its own experience after every question it answers, does it answer later questions better than the same model that never learns? The benchmark measures this directly. Each question is a natural-language question about a SQLite database of three product datasets (office products, electronics, musical instruments). The agent may run up to 15 exploratory queries, then submits one answer, and is told whether it was right and, if not, the correct answer. Reward for a question is 1 − queries/15 if the answer is correct and 0 if it is wrong, so a correct answer found in fewer queries earns more. After question 20 the database is migrated to a new schema, and questions 21–40 are asked against it.

Two numbers follow from that. Reward is the sum over the forty questions (at most 40). Gain is reward minus the reward of the same model answering the same questions with learning off. The benchmark calls this the stateless baseline. Gain is the benchmark's learning metric.

Reward rises as the learner trains on each answered question measured Running average of reward per question, 40 questions in order reward per question = 1 − queries/15 if correct, else 0. Mean over the questions answered so far 0 0.1 0.2 0.3 0.4 0.5 1 5 10 15 20 25 30 35 40 question (position in the run) running average reward database migration run 0 0.222 run 1 0.300 run 2 0.228 stateless baseline 0.035 Three runs of the same recipe in three question orders (the benchmark default schedule: order 0 canonical, orders 1 and 2 permuted within each half). The stateless baseline is the same model with learning off, scored once per question. Questions 21–40 follow a schema migration of the database. Source: data/de_lanes.json.
All three runs end above the stateless baseline and keep rising after the database migration at question 21. Final running averages: 0.222, 0.300 and 0.228 for the three runs against 0.035 for the baseline.

The three runs end at rewards 8.87, 12.00 and 9.13 against 1.40 for the baseline. The gain over the baseline is +7.47, +10.60 and +7.73.

Gain over the stateless baseline grows through the run measured Running average gain per question gain per question = reward of the training run − reward of the stateless baseline on the same question -0.2 -0.1 0 0.1 0.2 0.3 0.4 1 5 10 15 20 25 30 35 40 question (position in the run) running average gain database migration run 0 +0.187 run 1 +0.265 run 2 +0.193 Gain is the CL-Bench learning metric: each question is paired with the same question answered by the same model with learning off, so the difficulty of the question cancels. Source: data/de_lanes.json.
Gain turns positive by question 10 on every run and stays positive to the end. The total gain over 40 questions is +7.47, +10.60 and +7.73 for the three runs.

2. Setup

Model. Learner 1.0, the same model as in the ten-skill experiment, with a fixed number of trainable parameters set before the first question. The mechanism is proprietary and is not described here.

What the learner sees. At each question the learner receives the benchmark's own system prompt and the current question, and nothing else: no earlier questions, no notes, no retrieved examples. Its prompt is the same as the baseline's. Anything it carries from one question to the next is in its weights.

The loop. The learner answers a question. The benchmark scores it and releases its feedback. A teacher LLM reads the learner's trajectory on that question and the feedback, and writes procedures: what the question needed and which queries answer it. The recipe turns each procedure into training examples built on states the learner actually reached, and the learner is trained on them before the next question. The teacher never answers a question and never acts in the benchmark. Section 11 gives the full recipe.

Training. One pass, batch size one, no replay of earlier examples, no optimizer resets between questions. Across a run the learner trained on 25,519 to 76,560 supervised tokens.

Runs. The benchmark's default schedule asks the same forty questions in a fixed canonical order (run 0) and in permuted orders that keep the twenty pre-migration questions first (runs 1 and 2). The stateless baseline is run once. Because nothing carries over between its questions, its order does not matter. All four lanes used the same model and the same decoding.

3. A few example questions compared, before and after learning

Here is the same question answered by the baseline and by run 1, on the same database. The question asks about the office products review table, which has two product identifier columns.

Each column below shows one agent's actions in order. Each numbered step is one SQL query the agent sent, with the start of the result the database returned below it. The last line is the answer it submitted and whether it was right. The full trajectories of every run, all forty questions each, are in the replication repository.

Question. In the office products review data, there are two product identifier columns. Restricting to reviews with a star rating of 2.0 or below, how many more distinct values does the non-primary identifier have compared to the foreign key to the products table?

Stateless baseline

question 6 of 40 in this run · 7 queries
  1. list.tables
    attrs_g1
    attrs_g3
    fdbk_g1
    … 9 more lines
  2. listPRAGMA table_info(fdbk_g1);
    cid | name    | type    | notnull | dflt_value | pk
    0   | id      | INTEGER | 0       | NULL       | 1
    1   | item_id | TEXT    | 0       | NULL       | 0
    … 8 more lines
  3. listPRAGMA table_info(items_g1);
    cid | name        | type    | notnull | dflt_value | pk
    0   | ref_id      | TEXT    | 0       | NULL       | 1
    1   | ttl         | TEXT    | 0       | NULL       | 0
    … 8 more lines
  4. sampleSELECT main_cat FROM items_g1 LIMIT 5;
    main_cat
    All Electronics
    Amazon Home
    … 3 more lines
  5. computeSELECT main_cat, COUNT(*) FROM items_g1 GROUP BY main_cat;
    main_cat                     | COUNT(*)
    NULL                         | 741
    AMAZON FASHION               | 224
    All Beauty                   | 25
    All Electronics              | 1308
    Amazon Fire TV               | 1
    … 28 more lines
  6. computeSELECT COUNT(DISTINCT ref_id) FROM fdbk_g1 WHERE rtg <= 2.0 AND ref_id IN (SELECT ref_id FROM items_g1 WHERE main_cat = 'Office Products');
    COUNT(DISTINCT ref_id)
    5523
  7. computeSELECT COUNT(DISTINCT item_id) FROM fdbk_g1 WHERE rtg <= 2.0 AND ref_id IN (SELECT ref_id FROM items_g1 WHERE main_cat = 'Office Products');
    COUNT(DISTINCT item_id)
    6776
answered 1253 · wrong (correct: 1832)

Learner, run 1

question 17 of 40 in this run · 2 queries
  1. read schemaSELECT name, sql FROM sqlite_master WHERE type='table'
    name            | sql
    items_g1        | CREATE TABLE items_g1 (     ref_id TEXT PRIMARY KEY,     ttl TEXT,     avg_rtg REAL,     avg_rtg_30d R…
    fdbk_g1         | CREATE TABLE fdbk_g1 (     id INTEGER PRIMARY KEY AUTOINCREMENT,     item_id TEXT,     ref_id TEXT,   …
    … 10 more lines
  2. computeSELECT COUNT(DISTINCT item_id) AS n_item_id, COUNT(DISTINCT ref_id) AS n_ref_id, COUNT(DISTINCT item_id) - COUNT(DISTINCT ref_id) AS diff FROM fdbk_g1 WHERE rtg <= 2.0;
    n_item_id | n_ref_id | diff
    10059     | 8227     | 1832
answered 1832 · correct

The baseline lists the tables, inspects two of them, then filters the reviews to products whose catalog category is "Office Products". That filter is the mistake: the office dataset is the whole _g1 table group, and the category column also holds other categories. It answers 1253. Run 1 reaches this question 17th. It reads the schema in one query and computes the difference over the whole office review table in its second query: 1832, correct, two queries. Run 2 answers the same question the same way, also in two queries.

Three more questions, each asked before the migration, so both agents faced the same database:

Electronics reviews longer than 500 characters, products at most $40

Question. Among electronics reviews with a review body longer than 500 characters, considering only products whose list price is at most $40, what percentage are 1-star? Round to 2 decimal places.

Stateless baseline

question 16 of 40 in this run · 15 queries
  1. list.tables
    attrs_g1
    attrs_g3
    fdbk_g1
    … 9 more lines
  2. listPRAGMA table_info(items_g1);
    cid | name        | type    | notnull | dflt_value | pk
    0   | ref_id      | TEXT    | 0       | NULL       | 1
    1   | ttl         | TEXT    | 0       | NULL       | 0
    … 8 more lines
  3. sampleSELECT main_cat FROM items_g1 LIMIT 5;
    main_cat
    All Electronics
    Amazon Home
    … 3 more lines
  4. sampleSELECT main_cat FROM items_g2 LIMIT 5;
    main_cat
    Computers
    Computers
    … 3 more lines
  5. sampleSELECT main_cat FROM items_g3 LIMIT 5;
    main_cat
    Musical Instruments
    Musical Instruments
    … 3 more lines
  6. listPRAGMA table_info(fdbk_g1);
    cid | name    | type    | notnull | dflt_value | pk
    0   | id      | INTEGER | 0       | NULL       | 1
    1   | item_id | TEXT    | 0       | NULL       | 0
    … 8 more lines
  7. listPRAGMA table_info(fdbk_g2);
    cid | name    | type    | notnull | dflt_value | pk
    0   | id      | INTEGER | 0       | NULL       | 1
    1   | item_id | TEXT    | 0       | NULL       | 0
    … 8 more lines
  8. listPRAGMA table_info(fdbk_g3);
    cid | name   | type    | notnull | dflt_value | pk
    0   | id     | INTEGER | 0       | NULL       | 1
    1   | ref_id | TEXT    | 0       | NULL       | 0
    … 7 more lines
  9. sampleSELECT DISTINCT main_cat FROM items_g1 WHERE main_cat LIKE '%Electronics%' LIMIT 10;
    main_cat
    All Electronics
    Car Electronics
  10. sampleSELECT DISTINCT main_cat FROM items_g2 WHERE main_cat LIKE '%Electronics%' LIMIT 10;
    main_cat
    All Electronics
    Car Electronics
  11. sampleSELECT DISTINCT main_cat FROM items_g2 LIMIT 20;
    main_cat
    Computers
    Cell Phones & Accessories
    … 18 more lines
  12. computeSELECT COUNT(*) FROM items_g2 WHERE main_cat = 'All Electronics';
    COUNT(*)
    4562
  13. computeSELECT COUNT(*) FROM items_g1 WHERE main_cat = 'All Electronics';
    COUNT(*)
    1308
  14. computeSELECT COUNT(*) FROM items_g3 WHERE main_cat = 'All Electronics';
    COUNT(*)
    534
  15. computeSELECT ROUND(100.0 * SUM(CASE WHEN f.rtg = 1 THEN 1 ELSE 0 END) / COUNT(*), 2) AS pct_1star FROM fdbk_g2 f JOIN items_g2 i ON f.ref_id = i.ref_id WHERE i.main_cat = 'All Electronics' AND i.prc <= 40 AND LENGTH(f.body) > 500;
    pct_1star
    NULL
answered 0.00 · wrong (correct: 10.22)

Learner, run 1

question 13 of 40 in this run · 2 queries
  1. read schemaSELECT name, sql FROM sqlite_master WHERE type='table'
    name            | sql
    items_g1        | CREATE TABLE items_g1 (     ref_id TEXT PRIMARY KEY,     ttl TEXT,     avg_rtg REAL,     avg_rtg_30d R…
    fdbk_g1         | CREATE TABLE fdbk_g1 (     id INTEGER PRIMARY KEY AUTOINCREMENT,     item_id TEXT,     ref_id TEXT,   …
    … 10 more lines
  2. computeSELECT ROUND(100.0 * SUM(CASE WHEN rtg = 1 THEN 1 ELSE 0 END) / COUNT(*), 2) AS pct_1_star FROM fdbk_g2 JOIN items_g2 i ON fdbk_g2.ref_id = i.ref_id WHERE LENGTH(body) > 500 AND prc_usd <= 40
    pct_1_star
    10.22
answered 10.22 · correct

The baseline spends all fifteen queries working out which category values mean electronics, then filters on the column prc, which is empty in the electronics table, gets NULL and answers 0.00. Run 1 uses the electronics table group and its prc_usd price column: 10.22 in two queries. Section 4 traces where it learned that column.

Average review rating of office products above $50

Question. What is the average review rating given to office products that cost more than $50?

Stateless baseline

question 19 of 40 in this run · 4 queries
  1. list.tables
    attrs_g1
    attrs_g3
    fdbk_g1
    … 9 more lines
  2. listPRAGMA table_info(items_g1);
    cid | name        | type    | notnull | dflt_value | pk
    0   | ref_id      | TEXT    | 0       | NULL       | 1
    1   | ttl         | TEXT    | 0       | NULL       | 0
    … 8 more lines
  3. sampleSELECT DISTINCT main_cat FROM items_g1 LIMIT 10;
    main_cat
    All Electronics
    Amazon Home
    … 8 more lines
  4. computeSELECT AVG(avg_rtg) FROM items_g1 WHERE main_cat = 'Office Products' AND prc > 50;
    AVG(avg_rtg)
    4.16020942408377
answered 4.16020942408377 · wrong (correct: 4.11)

Learner, run 2

question 18 of 40 in this run · 3 queries
  1. read schemaSELECT name, sql FROM sqlite_master WHERE type='table'
    name            | sql
    items_g1        | CREATE TABLE items_g1 (     ref_id TEXT PRIMARY KEY,     ttl TEXT,     avg_rtg REAL,     avg_rtg_30d R…
    fdbk_g1         | CREATE TABLE fdbk_g1 (     id INTEGER PRIMARY KEY AUTOINCREMENT,     item_id TEXT,     ref_id TEXT,   …
    … 10 more lines
  2. computeSELECT COUNT(*) FROM items_g1 WHERE main_cat LIKE '%Office%';
    COUNT(*)
    13933
  3. computeSELECT AVG(f.rtg) AS avg_rating FROM fdbk_g1 f JOIN items_g1 i ON f.ref_id = i.ref_id WHERE i.prc > 50;
    avg_rating
    4.105430924150588
answered 4.105430924150588 · correct

The baseline averages the catalog's per-product rating column and answers 4.16. The question asks about the ratings given in reviews. Run 2 averages the review rows and answers 4.11.

Share of 'Computers & Accessories' products priced above the category average

Question. Among electronics products in the 'Computers & Accessories' top-level category, what percentage are priced above the category-wide average? Round to 2 decimal places. Only consider products with a non-zero price.

Stateless baseline

question 8 of 40 in this run · 9 queries
  1. list.tables
    attrs_g1
    attrs_g3
    fdbk_g1
    … 9 more lines
  2. listPRAGMA table_info(items_g1)
    cid | name        | type    | notnull | dflt_value | pk
    0   | ref_id      | TEXT    | 0       | NULL       | 1
    1   | ttl         | TEXT    | 0       | NULL       | 0
    … 8 more lines
  3. sampleSELECT DISTINCT main_cat FROM items_g1 LIMIT 10
    main_cat
    All Electronics
    Amazon Home
    … 8 more lines
  4. computeSELECT DISTINCT main_cat FROM items_g1 WHERE main_cat LIKE '%Computer%'
    main_cat
    Computers
  5. sampleSELECT DISTINCT main_cat FROM items_g2 LIMIT 10
    main_cat
    Computers
    Cell Phones & Accessories
    … 8 more lines
  6. sampleSELECT DISTINCT main_cat FROM items_g3 LIMIT 10
    main_cat
    Musical Instruments
    NULL
    … 8 more lines
  7. computeSELECT main_cat, COUNT(*) as cnt FROM items_g2 GROUP BY main_cat ORDER BY cnt DESC LIMIT 10
    main_cat                  | cnt
    Computers                 | 5247
    All Electronics           | 4562
    Camera & Photo            | 2789
    Cell Phones & Accessories | 1775
    NULL                      | 1304
    … 5 more lines
  8. computeSELECT AVG(prc) as avg_price FROM items_g2 WHERE main_cat = 'Computers' AND prc > 0
    avg_price
    14488.249662618084
  9. computeSELECT COUNT(*) as total_count, SUM(CASE WHEN prc > (SELECT AVG(prc) FROM items_g2 WHERE main_cat = 'Computers' AND prc > 0) THEN 1 ELSE 0 END) as above_avg_count FROM items_g2 WHERE main_cat = 'Computers' AND prc > 0
    total_count | above_avg_count
    1482        | 298
answered 20.11 · wrong (correct: 17.27)

Learner, run 2

question 14 of 40 in this run · 2 queries
  1. read schemaSELECT name, sql FROM sqlite_master WHERE type='table'
    name            | sql
    items_g1        | CREATE TABLE items_g1 (     ref_id TEXT PRIMARY KEY,     ttl TEXT,     avg_rtg REAL,     avg_rtg_30d R…
    fdbk_g1         | CREATE TABLE fdbk_g1 (     id INTEGER PRIMARY KEY AUTOINCREMENT,     item_id TEXT,     ref_id TEXT,   …
    … 10 more lines
  2. computeSELECT ROUND(100.0 * SUM(CASE WHEN prc > avg_prc THEN 1 ELSE 0 END) / COUNT(*), 2) AS pct_above_avg FROM (SELECT i.prc, AVG(i.prc) OVER () AS avg_prc FROM items_g2 i JOIN taxn_g2 t ON i.ref_id = t.ref_id WHERE t.cat_lvl = 1 AND t.cat_nm = 'Computers & Accessories' AND i.prc > 0);
    pct_above_avg
    17.27
answered 17.27 · correct

"Top-level category" refers to the product taxonomy table, not the catalog's category column. The baseline filters on the category value "Computers" and answers 20.11. Run 2 joins the taxonomy at level 1 and answers 17.27.

4. What it learned

The figure below draws every step of every question, in the order the run answered them, for the baseline and for run 1. Each row is a question. Each cell is one query, colored by what the query does.

Every step of every question: learning off and learning on measured read the full schema in one query list tables or columns look at sample rows compute toward the answer repeat an earlier identical query answer: correct (filled) or wrong (open) Stateless baseline (learning off) 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 migration 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 6 of 40 correct · 450 queries Learner, run 1 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 migration 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 16 of 40 correct · 223 queries Each row is one question, in the order the run answered it (row labels are positions 1–40). Each cell is one action, left to right. Up to 15 queries are allowed before the answer. The dashed line marks the schema migration after question 20. Classification of each query is by its text (scripts/step_classes.py). Source: data/traces.json.
The baseline opens every question by listing tables and sampling rows, and uses 450 queries over the run. Run 1 does the same for its first eleven questions, then opens every question with one full-schema read and answers most questions in two or three queries, 223 in total. It also falls into repeating an identical query on some questions. The baseline almost never does.

The baseline's rows look the same from question 1 to question 40: list the tables, look at sample rows, compute, answer. Run 1's rows look like the baseline's for its first eleven questions. From question 12 on, every row starts with one query that reads the full schema, and most rows end after two or three queries. The same switch happens in the other two runs, at question 6 in run 0 and question 9 in run 2:

The same change in the other two runs measured read the full schema in one query list tables or columns look at sample rows compute toward the answer repeat an earlier identical query answer: correct (filled) or wrong (open) Learner, run 0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 migration 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 13 of 40 correct · 306 queries Learner, run 2 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 migration 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 14 of 40 correct · 299 queries Each row is one question, in the order the run answered it (row labels are positions 1–40). Each cell is one action, left to right. Up to 15 queries are allowed before the answer. The dashed line marks the schema migration after question 20. Classification of each query is by its text (scripts/step_classes.py). Source: data/traces.json.
Run 0 switches to the full-schema opening at question 6 and run 2 at question 9. From question 21 both open every question that way. Run 2 repeats identical queries on several late questions, which is where its second-half score drops.

Counted per block of ten questions. A sample-row query looks at a few raw rows of a table (for example SELECT … LIMIT 5) instead of computing an answer. A full-schema read asks the database for all its table definitions at once. Touching a category column means filtering or grouping on the catalog's category column, the mistake shown in section 3.

per block of ten questions (positions 1–10, 11–20, 21–30, 31–40)baselinerun 0run 1run 2
questions opened with one full-schema read0 · 0 · 0 · 04 · 7 · 10 · 100 · 9 · 10 · 101 · 8 · 10 · 10
sample-row queries35 · 37 · 34 · 4830 · 4 · 5 · 024 · 1 · 0 · 022 · 0 · 0 · 1
questions that touch a category column9 · 9 · 10 · 99 · 6 · 9 · 310 · 4 · 3 · 38 · 4 · 2 · 6
queries in total95 · 108 · 127 · 120108 · 88 · 80 · 3090 · 54 · 42 · 37100 · 53 · 70 · 76
repeats of an identical earlier query0 · 3 · 0 · 021 · 11 · 4 · 019 · 17 · 11 · 83 · 20 · 28 · 42

Runs 0 and 1 need fewer queries with each block of ten questions: 108 down to 30 and 90 down to 37. The baseline never improves, from 95 to 127. Run 2 falls to 53 and then rises again, to 70 and 76, where its repeated queries start (section 9). In every run, sample-row queries drop from 22 to 30 in the first block to at most 5 in each later block.

The full-schema opening is one of the things the recipe teaches directly: every boundary adds a training example whose target is that query. It is trained from the first question on and the learner adopts it after five to eleven questions. The rest of what changes is specific to this database.

A lesson and where it was used

Run 1 answered the 'Computers & Accessories' question ninth and got it wrong: it used the office table group and the column prc, and answered 17.69 against the correct 17.27.

Learner, run 1, question 9

question 9 of 40 in this run · 5 queries
  1. list.tables
    attrs_g1
    attrs_g3
    fdbk_g1
    … 9 more lines
  2. listPRAGMA table_info(items_g1)
    cid | name        | type    | notnull | dflt_value | pk
    0   | ref_id      | TEXT    | 0       | NULL       | 1
    1   | ttl         | TEXT    | 0       | NULL       | 0
    … 8 more lines
  3. sampleSELECT DISTINCT main_cat FROM items_g1 LIMIT 10
    main_cat
    All Electronics
    Amazon Home
    … 8 more lines
  4. computeSELECT DISTINCT main_cat FROM items_g1 WHERE main_cat LIKE '%Computer%'
    main_cat
    Computers
  5. computeSELECT ROUND(100.0 * SUM(CASE WHEN prc > avg_prc THEN 1 ELSE 0 END) / COUNT(*), 2) AS pct_above_avg FROM items_g1, (SELECT AVG(prc) AS avg_prc FROM items_g1 WHERE main_cat = 'Computers' AND prc > 0) WHERE main_cat = 'Computers' AND prc > 0
    pct_above_avg
    17.69
answered 17.69 · wrong (correct: 17.27)

After that question the teacher wrote this procedure (verbatim):

Electronics-dataset category questions that name a taxonomy category ('Computers & Accessories') use the g2 table group: items_g2 joined to taxn_g2 by ref_id on cat_nm. They do not use items_g1.main_cat='Computers', which belongs to the office dataset and was rejected (17.69 vs released 17.27). items_g2 has two price columns (prc, prc_usd), so one bounded query computes the share of non-zero-priced products above the category average for both price columns and also reports their storage types.

The recipe built 7 training examples from it at that boundary. Each one is a state the learner actually reached in its question-9 trajectory, followed by the next action the procedure prescribes. One of them, as trained (the system prompt with the action schema is omitted here):

[user]
Question 9/40

Among electronics products in the 'Computers & Accessories' top-level category, what percentage are priced above the category-wide average? Round to 2 decimal places. Only consider products with a non-zero price.

You have access to a SQLite database that contains three separate product datasets covering different categories: office products, electronics, and musical instruments. Each dataset has its own set of tables. You will need to explore the schema to determine which tables belong to which product category.

You may run up to 15 exploratory SQL queries to examine the schema and data before submitting your final answer. Note: query results are capped at 50 rows — if a que … [401 more characters]

[assistant]
{"action":"QUERY","content":".tables"}

[user]
FEEDBACK: Query result (1/15 queries used, 14 remaining):

attrs_g1
attrs_g3
fdbk_g1
fdbk_g2
fdbk_g3
fdbk_stats_g1
items_g1
items_g2
items_g3
sqlite_sequence
taxn_g1
taxn_g2

[user]
Question 9/40

Among electronics products in the 'Computers & Accessories' top-level category, what percentage are priced above the category-wide average? Round to 2 decimal places. Only consider products with a non-zero price.

You have access to a SQLite database that contains three separate product datasets covering different categories: office products, electronics, and musical instruments. Each dataset has its own set of tables. You will need to explore the schema to determine which tables belong to which product category.

You may run up to 15 exploratory SQL queries to examine the schema and data before submitting your final answer. Note: query results are capped at 50 rows — if a que … [401 more characters]

[target, trained]
{"action": "QUERY", "content": "WITH pop AS (SELECT DISTINCT i.ref_id, i.prc, i.prc_usd FROM items_g2 i JOIN taxn_g2 t ON t.ref_id = i.ref_id WHERE t.cat_nm = 'Computers & Accessories') SELECT COUNT(*) AS n_items, MAX(typeof(prc)) AS prc_type, MAX(typeof(prc_usd)) AS prc_usd_type, SUM(prc > 0) AS n_prc_pos, ROUND(100.0 * SUM(prc > (SELECT AVG(prc) FROM pop WHERE prc > 0)) / SUM(prc > 0), 2) AS pct_above_avg_prc, SUM(prc_usd > 0) AS n_prc_usd_pos, ROUND(100.0 * SUM(prc_usd > (SELECT AVG(prc_usd) FROM pop WHERE prc_usd > 0)) / SUM(prc_usd > 0), 2) AS pct_above_avg_prc_usd FROM pop;"}

Four questions later run 1 was asked a different question about electronics, the one in section 3 about reviews longer than 500 characters. Its second query used the electronics table group and prc_usd, and the answer was correct. Before question 9, run 1 had not written prc_usd in any query.

5. Learner 1.0 learns efficiently online from very few examples

A run trains once after each question, on examples built from that question alone. Over forty questions run 1 trained on 256 examples and 25,519 supervised tokens, about 654 tokens per question. Runs 0 and 2 used 76,560 and 43,456. The best run used the least.

Sixteen correct answers from 25,519 supervised tokens measured Correct answers against supervised tokens trained so far supervised tokens = target tokens of the training examples. Prompt tokens are not counted 0 4 8 12 16 0 20 40 60 80 supervised tokens trained before the question (thousands) cumulative correct answers run 0 13 correct, 76.6k tokens run 1 16 correct, 25.5k tokens run 2 14 correct, 43.5k tokens Each line follows one run through its 40 questions. Dots mark correct answers. Training happens once after each question, on examples built from that question. Nothing is replayed. Source: data/losses.json, data/de_lanes.json.
The whole run trains on 25,519 to 76,560 supervised tokens across 256 to 370 examples. Run 1, the best run, used the fewest.

Individual lessons are learned from a handful of examples. The table lists names the learner did not write in any query until a training example used them. Each name was already visible in the schema listings the learner had read. What the examples taught was to use it.

runnamewhat it istaught after questionexamplessupervised tokenslater questions using itcorrect
run 1product_attributes_g3musical-instrument attribute table added by the migration22561054
run 0product_attributes_g3musical-instrument attribute table added by the migration2452,22842
run 2verified_statusoffice review verification label added by the migration31128011
run 1prc_usdelectronics price in dollars951,10571
run 1attrs_g1office product attribute table1451,22022
run 1cat_lvltaxonomy level column771,83441
run 2item_idsecond product identifier in office reviews2360321

The migration's new attribute table is the clearest case. Run 1 read it in the schema listing at questions 21 and 22 and did not query it. After question 22 it trained on five examples that did, 610 supervised tokens in total, and from question 25 on it used the table in five questions and answered four of them correctly. Run 2 learned the migration's new office verification label from a single example of 280 tokens and used it correctly two questions later.

The training loss shows the same thing. Each training example is used for one update and never again. Its loss is recorded in that same step, before the update, so every point below is the model's error on an example it has not trained on yet. A falling line means the model gets better at examples it has never trained on: it is learning how to act on this database, not memorizing the examples it was shown.

Each new lesson surprises the model less as the run goes on measured Training loss on each new example, measured before the model trains on it mean over the examples written after each question. No example is trained twice, so the model has not trained on it when this loss is measured 0 0.2 0.4 0.6 0.8 1 5 10 15 20 25 30 35 39 question after which the examples were written mean loss (nats per target token) database migration run 0 0.29 → 0.12 run 1 0.37 → 0.11 run 2 0.26 → 0.14 Thin lines: mean loss of the examples written after each question. Thick lines and dots: mean over blocks of ten questions (1–10, 11–20, 21–30, 31–39). Each example is one optimizer step at batch size 1 in a single pass. The loss is the one computed for that step, before the update. There is no separate validation set. Source: data/losses.json.
In all three runs the average loss on a new example falls between the first ten questions and the last ten, including after the database migration: run 0 by 59%, run 1 by 71%, run 2 by 46%.
A lesson is absorbed within a handful of examples measured Loss on the k-th example of the same lesson, before training on it examples the teacher wrote for one lesson after one question. Median over all lessons in the three runs 0 0.1 0.2 0.3 0.4 1 2 3 4 5 6 7 8 9 10 example number within the lesson median loss (nats per target token) 0.00 0.32 0.23 0.21 0.16 0.17 0.08 0.06 0.05 0.06 Examples at each position, 1 to 10: 134, 131, 125, 96, 81, 62, 47, 40, 21, 11. Only examples whose target the teacher wrote for a state the learner actually reached. The first example of a lesson is usually the step the learner already took correctly, so its loss is near zero (grey). Each example is used once. Source: data/losses.json, lane schedules.
From the second example of a lesson to the eighth, the median loss falls from 0.32 to 0.06: the model has largely learned the lesson before the last of its examples arrives.

6. After the migration

At question 21 the database is migrated: the musical-instrument attribute table is renamed and replaced by a new one, musical-instrument timestamps are split into year and month columns, office reviews gain a partially filled verification label, and electronics gain a second price column. The paper that introduces the benchmark (Asawa et al., 2026, arXiv:2606.05661) illustrates the failure this is meant to catch: an agent that reuses what it learned before the migration and queries tables that no longer exist.

None of the three runs did that. Every one of their sixty post-migration questions opened by reading the full schema, and no run sent a single query to the renamed table under its old name. The full-schema opening, learned in the first half, is what exposes the new schema.

After the migration, two of three product groups recover measured Share of questions answered correctly, before and after the database migration by product group. Each pair of bars is one run, before (light) and after (dark) the migration 0% 25% 50% 75% 100% correct r0 4/6 r1 4/6 r2 4/6 Office r0 2/7 r1 5/7 r2 2/7 Musical instruments r0 1/7 r1 0/7 r2 0/7 Electronics Before: questions 1–20 (Office 8, Musical instruments 3, Electronics 8 per run). After: questions 21–40 on the migrated database (Office 6, Musical instruments 7, Electronics 7). One pre-migration question spans groups and is left out. Source: data/de_lanes.json, data/traces.json.
Office questions go from 3–4 of 8 correct before the migration to 4 of 6 after it in every run, and musical instruments to 2–5 of 7. Electronics drops to 0–1 of 7, where the migration added a second price column that no run resolved.

Two of the three product groups recovered. Run 1 shows how. On question 22 it looked for brand in the product table's brand column, which is empty for every instrument, and failed:

Learner, run 1, question 22

question 22 of 40 in this run · 13 queries
  1. read schemaSELECT name, sql FROM sqlite_master WHERE type='table'
    name                  | sql
    items_g1              | CREATE TABLE items_g1 (     ref_id TEXT PRIMARY KEY,     ttl TEXT,     avg_rtg REAL,     avg_rtg…
    fdbk_g1               | CREATE TABLE fdbk_g1 (     id INTEGER PRIMARY KEY AUTOINCREMENT,     item_id TEXT,     ref_id TE…
    … 11 more lines
  2. computeWITH n AS (SELECT COUNT(DISTINCT main_cat) AS n_cats FROM items_g1 UNION ALL SELECT COUNT(DISTINCT main_cat) AS n_cats FROM items_g2 UNION ALL SELECT COUNT(DISTINCT main_cat) AS n_cats FROM items_g3), r AS (SELECT 'g1' AS grp, main_cat, COUNT(DISTINCT ref_id) AS n_items FROM items_g1 GROUP BY main_cat UNION ALL SELECT 'g2' AS grp, main_cat, COUNT(DISTINCT ref_id) AS n_items FROM items_g2 GROUP BY main_cat UNION ALL SELECT 'g3' AS grp, main_cat, COUNT(DISTINCT ref_id) AS n_items FROM items_g3 GROUP BY main_cat) SELECT r.* FROM r JOIN n ON r.grp = CASE WHEN n.n_cats = r.n_items THEN 'g1' ELSE 'g2' END LIMIT 10;
    grp | main_cat              | n_items
    g2  | NULL                  | 1304
    g2  | AMAZON FASHION        | 232
    g2  | All Beauty            | 25
    g2  | All Electronics       | 4562
    g2  | Amazon Devices        | 106
    … 5 more lines
  3. computeWITH n AS (SELECT COUNT(DISTINCT main_cat) AS n_cats FROM items_g1 UNION ALL SELECT COUNT(DISTINCT main_cat) AS n_cats FROM items_g2 UNION ALL SELECT COUNT(DISTINCT main_cat) AS n_cats FROM items_g3), r AS (SELECT 'g1' AS grp, main_cat, COUNT(DISTINCT ref_id) AS n_items FROM items_g1 GROUP BY main_cat UNION ALL SELECT 'g2' AS grp, main_cat, COUNT(DISTINCT ref_id) AS n_items FROM items_g2 GROUP BY main_cat UNION ALL SELECT 'g3' AS grp, main_cat, COUNT(DISTINCT ref_id) AS n_items FROM items_g3 GROUP BY main_cat) SELECT r.* FROM r JOIN n ON r.grp = CASE WHEN n.n_cats = r.n_items THEN 'g1' ELSE 'g2' END WHERE r.main_cat LIKE '%Music%' OR r.main_cat LIKE '%Instrument%' LIMIT 20;
    grp | main_cat            | n_items
    g2  | Musical Instruments | 85
    g2  | Musical Instruments | 85
    g2  | Musical Instruments | 85
  4. computeWITH brands AS (SELECT str_nm AS brand, COUNT(DISTINCT ref_id) AS n_products, AVG(prc_usd) AS avg_price FROM items_g2 WHERE main_cat = 'Musical Instruments' GROUP BY str_nm HAVING n_products >= 50) SELECT brand, n_products, avg_price FROM brands ORDER BY avg_price DESC, brand ASC LIMIT 10;
    (no results)
  5. computeWITH BrandStats AS (SELECT str_nm AS brand, COUNT(DISTINCT ref_id) AS n_products, AVG(prc) AS avg_price FROM items_g3 GROUP BY str_nm HAVING n_products >= 50) SELECT brand, n_products, avg_price FROM BrandStats ORDER BY avg_price DESC, brand ASC LIMIT 10;
    brand | n_products | avg_price
    NULL  | 20000      | 123.38623097244937
  6. computeWITH BrandStats AS (SELECT str_nm AS brand, COUNT(DISTINCT ref_id) AS n_products, AVG(prc) AS avg_price FROM items_g3 WHERE str_nm IS NOT NULL GROUP BY str_nm HAVING n_products >= 50) SELECT brand, n_products, avg_price FROM BrandStats ORDER BY avg_price DESC, brand ASC LIMIT 10;
    (no results)
  7. computeWITH BrandStats AS (SELECT str_nm AS brand, COUNT(DISTINCT ref_id) AS n_products, AVG(prc) AS avg_price FROM items_g1 GROUP BY str_nm HAVING n_products >= 50) SELECT brand, n_products, avg_price FROM BrandStats ORDER BY avg_price DESC, brand ASC LIMIT 10;
    brand          | n_products | avg_price
    HP             | 104        | 374.79033898305084
    Canon          | 66         | 227.96085714285715
    Amazon Renewed | 66         | 187.9235294117647
    Brother        | 66         | 137.94026315789472
    Epson          | 65         | 123.12
    … 5 more lines
  8. computeWITH BrandStats AS (SELECT str_nm AS brand, COUNT(DISTINCT ref_id) AS n_products, AVG(prc) AS avg_price FROM items_g1 WHERE main_cat = 'Musical Instruments' GROUP BY str_nm HAVING n_products >= 50) SELECT brand, n_products, avg_price FROM BrandStats ORDER BY avg_price DESC, brand ASC LIMIT 10;
    (no results)
  9. computeWITH BrandStats AS (SELECT str_nm AS brand, COUNT(DISTINCT ref_id) AS n_products, AVG(prc_usd) AS avg_price FROM items_g2 WHERE main_cat = 'Musical Instruments' AND str_nm IS NOT NULL GROUP BY str_nm HAVING n_products >= 50) SELECT brand, n_products, avg_price FROM BrandStats ORDER BY avg_price DESC, brand ASC LIMIT 10;
    (no results)
  10. repeat ×4WITH BrandStats AS (SELECT str_nm AS brand, COUNT(DISTINCT ref_id) AS n_products, AVG(prc) AS avg_price FROM items_g3 WHERE str_nm IS NOT NULL GROUP BY str_nm HAVING n_products >= 50) SELECT brand, n_products, avg_price FROM BrandStats ORDER BY avg_price DESC, brand ASC LIMIT 10;
    the same query, sent 4 times in a row (steps 10–13)
answered Yamaha · wrong (correct: Fender)

The lesson written after that question said where brand now lives: the new attribute table. Three questions later run 1 used it in its second query and answered correctly:

Learner, run 1, question 25

question 25 of 40 in this run · 3 queries
  1. read schemaSELECT name, sql FROM sqlite_master WHERE type='table'
    name                  | sql
    items_g1              | CREATE TABLE items_g1 (     ref_id TEXT PRIMARY KEY,     ttl TEXT,     avg_rtg REAL,     avg_rtg…
    fdbk_g1               | CREATE TABLE fdbk_g1 (     id INTEGER PRIMARY KEY AUTOINCREMENT,     item_id TEXT,     ref_id TE…
    … 11 more lines
  2. computeSELECT COUNT(DISTINCT F.ref_id) AS n_products FROM product_attributes_g3 F LEFT JOIN attrs_g3_legacy L ON F.ref_id = L.ref_id AND L.attr_key = 'Brand' WHERE F.attr_key = 'Brand' AND F.attr_val IS NOT NULL AND F.attr_val != '' AND L.ref_id IS NULL
    ERROR: query exceeded 10s timeout
  3. computeWITH F AS (SELECT ref_id FROM product_attributes_g3 WHERE attr_key = 'Brand' AND attr_val IS NOT NULL AND attr_val != ''), L AS (SELECT DISTINCT ref_id FROM attrs_g3_legacy WHERE attr_key = 'Brand') SELECT (SELECT COUNT(*) FROM F) - (SELECT COUNT(*) FROM F WHERE ref_id IN (SELECT ref_id FROM L)) AS n_products
    n_products
    1145
answered 1145 · correct

Electronics did not recover. The migration added a new price column to the electronics table, next to the dollar column the learner had learned. The teacher's lessons after questions 27, 31, 34 and 35 call the competing price columns unresolved, and no run answered more than one post-migration electronics question correctly.

7. Results

lanecorrect of 40reward (of 40)gain over baselinequeriessupervised tokens trainedvs baseline: gained / lostexact McNemar p
Run 0138.87+7.4730676,56010 / 30.092
Run 11612.00+10.6022325,51912 / 20.013
Run 2149.13+7.7329943,45611 / 30.057
mean of the three runs14.33 ± 0.8810.00+8.6027648,512
Stateless baseline61.4004500

Each run is compared with the baseline question by question. "Gained" counts questions the run answers correctly and the baseline gets wrong. "lost" counts the reverse. Pooled over the three runs, the learner gains 33 questions and loses 8. Every question the baseline answers correctly is also answered correctly by at least one run.

Every question, every run measured Correct (filled) or wrong (open) on each of the 40 questions, by question id questions 1–20 before the migration, 101–120 after it. Each run answered them in its own order baseline 6/40 run 0 13/40 run 1 16/40 run 2 14/40 1 3 5 7 9 11 13 15 17 19 101 103 105 107 109 111 113 115 117 119 Source: data/de_lanes.json.
The baseline solves 6 questions and the three runs solve 13, 16 and 14. Nine questions are solved by at least two runs while the baseline misses them. Every question the baseline solves is also solved by at least one run.
The learner needs fewer exploratory queries as it trains measured Exploratory queries per question, moving average over five questions each question allows up to 15 queries before the answer 0 3 6 9 12 15 1 5 10 15 20 25 30 35 40 question (position in the run) queries per question database migration run 0 4.0 run 1 5.4 run 2 7.8 stateless baseline 11.4 Mean queries per question over all 40: 7.65, 5.58 and 7.48 for the three runs, 11.25 for the baseline. Source: data/de_lanes.json.
The baseline spends about 11 queries per question throughout. Runs 0 and 1 fall from 9.8 and 7.2 queries per question in the first half to 5.5 and 4.0 in the second. Run 2 stays near 7.5.
Training examples written after each question measured Cumulative supervised tokens trained, per run one boundary after each answered question 1–39. One pass, one example per update, no replay 0 20,000 40,000 60,000 80,000 1 5 10 15 20 25 30 35 39 boundary (after question n) supervised tokens run 0 76,560 total run 1 25,519 total run 2 43,456 total Supervised tokens are the answer tokens the loss is computed on. Source: data/de_lanes.json (training).
Each run trains on a few hundred to a few thousand supervised tokens after each question, 25,519 to 76,560 over the run. The run that trained least scored highest.

8. Against the published leaderboard

Learner 1.0 is the only system on this leaderboard that learns strictly through its weights. Its context is reset before every question, so no in-context learning is allowed between questions. Every other system carries its experience forward in its context, a notepad, a memory system or a coding agent's workspace.

The benchmark publishes results for twelve systems on this task, each pairing a frontier model with a way of carrying experience forward: in context, through a notepad, through a memory system, or through a coding agent's workspace. Our three runs average a reward of 10.00 and a gain of +8.60.

Database exploration: our runs against the published leaderboard measured Reward over 40 questions (mean ± standard error) 0 5 10 15 20 25 Claude Code · Sonnet 4.6 22.1 Mem0 · GPT-5.4 17.2 ICL · Claude Opus 4.7 15.7 ICL · Gemini 3 Flash 15.0 ICL · Claude Sonnet 4.6 15.0 ICL · GPT-5.4 13.9 ICL Notepad · GPT-5.4 12.4 ICL · Gemini 3.1 Pro Preview 11.6 ICL Notepad · Claude Sonnet 4.6 11.0 Learner 1.0 (ours, 3 runs) 10.0 Codex · GPT-5.4 9.6 ICL Notepad · Gemini 3.1 Pro Preview 8.5 ACE · GPT-5.4 7.9 Normalized gain: gain ÷ (40 − own baseline) 0 0.1 0.2 0.3 0.4 0.5 Claude Code · Sonnet 4.6 0.436 Mem0 · GPT-5.4 0.362 ICL · Gemini 3 Flash 0.315 ICL · Claude Opus 4.7 0.283 ICL · Claude Sonnet 4.6 0.253 ICL · GPT-5.4 0.242 Learner 1.0 (ours, 3 runs) 0.223 ICL · Gemini 3.1 Pro Preview 0.194 ICL Notepad · GPT-5.4 0.187 Codex · GPT-5.4 0.168 ICL Notepad · Gemini 3.1 Pro Preview 0.121 ICL Notepad · Claude Sonnet 4.6 0.112 ACE · GPT-5.4 0.069 Published systems: CL-Bench leaderboard data of 2026-07-18, database_exploration task, 5 runs each (Codex · GPT-5.4: 1 run). Ours: 3 runs (orders 0, 1, 2) and one stateless baseline. Gain is normalized as the leaderboard does: gain over the system's own stateless baseline divided by the headroom above it (40 − baseline). Sources: data/leaderboard_data_2026-07-18.json, data/de_lanes.json.
Our three runs average a reward of 10.00 (10 of 13) and a normalized gain of 0.223 (7 of 13).
systemrunsrewardreward rankgainnormalized gaingain rank
Claude Code · Sonnet 4.6522.051+13.850.4361
Mem0 · GPT-5.4517.242+12.910.3622
ICL · Claude Opus 4.7515.653+9.590.2834
ICL · Gemini 3 Flash515.034+11.490.3153
ICL · Claude Sonnet 4.6515.015+8.480.2535
ICL · GPT-5.4513.886+8.350.2426
ICL Notepad · GPT-5.4512.377+6.370.1879
ICL · Gemini 3.1 Pro Preview511.568+6.830.1948
ICL Notepad · Claude Sonnet 4.6511.009+3.670.11212
Learner 1.0 (ours)310.0010+8.600.2237
Codex · GPT-5.419.6011+6.130.16810
ICL Notepad · Gemini 3.1 Pro Preview58.5212+4.320.12111
ACE · GPT-5.457.8513+2.390.06913

On reward we rank 10 of 13. Gain is ranked as the leaderboard ranks it, normalized by each system's headroom (gain divided by 40 minus the system's own baseline). On that measure we rank 7 of 13. Our baseline is weaker than the frontier baselines (1.40 against 5.53 for ICL · GPT-5.4), so it has more headroom, which the normalization takes into account.

9. Where it fails

The learner has a failure mode the baseline does not: it sends a query it already sent in the same question. The database is read-only, so the second copy returns the same result and spends one of the fifteen queries for nothing. The baseline repeats an identical query 3 times in 450 queries. The three runs repeat 36, 55 and 93 times. Once the learner has repeated, its next query is again a repeat 78 to 98 percent of the time, and questions with a repeat score 0.01 to 0.08 on average, against 0.27 to 0.36 for the rest. Here is run 1 on its sixth question:

Learner, run 1, question 6

question 6 of 40 in this run · 15 queries
  1. listSELECT name FROM sqlite_master WHERE type='table';
    name
    items_g1
    fdbk_g1
    … 10 more lines
  2. sampleSELECT * FROM items_g1 LIMIT 5;
    ref_id     | ttl                                                                                                        …
    B001DZWCDO | LD Products Compatible Toner Cartridge Replacement for Samsung ML-2850 Series ML-D2850B High Yield (Black) …
    B096Z4NJLV | Hodiczk Small Desk Shelf Organizer White with Drawer Cute Desktop Organizer Shelf Storage Plastic Kawai Des…
    … 3 more lines
  3. computeSELECT COUNT(*) FROM fdbk_g1 WHERE main_cat = 'Office Products' AND prc > 20 AND prc < 200 AND strftime('%Y', dt) = '2022';
    ERROR: no such column: main_cat
  4. computeSELECT COUNT(*) FROM fdbk_g1 WHERE ref_id IN (SELECT ref_id FROM items_g1 WHERE main_cat = 'Office Products' AND prc > 20 AND prc < 200) AND dt >= '2022-01-01' AND dt < '2023-01-01';
    ERROR: no such column: dt
  5. repeat ×11SELECT COUNT(*) FROM fdbk_g1 WHERE ref_id IN (SELECT ref_id FROM items_g1 WHERE main_cat = 'Office Products' AND prc > 20 AND prc < 200) AND dt >= '2022-01-01' AND dt < '2023-01-01';
    the same query, sent 11 times in a row (steps 5–15)
answered 0 · wrong (correct: 13766)

Run 2's second half has the most repeats, and its score drops there: 8 correct in the first twenty questions and 6 in the second, while runs 0 and 1 improve from the first half to the second.

What causes it. A repeat usually starts right after a result that told the learner nothing new: an error, an empty table, or a count of zero. At those moments the untrained model never repeats (0 of 25), while the trained runs repeat at 30 to 67 percent of them (after an SQL error: 3 of 20, 12 of 19 and 22 of 33 in runs 0, 1 and 2). Replaying the exact recorded conversations, byte for byte, confirms that the trained weights make the choice: at the 30 moments where a trained run began a repeat, the trained checkpoint repeats all 30 times, and the untrained model, given the same conversation, sends a new query 26 times. The repeats are not memorized lessons either: repeated queries resemble the trained targets less than the learner's other queries do.

The cause is what the training leaves out. This recipe is heavy on knowledge and light on state: it teaches what to query for a kind of question, and almost never what to do after a query fails or returns nothing. Over a whole run it trains 22 to 25 examples about the database's state, none of them after an empty or zero result. An earlier controlled study on a synthetic 40-episode panel shows why that matters: a knowledge-only recipe produced 335 repeated failing queries, adding 1,024 more knowledge examples still left 203, and adding instead 704 examples of the next step at failure states left none.

10. The recipe

After each question. The teacher receives the learner's trajectory on the question just answered, the benchmark's released feedback for it, the learner's trajectories on earlier questions as evidence, and the procedures it wrote before. It returns procedures in a fixed JSON format: what the question needed, the states that lead to an answer, and the exact next query or answer at each state. Its full prompt is published in the replication repository (recipe/teacher_prompt.txt).

From procedures to examples. The recipe builds each example from a state the learner actually reached: the conversation up to that point, with the recorded query results, and the next action as the target. It keeps only examples grounded in recorded results, adds three fixed examples about reading the database's state (the full-schema opening among them), and caps any single procedure at a quarter of a boundary's examples. Every example is trained once.

Training and limits. One pass, batch one, no replay. The learner's context is 32,000 tokens and its answers are limited to 4,096 tokens. The teacher call is limited to 15 minutes.

11. Interpretation, and how to check it

Training on each answered question changed how the learner explores this database: it stops sampling tables, reads the schema once, and goes straight to the tables and columns earlier questions taught it. That produced more correct answers in fewer queries on all three runs. The size of the effect varies by run (13 to 16 correct), and the learner picks up a repetition failure the baseline does not have.

The baseline answered questions 21–40 on the pre-migration database, as the benchmark's baseline protocol does. The section-3 examples are all pre-migration questions for that reason.

To check: every trajectory of every lane, the feedback, the training examples and the scores are in the replication repository, and the figures regenerate from them.