Demonstrations · Continual Learning Bench
Database exploration: the earlier recipe
On the database exploration task of Continual Learning Bench, our Learner 1.0 model answers forty questions about one product database in sequence and is trained on each question after it answers it. Over three runs it answers 13, 16 and 14 questions correctly where the same model with learning off answers 6, and it uses 39% fewer queries to get there.
Model Learner 1.0
Benchmark Continual Learning Bench, database exploration, default schedule
Questions 40 per run, 20 before and 20 after a schema migration
Runs 3 question orders (0 canonical, 1 and 2 permuted) and one stateless baseline
Learning after every answered question. One pass, one example per update, no replay
What the learner sees only the current question, exactly as the benchmark presents it
Teacher a teacher LLM writes the training examples after each question
Compute one H200 per run
Learner 1.0 is the only system on this benchmark that learns strictly through its weights. Its context is reset before every question, so there is no in-context learning between questions: anything it carries from one question to the next is in its weights. Every other system on the public leaderboard carries its experience forward in its context, a notepad, a memory store or a coding agent’s workspace.
1. The question and what was measured
The question: if a model is trained on its own experience after every question it answers, does it answer later questions better than the same model that never learns? The benchmark measures this directly. Each question is a natural-language question about a SQLite database of three product datasets (office products, electronics, musical instruments). The agent may run up to 15 exploratory queries, then submits one answer, and is told whether it was right and, if not, the correct answer. Reward for a question is 1 − queries/15 if the answer is correct and 0 if it is wrong, so a correct answer found in fewer queries earns more. After question 20 the database is migrated to a new schema, and questions 21–40 are asked against it.
Two numbers follow from that. Reward is the sum over the forty questions (at most 40). Gain is reward minus the reward of the same model answering the same questions with learning off. The benchmark calls this the stateless baseline. Gain is the benchmark's learning metric.
The three runs end at rewards 8.87, 12.00 and 9.13 against 1.40 for the baseline. The gain over the baseline is +7.47, +10.60 and +7.73.
2. Setup
Model. Learner 1.0, the same model as in the ten-skill experiment, with a fixed number of trainable parameters set before the first question. The mechanism is proprietary and is not described here.
What the learner sees. At each question the learner receives the benchmark's own system prompt and the current question, and nothing else: no earlier questions, no notes, no retrieved examples. Its prompt is the same as the baseline's. Anything it carries from one question to the next is in its weights.
The loop. The learner answers a question. The benchmark scores it and releases its feedback. A teacher LLM reads the learner's trajectory on that question and the feedback, and writes procedures: what the question needed and which queries answer it. The recipe turns each procedure into training examples built on states the learner actually reached, and the learner is trained on them before the next question. The teacher never answers a question and never acts in the benchmark. Section 11 gives the full recipe.
Training. One pass, batch size one, no replay of earlier examples, no optimizer resets between questions. Across a run the learner trained on 25,519 to 76,560 supervised tokens.
Runs. The benchmark's default schedule asks the same forty questions in a fixed canonical order (run 0) and in permuted orders that keep the twenty pre-migration questions first (runs 1 and 2). The stateless baseline is run once. Because nothing carries over between its questions, its order does not matter. All four lanes used the same model and the same decoding.
3. A few example questions compared, before and after learning
Here is the same question answered by the baseline and by run 1, on the same database. The question asks about the office products review table, which has two product identifier columns.
Each column below shows one agent's actions in order. Each numbered step is one SQL query the agent sent, with the start of the result the database returned below it. The last line is the answer it submitted and whether it was right. The full trajectories of every run, all forty questions each, are in the replication repository.
Question. In the office products review data, there are two product identifier columns. Restricting to reviews with a star rating of 2.0 or below, how many more distinct values does the non-primary identifier have compared to the foreign key to the products table?
Stateless baseline
- list
.tablesattrs_g1 attrs_g3 fdbk_g1 … 9 more lines
- list
PRAGMA table_info(fdbk_g1);cid | name | type | notnull | dflt_value | pk 0 | id | INTEGER | 0 | NULL | 1 1 | item_id | TEXT | 0 | NULL | 0 … 8 more lines
- list
PRAGMA table_info(items_g1);cid | name | type | notnull | dflt_value | pk 0 | ref_id | TEXT | 0 | NULL | 1 1 | ttl | TEXT | 0 | NULL | 0 … 8 more lines
- sample
SELECT main_cat FROM items_g1 LIMIT 5;main_cat All Electronics Amazon Home … 3 more lines
- compute
SELECT main_cat, COUNT(*) FROM items_g1 GROUP BY main_cat;main_cat | COUNT(*) NULL | 741 AMAZON FASHION | 224 All Beauty | 25 All Electronics | 1308 Amazon Fire TV | 1 … 28 more lines
- compute
SELECT COUNT(DISTINCT ref_id) FROM fdbk_g1 WHERE rtg <= 2.0 AND ref_id IN (SELECT ref_id FROM items_g1 WHERE main_cat = 'Office Products');COUNT(DISTINCT ref_id) 5523
- compute
SELECT COUNT(DISTINCT item_id) FROM fdbk_g1 WHERE rtg <= 2.0 AND ref_id IN (SELECT ref_id FROM items_g1 WHERE main_cat = 'Office Products');COUNT(DISTINCT item_id) 6776
Learner, run 1
- read schema
SELECT name, sql FROM sqlite_master WHERE type='table'name | sql items_g1 | CREATE TABLE items_g1 ( ref_id TEXT PRIMARY KEY, ttl TEXT, avg_rtg REAL, avg_rtg_30d R… fdbk_g1 | CREATE TABLE fdbk_g1 ( id INTEGER PRIMARY KEY AUTOINCREMENT, item_id TEXT, ref_id TEXT, … … 10 more lines
- compute
SELECT COUNT(DISTINCT item_id) AS n_item_id, COUNT(DISTINCT ref_id) AS n_ref_id, COUNT(DISTINCT item_id) - COUNT(DISTINCT ref_id) AS diff FROM fdbk_g1 WHERE rtg <= 2.0;n_item_id | n_ref_id | diff 10059 | 8227 | 1832
The baseline lists the tables, inspects two of them, then filters the reviews to products whose catalog category is "Office Products". That filter is the mistake: the office dataset is the whole _g1 table group, and the category column also holds other categories. It answers 1253. Run 1 reaches this question 17th. It reads the schema in one query and computes the difference over the whole office review table in its second query: 1832, correct, two queries. Run 2 answers the same question the same way, also in two queries.
Three more questions, each asked before the migration, so both agents faced the same database:
Electronics reviews longer than 500 characters, products at most $40
Question. Among electronics reviews with a review body longer than 500 characters, considering only products whose list price is at most $40, what percentage are 1-star? Round to 2 decimal places.
Stateless baseline
- list
.tablesattrs_g1 attrs_g3 fdbk_g1 … 9 more lines
- list
PRAGMA table_info(items_g1);cid | name | type | notnull | dflt_value | pk 0 | ref_id | TEXT | 0 | NULL | 1 1 | ttl | TEXT | 0 | NULL | 0 … 8 more lines
- sample
SELECT main_cat FROM items_g1 LIMIT 5;main_cat All Electronics Amazon Home … 3 more lines
- sample
SELECT main_cat FROM items_g2 LIMIT 5;main_cat Computers Computers … 3 more lines
- sample
SELECT main_cat FROM items_g3 LIMIT 5;main_cat Musical Instruments Musical Instruments … 3 more lines
- list
PRAGMA table_info(fdbk_g1);cid | name | type | notnull | dflt_value | pk 0 | id | INTEGER | 0 | NULL | 1 1 | item_id | TEXT | 0 | NULL | 0 … 8 more lines
- list
PRAGMA table_info(fdbk_g2);cid | name | type | notnull | dflt_value | pk 0 | id | INTEGER | 0 | NULL | 1 1 | item_id | TEXT | 0 | NULL | 0 … 8 more lines
- list
PRAGMA table_info(fdbk_g3);cid | name | type | notnull | dflt_value | pk 0 | id | INTEGER | 0 | NULL | 1 1 | ref_id | TEXT | 0 | NULL | 0 … 7 more lines
- sample
SELECT DISTINCT main_cat FROM items_g1 WHERE main_cat LIKE '%Electronics%' LIMIT 10;main_cat All Electronics Car Electronics
- sample
SELECT DISTINCT main_cat FROM items_g2 WHERE main_cat LIKE '%Electronics%' LIMIT 10;main_cat All Electronics Car Electronics
- sample
SELECT DISTINCT main_cat FROM items_g2 LIMIT 20;main_cat Computers Cell Phones & Accessories … 18 more lines
- compute
SELECT COUNT(*) FROM items_g2 WHERE main_cat = 'All Electronics';COUNT(*) 4562
- compute
SELECT COUNT(*) FROM items_g1 WHERE main_cat = 'All Electronics';COUNT(*) 1308
- compute
SELECT COUNT(*) FROM items_g3 WHERE main_cat = 'All Electronics';COUNT(*) 534
- compute
SELECT ROUND(100.0 * SUM(CASE WHEN f.rtg = 1 THEN 1 ELSE 0 END) / COUNT(*), 2) AS pct_1star FROM fdbk_g2 f JOIN items_g2 i ON f.ref_id = i.ref_id WHERE i.main_cat = 'All Electronics' AND i.prc <= 40 AND LENGTH(f.body) > 500;pct_1star NULL
Learner, run 1
- read schema
SELECT name, sql FROM sqlite_master WHERE type='table'name | sql items_g1 | CREATE TABLE items_g1 ( ref_id TEXT PRIMARY KEY, ttl TEXT, avg_rtg REAL, avg_rtg_30d R… fdbk_g1 | CREATE TABLE fdbk_g1 ( id INTEGER PRIMARY KEY AUTOINCREMENT, item_id TEXT, ref_id TEXT, … … 10 more lines
- compute
SELECT ROUND(100.0 * SUM(CASE WHEN rtg = 1 THEN 1 ELSE 0 END) / COUNT(*), 2) AS pct_1_star FROM fdbk_g2 JOIN items_g2 i ON fdbk_g2.ref_id = i.ref_id WHERE LENGTH(body) > 500 AND prc_usd <= 40pct_1_star 10.22
The baseline spends all fifteen queries working out which category values mean electronics, then filters on the column prc, which is empty in the electronics table, gets NULL and answers 0.00. Run 1 uses the electronics table group and its prc_usd price column: 10.22 in two queries. Section 4 traces where it learned that column.
Average review rating of office products above $50
Question. What is the average review rating given to office products that cost more than $50?
Stateless baseline
- list
.tablesattrs_g1 attrs_g3 fdbk_g1 … 9 more lines
- list
PRAGMA table_info(items_g1);cid | name | type | notnull | dflt_value | pk 0 | ref_id | TEXT | 0 | NULL | 1 1 | ttl | TEXT | 0 | NULL | 0 … 8 more lines
- sample
SELECT DISTINCT main_cat FROM items_g1 LIMIT 10;main_cat All Electronics Amazon Home … 8 more lines
- compute
SELECT AVG(avg_rtg) FROM items_g1 WHERE main_cat = 'Office Products' AND prc > 50;AVG(avg_rtg) 4.16020942408377
Learner, run 2
- read schema
SELECT name, sql FROM sqlite_master WHERE type='table'name | sql items_g1 | CREATE TABLE items_g1 ( ref_id TEXT PRIMARY KEY, ttl TEXT, avg_rtg REAL, avg_rtg_30d R… fdbk_g1 | CREATE TABLE fdbk_g1 ( id INTEGER PRIMARY KEY AUTOINCREMENT, item_id TEXT, ref_id TEXT, … … 10 more lines
- compute
SELECT COUNT(*) FROM items_g1 WHERE main_cat LIKE '%Office%';COUNT(*) 13933
- compute
SELECT AVG(f.rtg) AS avg_rating FROM fdbk_g1 f JOIN items_g1 i ON f.ref_id = i.ref_id WHERE i.prc > 50;avg_rating 4.105430924150588
The baseline averages the catalog's per-product rating column and answers 4.16. The question asks about the ratings given in reviews. Run 2 averages the review rows and answers 4.11.
Share of 'Computers & Accessories' products priced above the category average
Question. Among electronics products in the 'Computers & Accessories' top-level category, what percentage are priced above the category-wide average? Round to 2 decimal places. Only consider products with a non-zero price.
Stateless baseline
- list
.tablesattrs_g1 attrs_g3 fdbk_g1 … 9 more lines
- list
PRAGMA table_info(items_g1)cid | name | type | notnull | dflt_value | pk 0 | ref_id | TEXT | 0 | NULL | 1 1 | ttl | TEXT | 0 | NULL | 0 … 8 more lines
- sample
SELECT DISTINCT main_cat FROM items_g1 LIMIT 10main_cat All Electronics Amazon Home … 8 more lines
- compute
SELECT DISTINCT main_cat FROM items_g1 WHERE main_cat LIKE '%Computer%'main_cat Computers
- sample
SELECT DISTINCT main_cat FROM items_g2 LIMIT 10main_cat Computers Cell Phones & Accessories … 8 more lines
- sample
SELECT DISTINCT main_cat FROM items_g3 LIMIT 10main_cat Musical Instruments NULL … 8 more lines
- compute
SELECT main_cat, COUNT(*) as cnt FROM items_g2 GROUP BY main_cat ORDER BY cnt DESC LIMIT 10main_cat | cnt Computers | 5247 All Electronics | 4562 Camera & Photo | 2789 Cell Phones & Accessories | 1775 NULL | 1304 … 5 more lines
- compute
SELECT AVG(prc) as avg_price FROM items_g2 WHERE main_cat = 'Computers' AND prc > 0avg_price 14488.249662618084
- compute
SELECT COUNT(*) as total_count, SUM(CASE WHEN prc > (SELECT AVG(prc) FROM items_g2 WHERE main_cat = 'Computers' AND prc > 0) THEN 1 ELSE 0 END) as above_avg_count FROM items_g2 WHERE main_cat = 'Computers' AND prc > 0total_count | above_avg_count 1482 | 298
Learner, run 2
- read schema
SELECT name, sql FROM sqlite_master WHERE type='table'name | sql items_g1 | CREATE TABLE items_g1 ( ref_id TEXT PRIMARY KEY, ttl TEXT, avg_rtg REAL, avg_rtg_30d R… fdbk_g1 | CREATE TABLE fdbk_g1 ( id INTEGER PRIMARY KEY AUTOINCREMENT, item_id TEXT, ref_id TEXT, … … 10 more lines
- compute
SELECT ROUND(100.0 * SUM(CASE WHEN prc > avg_prc THEN 1 ELSE 0 END) / COUNT(*), 2) AS pct_above_avg FROM (SELECT i.prc, AVG(i.prc) OVER () AS avg_prc FROM items_g2 i JOIN taxn_g2 t ON i.ref_id = t.ref_id WHERE t.cat_lvl = 1 AND t.cat_nm = 'Computers & Accessories' AND i.prc > 0);pct_above_avg 17.27
"Top-level category" refers to the product taxonomy table, not the catalog's category column. The baseline filters on the category value "Computers" and answers 20.11. Run 2 joins the taxonomy at level 1 and answers 17.27.
4. What it learned
The figure below draws every step of every question, in the order the run answered them, for the baseline and for run 1. Each row is a question. Each cell is one query, colored by what the query does.
The baseline's rows look the same from question 1 to question 40: list the tables, look at sample rows, compute, answer. Run 1's rows look like the baseline's for its first eleven questions. From question 12 on, every row starts with one query that reads the full schema, and most rows end after two or three queries. The same switch happens in the other two runs, at question 6 in run 0 and question 9 in run 2:
Counted per block of ten questions. A sample-row query looks at a few raw rows of a table (for example SELECT … LIMIT 5) instead of computing an answer. A full-schema read asks the database for all its table definitions at once. Touching a category column means filtering or grouping on the catalog's category column, the mistake shown in section 3.
| per block of ten questions (positions 1–10, 11–20, 21–30, 31–40) | baseline | run 0 | run 1 | run 2 |
|---|---|---|---|---|
| questions opened with one full-schema read | 0 · 0 · 0 · 0 | 4 · 7 · 10 · 10 | 0 · 9 · 10 · 10 | 1 · 8 · 10 · 10 |
| sample-row queries | 35 · 37 · 34 · 48 | 30 · 4 · 5 · 0 | 24 · 1 · 0 · 0 | 22 · 0 · 0 · 1 |
| questions that touch a category column | 9 · 9 · 10 · 9 | 9 · 6 · 9 · 3 | 10 · 4 · 3 · 3 | 8 · 4 · 2 · 6 |
| queries in total | 95 · 108 · 127 · 120 | 108 · 88 · 80 · 30 | 90 · 54 · 42 · 37 | 100 · 53 · 70 · 76 |
| repeats of an identical earlier query | 0 · 3 · 0 · 0 | 21 · 11 · 4 · 0 | 19 · 17 · 11 · 8 | 3 · 20 · 28 · 42 |
Runs 0 and 1 need fewer queries with each block of ten questions: 108 down to 30 and 90 down to 37. The baseline never improves, from 95 to 127. Run 2 falls to 53 and then rises again, to 70 and 76, where its repeated queries start (section 9). In every run, sample-row queries drop from 22 to 30 in the first block to at most 5 in each later block.
The full-schema opening is one of the things the recipe teaches directly: every boundary adds a training example whose target is that query. It is trained from the first question on and the learner adopts it after five to eleven questions. The rest of what changes is specific to this database.
A lesson and where it was used
Run 1 answered the 'Computers & Accessories' question ninth and got it wrong: it used the office table group and the column prc, and answered 17.69 against the correct 17.27.
Learner, run 1, question 9
- list
.tablesattrs_g1 attrs_g3 fdbk_g1 … 9 more lines
- list
PRAGMA table_info(items_g1)cid | name | type | notnull | dflt_value | pk 0 | ref_id | TEXT | 0 | NULL | 1 1 | ttl | TEXT | 0 | NULL | 0 … 8 more lines
- sample
SELECT DISTINCT main_cat FROM items_g1 LIMIT 10main_cat All Electronics Amazon Home … 8 more lines
- compute
SELECT DISTINCT main_cat FROM items_g1 WHERE main_cat LIKE '%Computer%'main_cat Computers
- compute
SELECT ROUND(100.0 * SUM(CASE WHEN prc > avg_prc THEN 1 ELSE 0 END) / COUNT(*), 2) AS pct_above_avg FROM items_g1, (SELECT AVG(prc) AS avg_prc FROM items_g1 WHERE main_cat = 'Computers' AND prc > 0) WHERE main_cat = 'Computers' AND prc > 0pct_above_avg 17.69
After that question the teacher wrote this procedure (verbatim):
Electronics-dataset category questions that name a taxonomy category ('Computers & Accessories') use the g2 table group: items_g2 joined to taxn_g2 by ref_id on cat_nm. They do not use items_g1.main_cat='Computers', which belongs to the office dataset and was rejected (17.69 vs released 17.27). items_g2 has two price columns (prc, prc_usd), so one bounded query computes the share of non-zero-priced products above the category average for both price columns and also reports their storage types.
The recipe built 7 training examples from it at that boundary. Each one is a state the learner actually reached in its question-9 trajectory, followed by the next action the procedure prescribes. One of them, as trained (the system prompt with the action schema is omitted here):
[user]
Question 9/40
Among electronics products in the 'Computers & Accessories' top-level category, what percentage are priced above the category-wide average? Round to 2 decimal places. Only consider products with a non-zero price.
You have access to a SQLite database that contains three separate product datasets covering different categories: office products, electronics, and musical instruments. Each dataset has its own set of tables. You will need to explore the schema to determine which tables belong to which product category.
You may run up to 15 exploratory SQL queries to examine the schema and data before submitting your final answer. Note: query results are capped at 50 rows — if a que … [401 more characters]
[assistant]
{"action":"QUERY","content":".tables"}
[user]
FEEDBACK: Query result (1/15 queries used, 14 remaining):
attrs_g1
attrs_g3
fdbk_g1
fdbk_g2
fdbk_g3
fdbk_stats_g1
items_g1
items_g2
items_g3
sqlite_sequence
taxn_g1
taxn_g2
[user]
Question 9/40
Among electronics products in the 'Computers & Accessories' top-level category, what percentage are priced above the category-wide average? Round to 2 decimal places. Only consider products with a non-zero price.
You have access to a SQLite database that contains three separate product datasets covering different categories: office products, electronics, and musical instruments. Each dataset has its own set of tables. You will need to explore the schema to determine which tables belong to which product category.
You may run up to 15 exploratory SQL queries to examine the schema and data before submitting your final answer. Note: query results are capped at 50 rows — if a que … [401 more characters]
[target, trained]
{"action": "QUERY", "content": "WITH pop AS (SELECT DISTINCT i.ref_id, i.prc, i.prc_usd FROM items_g2 i JOIN taxn_g2 t ON t.ref_id = i.ref_id WHERE t.cat_nm = 'Computers & Accessories') SELECT COUNT(*) AS n_items, MAX(typeof(prc)) AS prc_type, MAX(typeof(prc_usd)) AS prc_usd_type, SUM(prc > 0) AS n_prc_pos, ROUND(100.0 * SUM(prc > (SELECT AVG(prc) FROM pop WHERE prc > 0)) / SUM(prc > 0), 2) AS pct_above_avg_prc, SUM(prc_usd > 0) AS n_prc_usd_pos, ROUND(100.0 * SUM(prc_usd > (SELECT AVG(prc_usd) FROM pop WHERE prc_usd > 0)) / SUM(prc_usd > 0), 2) AS pct_above_avg_prc_usd FROM pop;"}
Four questions later run 1 was asked a different question about electronics, the one in section 3 about reviews longer than 500 characters. Its second query used the electronics table group and prc_usd, and the answer was correct. Before question 9, run 1 had not written prc_usd in any query.
5. Learner 1.0 learns efficiently online from very few examples
A run trains once after each question, on examples built from that question alone. Over forty questions run 1 trained on 256 examples and 25,519 supervised tokens, about 654 tokens per question. Runs 0 and 2 used 76,560 and 43,456. The best run used the least.
Individual lessons are learned from a handful of examples. The table lists names the learner did not write in any query until a training example used them. Each name was already visible in the schema listings the learner had read. What the examples taught was to use it.
| run | name | what it is | taught after question | examples | supervised tokens | later questions using it | correct |
|---|---|---|---|---|---|---|---|
| run 1 | product_attributes_g3 | musical-instrument attribute table added by the migration | 22 | 5 | 610 | 5 | 4 |
| run 0 | product_attributes_g3 | musical-instrument attribute table added by the migration | 24 | 5 | 2,228 | 4 | 2 |
| run 2 | verified_status | office review verification label added by the migration | 31 | 1 | 280 | 1 | 1 |
| run 1 | prc_usd | electronics price in dollars | 9 | 5 | 1,105 | 7 | 1 |
| run 1 | attrs_g1 | office product attribute table | 14 | 5 | 1,220 | 2 | 2 |
| run 1 | cat_lvl | taxonomy level column | 7 | 7 | 1,834 | 4 | 1 |
| run 2 | item_id | second product identifier in office reviews | 2 | 3 | 603 | 2 | 1 |
The migration's new attribute table is the clearest case. Run 1 read it in the schema listing at questions 21 and 22 and did not query it. After question 22 it trained on five examples that did, 610 supervised tokens in total, and from question 25 on it used the table in five questions and answered four of them correctly. Run 2 learned the migration's new office verification label from a single example of 280 tokens and used it correctly two questions later.
The training loss shows the same thing. Each training example is used for one update and never again. Its loss is recorded in that same step, before the update, so every point below is the model's error on an example it has not trained on yet. A falling line means the model gets better at examples it has never trained on: it is learning how to act on this database, not memorizing the examples it was shown.
6. After the migration
At question 21 the database is migrated: the musical-instrument attribute table is renamed and replaced by a new one, musical-instrument timestamps are split into year and month columns, office reviews gain a partially filled verification label, and electronics gain a second price column. The paper that introduces the benchmark (Asawa et al., 2026, arXiv:2606.05661) illustrates the failure this is meant to catch: an agent that reuses what it learned before the migration and queries tables that no longer exist.
None of the three runs did that. Every one of their sixty post-migration questions opened by reading the full schema, and no run sent a single query to the renamed table under its old name. The full-schema opening, learned in the first half, is what exposes the new schema.
Two of the three product groups recovered. Run 1 shows how. On question 22 it looked for brand in the product table's brand column, which is empty for every instrument, and failed:
Learner, run 1, question 22
- read schema
SELECT name, sql FROM sqlite_master WHERE type='table'name | sql items_g1 | CREATE TABLE items_g1 ( ref_id TEXT PRIMARY KEY, ttl TEXT, avg_rtg REAL, avg_rtg… fdbk_g1 | CREATE TABLE fdbk_g1 ( id INTEGER PRIMARY KEY AUTOINCREMENT, item_id TEXT, ref_id TE… … 11 more lines
- compute
WITH n AS (SELECT COUNT(DISTINCT main_cat) AS n_cats FROM items_g1 UNION ALL SELECT COUNT(DISTINCT main_cat) AS n_cats FROM items_g2 UNION ALL SELECT COUNT(DISTINCT main_cat) AS n_cats FROM items_g3), r AS (SELECT 'g1' AS grp, main_cat, COUNT(DISTINCT ref_id) AS n_items FROM items_g1 GROUP BY main_cat UNION ALL SELECT 'g2' AS grp, main_cat, COUNT(DISTINCT ref_id) AS n_items FROM items_g2 GROUP BY main_cat UNION ALL SELECT 'g3' AS grp, main_cat, COUNT(DISTINCT ref_id) AS n_items FROM items_g3 GROUP BY main_cat) SELECT r.* FROM r JOIN n ON r.grp = CASE WHEN n.n_cats = r.n_items THEN 'g1' ELSE 'g2' END LIMIT 10;grp | main_cat | n_items g2 | NULL | 1304 g2 | AMAZON FASHION | 232 g2 | All Beauty | 25 g2 | All Electronics | 4562 g2 | Amazon Devices | 106 … 5 more lines
- compute
WITH n AS (SELECT COUNT(DISTINCT main_cat) AS n_cats FROM items_g1 UNION ALL SELECT COUNT(DISTINCT main_cat) AS n_cats FROM items_g2 UNION ALL SELECT COUNT(DISTINCT main_cat) AS n_cats FROM items_g3), r AS (SELECT 'g1' AS grp, main_cat, COUNT(DISTINCT ref_id) AS n_items FROM items_g1 GROUP BY main_cat UNION ALL SELECT 'g2' AS grp, main_cat, COUNT(DISTINCT ref_id) AS n_items FROM items_g2 GROUP BY main_cat UNION ALL SELECT 'g3' AS grp, main_cat, COUNT(DISTINCT ref_id) AS n_items FROM items_g3 GROUP BY main_cat) SELECT r.* FROM r JOIN n ON r.grp = CASE WHEN n.n_cats = r.n_items THEN 'g1' ELSE 'g2' END WHERE r.main_cat LIKE '%Music%' OR r.main_cat LIKE '%Instrument%' LIMIT 20;grp | main_cat | n_items g2 | Musical Instruments | 85 g2 | Musical Instruments | 85 g2 | Musical Instruments | 85
- compute
WITH brands AS (SELECT str_nm AS brand, COUNT(DISTINCT ref_id) AS n_products, AVG(prc_usd) AS avg_price FROM items_g2 WHERE main_cat = 'Musical Instruments' GROUP BY str_nm HAVING n_products >= 50) SELECT brand, n_products, avg_price FROM brands ORDER BY avg_price DESC, brand ASC LIMIT 10;(no results)
- compute
WITH BrandStats AS (SELECT str_nm AS brand, COUNT(DISTINCT ref_id) AS n_products, AVG(prc) AS avg_price FROM items_g3 GROUP BY str_nm HAVING n_products >= 50) SELECT brand, n_products, avg_price FROM BrandStats ORDER BY avg_price DESC, brand ASC LIMIT 10;brand | n_products | avg_price NULL | 20000 | 123.38623097244937
- compute
WITH BrandStats AS (SELECT str_nm AS brand, COUNT(DISTINCT ref_id) AS n_products, AVG(prc) AS avg_price FROM items_g3 WHERE str_nm IS NOT NULL GROUP BY str_nm HAVING n_products >= 50) SELECT brand, n_products, avg_price FROM BrandStats ORDER BY avg_price DESC, brand ASC LIMIT 10;(no results)
- compute
WITH BrandStats AS (SELECT str_nm AS brand, COUNT(DISTINCT ref_id) AS n_products, AVG(prc) AS avg_price FROM items_g1 GROUP BY str_nm HAVING n_products >= 50) SELECT brand, n_products, avg_price FROM BrandStats ORDER BY avg_price DESC, brand ASC LIMIT 10;brand | n_products | avg_price HP | 104 | 374.79033898305084 Canon | 66 | 227.96085714285715 Amazon Renewed | 66 | 187.9235294117647 Brother | 66 | 137.94026315789472 Epson | 65 | 123.12 … 5 more lines
- compute
WITH BrandStats AS (SELECT str_nm AS brand, COUNT(DISTINCT ref_id) AS n_products, AVG(prc) AS avg_price FROM items_g1 WHERE main_cat = 'Musical Instruments' GROUP BY str_nm HAVING n_products >= 50) SELECT brand, n_products, avg_price FROM BrandStats ORDER BY avg_price DESC, brand ASC LIMIT 10;(no results)
- compute
WITH BrandStats AS (SELECT str_nm AS brand, COUNT(DISTINCT ref_id) AS n_products, AVG(prc_usd) AS avg_price FROM items_g2 WHERE main_cat = 'Musical Instruments' AND str_nm IS NOT NULL GROUP BY str_nm HAVING n_products >= 50) SELECT brand, n_products, avg_price FROM BrandStats ORDER BY avg_price DESC, brand ASC LIMIT 10;(no results)
- repeat ×4
WITH BrandStats AS (SELECT str_nm AS brand, COUNT(DISTINCT ref_id) AS n_products, AVG(prc) AS avg_price FROM items_g3 WHERE str_nm IS NOT NULL GROUP BY str_nm HAVING n_products >= 50) SELECT brand, n_products, avg_price FROM BrandStats ORDER BY avg_price DESC, brand ASC LIMIT 10;the same query, sent 4 times in a row (steps 10–13)
The lesson written after that question said where brand now lives: the new attribute table. Three questions later run 1 used it in its second query and answered correctly:
Learner, run 1, question 25
- read schema
SELECT name, sql FROM sqlite_master WHERE type='table'name | sql items_g1 | CREATE TABLE items_g1 ( ref_id TEXT PRIMARY KEY, ttl TEXT, avg_rtg REAL, avg_rtg… fdbk_g1 | CREATE TABLE fdbk_g1 ( id INTEGER PRIMARY KEY AUTOINCREMENT, item_id TEXT, ref_id TE… … 11 more lines
- compute
SELECT COUNT(DISTINCT F.ref_id) AS n_products FROM product_attributes_g3 F LEFT JOIN attrs_g3_legacy L ON F.ref_id = L.ref_id AND L.attr_key = 'Brand' WHERE F.attr_key = 'Brand' AND F.attr_val IS NOT NULL AND F.attr_val != '' AND L.ref_id IS NULLERROR: query exceeded 10s timeout
- compute
WITH F AS (SELECT ref_id FROM product_attributes_g3 WHERE attr_key = 'Brand' AND attr_val IS NOT NULL AND attr_val != ''), L AS (SELECT DISTINCT ref_id FROM attrs_g3_legacy WHERE attr_key = 'Brand') SELECT (SELECT COUNT(*) FROM F) - (SELECT COUNT(*) FROM F WHERE ref_id IN (SELECT ref_id FROM L)) AS n_productsn_products 1145
Electronics did not recover. The migration added a new price column to the electronics table, next to the dollar column the learner had learned. The teacher's lessons after questions 27, 31, 34 and 35 call the competing price columns unresolved, and no run answered more than one post-migration electronics question correctly.
7. Results
| lane | correct of 40 | reward (of 40) | gain over baseline | queries | supervised tokens trained | vs baseline: gained / lost | exact McNemar p |
|---|---|---|---|---|---|---|---|
| Run 0 | 13 | 8.87 | +7.47 | 306 | 76,560 | 10 / 3 | 0.092 |
| Run 1 | 16 | 12.00 | +10.60 | 223 | 25,519 | 12 / 2 | 0.013 |
| Run 2 | 14 | 9.13 | +7.73 | 299 | 43,456 | 11 / 3 | 0.057 |
| mean of the three runs | 14.33 ± 0.88 | 10.00 | +8.60 | 276 | 48,512 | ||
| Stateless baseline | 6 | 1.40 | 0 | 450 | 0 |
Each run is compared with the baseline question by question. "Gained" counts questions the run answers correctly and the baseline gets wrong. "lost" counts the reverse. Pooled over the three runs, the learner gains 33 questions and loses 8. Every question the baseline answers correctly is also answered correctly by at least one run.
8. Against the published leaderboard
Learner 1.0 is the only system on this leaderboard that learns strictly through its weights. Its context is reset before every question, so no in-context learning is allowed between questions. Every other system carries its experience forward in its context, a notepad, a memory system or a coding agent's workspace.
The benchmark publishes results for twelve systems on this task, each pairing a frontier model with a way of carrying experience forward: in context, through a notepad, through a memory system, or through a coding agent's workspace. Our three runs average a reward of 10.00 and a gain of +8.60.
| system | runs | reward | reward rank | gain | normalized gain | gain rank |
|---|---|---|---|---|---|---|
| Claude Code · Sonnet 4.6 | 5 | 22.05 | 1 | +13.85 | 0.436 | 1 |
| Mem0 · GPT-5.4 | 5 | 17.24 | 2 | +12.91 | 0.362 | 2 |
| ICL · Claude Opus 4.7 | 5 | 15.65 | 3 | +9.59 | 0.283 | 4 |
| ICL · Gemini 3 Flash | 5 | 15.03 | 4 | +11.49 | 0.315 | 3 |
| ICL · Claude Sonnet 4.6 | 5 | 15.01 | 5 | +8.48 | 0.253 | 5 |
| ICL · GPT-5.4 | 5 | 13.88 | 6 | +8.35 | 0.242 | 6 |
| ICL Notepad · GPT-5.4 | 5 | 12.37 | 7 | +6.37 | 0.187 | 9 |
| ICL · Gemini 3.1 Pro Preview | 5 | 11.56 | 8 | +6.83 | 0.194 | 8 |
| ICL Notepad · Claude Sonnet 4.6 | 5 | 11.00 | 9 | +3.67 | 0.112 | 12 |
| Learner 1.0 (ours) | 3 | 10.00 | 10 | +8.60 | 0.223 | 7 |
| Codex · GPT-5.4 | 1 | 9.60 | 11 | +6.13 | 0.168 | 10 |
| ICL Notepad · Gemini 3.1 Pro Preview | 5 | 8.52 | 12 | +4.32 | 0.121 | 11 |
| ACE · GPT-5.4 | 5 | 7.85 | 13 | +2.39 | 0.069 | 13 |
On reward we rank 10 of 13. Gain is ranked as the leaderboard ranks it, normalized by each system's headroom (gain divided by 40 minus the system's own baseline). On that measure we rank 7 of 13. Our baseline is weaker than the frontier baselines (1.40 against 5.53 for ICL · GPT-5.4), so it has more headroom, which the normalization takes into account.
9. Where it fails
The learner has a failure mode the baseline does not: it sends a query it already sent in the same question. The database is read-only, so the second copy returns the same result and spends one of the fifteen queries for nothing. The baseline repeats an identical query 3 times in 450 queries. The three runs repeat 36, 55 and 93 times. Once the learner has repeated, its next query is again a repeat 78 to 98 percent of the time, and questions with a repeat score 0.01 to 0.08 on average, against 0.27 to 0.36 for the rest. Here is run 1 on its sixth question:
Learner, run 1, question 6
- list
SELECT name FROM sqlite_master WHERE type='table';name items_g1 fdbk_g1 … 10 more lines
- sample
SELECT * FROM items_g1 LIMIT 5;ref_id | ttl … B001DZWCDO | LD Products Compatible Toner Cartridge Replacement for Samsung ML-2850 Series ML-D2850B High Yield (Black) … B096Z4NJLV | Hodiczk Small Desk Shelf Organizer White with Drawer Cute Desktop Organizer Shelf Storage Plastic Kawai Des… … 3 more lines
- compute
SELECT COUNT(*) FROM fdbk_g1 WHERE main_cat = 'Office Products' AND prc > 20 AND prc < 200 AND strftime('%Y', dt) = '2022';ERROR: no such column: main_cat
- compute
SELECT COUNT(*) FROM fdbk_g1 WHERE ref_id IN (SELECT ref_id FROM items_g1 WHERE main_cat = 'Office Products' AND prc > 20 AND prc < 200) AND dt >= '2022-01-01' AND dt < '2023-01-01';ERROR: no such column: dt
- repeat ×11
SELECT COUNT(*) FROM fdbk_g1 WHERE ref_id IN (SELECT ref_id FROM items_g1 WHERE main_cat = 'Office Products' AND prc > 20 AND prc < 200) AND dt >= '2022-01-01' AND dt < '2023-01-01';the same query, sent 11 times in a row (steps 5–15)
Run 2's second half has the most repeats, and its score drops there: 8 correct in the first twenty questions and 6 in the second, while runs 0 and 1 improve from the first half to the second.
What causes it. A repeat usually starts right after a result that told the learner nothing new: an error, an empty table, or a count of zero. At those moments the untrained model never repeats (0 of 25), while the trained runs repeat at 30 to 67 percent of them (after an SQL error: 3 of 20, 12 of 19 and 22 of 33 in runs 0, 1 and 2). Replaying the exact recorded conversations, byte for byte, confirms that the trained weights make the choice: at the 30 moments where a trained run began a repeat, the trained checkpoint repeats all 30 times, and the untrained model, given the same conversation, sends a new query 26 times. The repeats are not memorized lessons either: repeated queries resemble the trained targets less than the learner's other queries do.
The cause is what the training leaves out. This recipe is heavy on knowledge and light on state: it teaches what to query for a kind of question, and almost never what to do after a query fails or returns nothing. Over a whole run it trains 22 to 25 examples about the database's state, none of them after an empty or zero result. An earlier controlled study on a synthetic 40-episode panel shows why that matters: a knowledge-only recipe produced 335 repeated failing queries, adding 1,024 more knowledge examples still left 203, and adding instead 704 examples of the next step at failure states left none.
10. The recipe
After each question. The teacher receives the learner's trajectory on the question just answered, the benchmark's released feedback for it, the learner's trajectories on earlier questions as evidence, and the procedures it wrote before. It returns procedures in a fixed JSON format: what the question needed, the states that lead to an answer, and the exact next query or answer at each state. Its full prompt is published in the replication repository (recipe/teacher_prompt.txt).
From procedures to examples. The recipe builds each example from a state the learner actually reached: the conversation up to that point, with the recorded query results, and the next action as the target. It keeps only examples grounded in recorded results, adds three fixed examples about reading the database's state (the full-schema opening among them), and caps any single procedure at a quarter of a boundary's examples. Every example is trained once.
Training and limits. One pass, batch one, no replay. The learner's context is 32,000 tokens and its answers are limited to 4,096 tokens. The teacher call is limited to 15 minutes.
11. Interpretation, and how to check it
Training on each answered question changed how the learner explores this database: it stops sampling tables, reads the schema once, and goes straight to the tables and columns earlier questions taught it. That produced more correct answers in fewer queries on all three runs. The size of the effect varies by run (13 to 16 correct), and the learner picks up a repetition failure the baseline does not have.
The baseline answered questions 21–40 on the pre-migration database, as the benchmark's baseline protocol does. The section-3 examples are all pre-migration questions for that reason.
To check: every trajectory of every lane, the feedback, the training examples and the scores are in the replication repository, and the figures regenerate from them.