Demonstrations · Continual Learning Bench
Forty database questions, one model learning as it answers
On the database exploration task of Continual Learning Bench, our Learner 1.0 model answers forty questions about one product database in sequence and is trained on each question after it answers it. Over three runs it answers 21, 15 and 16 questions correctly where the same model with learning off answers 7.
Model Learner 1.0
Benchmark Continual Learning Bench, database exploration, default schedule
Questions 40 per run, 20 before and 20 after a schema migration
Runs 3 question orders (0 canonical, 1 and 2 permuted) and one stateless baseline
Learning after every answered question. One pass, one example per update, no replay
What the learner sees only the current question, exactly as the benchmark presents it
Teacher a teacher LLM writes the training examples after each question
Compute one H200 per run
Learner 1.0 is the only system on this benchmark that learns strictly through its weights. Its context is reset before every question, so there is no in-context learning between questions: anything it carries from one question to the next is in its weights. Every other system on the public leaderboard carries its experience forward in its context, a notepad, a memory store or a coding agent’s workspace.
1. The question and what was measured
The question: can our continual learning agent learn from its own experience by updating its weights? There are 40 questions in the benchmark, and the agent must learn from its own experience incrementally and sequentially, and improve its performance. We use only weight-based learning and reset the context between questions, so there is no in-context learning between questions. If the agent improves over the baseline, it must be due to online parametric continual learning.
There are two challenges. 1) Data efficiency: each question of the benchmark generates very little data, so the agent must learn efficiently from just a handful of examples. 2) It must retain its learnings in the weights as it is incrementally trained. As per our understanding, no other technique today can achieve these two traits when it comes to parametric continual learning.
Each question is a natural-language question about a SQLite database of three product datasets (office products, electronics, musical instruments). The agent may run up to 15 exploratory queries, then submits one answer, and is told whether it was right and, if not, the correct answer. Reward for a question is 1 − queries/15 if the answer is correct and 0 if it is wrong. After question 20 the database is migrated to a new schema, and questions 21–40 are asked against it.
Reward is the sum over the forty questions (at most 40). Gain is reward minus the reward of the same model answering the same questions with learning off, which the benchmark calls the stateless baseline. Gain is the benchmark's learning metric.
The three runs end at rewards 12.13, 7.27 and 10.67 against 1.87 for the baseline.
2. Setup
Model. Learner 1.0, the same model as in the ten-skill experiment, with a fixed number of trainable parameters set before the first question. The mechanism is proprietary and is not described here.
What the learner sees. At each question the learner receives the benchmark's own system prompt and the current question, and nothing else: no earlier questions, no notes, no retrieved examples. Its prompt is the same as the baseline's. Anything it carries from one question to the next is in its weights.
Thinking when stuck. Before each query the learner normally writes its next action directly. When it is stuck, it first writes up to 512 tokens of reasoning that the benchmark does not see, and then the action. It counts as stuck when its last query returned an error or no rows, or when it has used 8 or more of its 15 queries on the question. The rule reads only the text the benchmark returns, and the baseline uses the same rule. The learner thought before 53, 66 and 39 of its 299, 291 and 228 actions in the three runs. The baseline before 198 of 522.
The loop. The learner answers a question. The benchmark scores it and releases its feedback. A teacher LLM reads the learner's trajectory on that question and the feedback, and writes procedures: what the question needed and which queries answer it. The recipe turns each procedure into training examples built on states the learner actually reached, and the learner is trained on them before the next question. The teacher never answers a question and never acts in the benchmark.
Training. One pass, batch size one, no replay of earlier examples, no optimizer resets between questions. Across a run the learner trained on 240 to 539 examples and 20,547 to 45,096 supervised tokens.
Runs. The benchmark's default schedule asks the same forty questions in a fixed canonical order (run 0) and in permuted orders that keep the twenty pre-migration questions first (runs 1 and 2). The stateless baseline is run once. Nothing carries over between its questions, so its order does not matter.
3. A few example questions compared, before and after learning
Three questions asked before the migration, so both agents faced the same database. Each is one the baseline gets wrong and run 0 gets right.
For each question the two columns show the agent's actions in order. Each numbered step is one SQL query the agent sent, with the start of the result the database returned below it. The last line is the answer it submitted and whether it was right. The left column is the stateless baseline, the same model with learning off. The right column is the learner. The full trajectories of every run, all forty questions each, are in the replication repository.
Question. What is the average review rating given to office products that cost more than $50?
Stateless baseline
- list
.tablesattrs_g1 attrs_g3 fdbk_g1 … 9 more lines
- list
PRAGMA table_info(items_g1);cid | name | type | notnull | dflt_value | pk 0 | ref_id | TEXT | 0 | NULL | 1 1 | ttl | TEXT | 0 | NULL | 0 … 8 more lines
- sample
SELECT DISTINCT main_cat FROM items_g1 LIMIT 10;main_cat All Electronics Amazon Home … 8 more lines
- compute
SELECT AVG(avg_rtg) FROM items_g1 WHERE main_cat = 'Office Products' AND prc > 50;AVG(avg_rtg) 4.16020942408377
Learner, run 0
- list
.tablesattrs_g1 attrs_g3 fdbk_g1 … 9 more lines
- read schema
SELECT name, sql FROM sqlite_master WHERE type='table' AND name LIKE '%g1%'name | sql items_g1 | CREATE TABLE items_g1 ( ref_id TEXT PRIMARY KEY, ttl TEXT, avg_rtg REAL, avg_rtg_30d REA… fdbk_g1 | CREATE TABLE fdbk_g1 ( id INTEGER PRIMARY KEY AUTOINCREMENT, item_id TEXT, ref_id TEXT, … … 3 more lines
- compute
SELECT AVG(rtg) FROM fdbk_g1 WHERE ref_id IN (SELECT ref_id FROM items_g1 WHERE prc > 50.0);AVG(rtg) 4.105430924150588
Question. Among electronics reviews with a review body longer than 500 characters, considering only products whose list price is at most $40, what percentage are 1-star? Round to 2 decimal places.
Stateless baseline
- list
.tablesattrs_g1 attrs_g3 fdbk_g1 … 9 more lines
- list
PRAGMA table_info(items_g1);cid | name | type | notnull | dflt_value | pk 0 | ref_id | TEXT | 0 | NULL | 1 1 | ttl | TEXT | 0 | NULL | 0 … 8 more lines
- sample
SELECT main_cat FROM items_g1 LIMIT 5;main_cat All Electronics Amazon Home … 3 more lines
- sample
SELECT main_cat FROM items_g2 LIMIT 5;main_cat Computers Computers … 3 more lines
- sample
SELECT main_cat FROM items_g3 LIMIT 5;main_cat Musical Instruments Musical Instruments … 3 more lines
- list
PRAGMA table_info(fdbk_g1);cid | name | type | notnull | dflt_value | pk 0 | id | INTEGER | 0 | NULL | 1 1 | item_id | TEXT | 0 | NULL | 0 … 8 more lines
- list
PRAGMA table_info(fdbk_g2);cid | name | type | notnull | dflt_value | pk 0 | id | INTEGER | 0 | NULL | 1 1 | item_id | TEXT | 0 | NULL | 0 … 8 more lines
- list
PRAGMA table_info(fdbk_g3);cid | name | type | notnull | dflt_value | pk 0 | id | INTEGER | 0 | NULL | 1 1 | ref_id | TEXT | 0 | NULL | 0 … 7 more lines
- sample
SELECT DISTINCT main_cat FROM items_g1;main_cat All Electronics Amazon Home … 31 more lines
- sample
SELECT DISTINCT main_cat FROM items_g2;main_cat Computers Cell Phones & Accessories … 31 more lines
- compute
SELECT COUNT(*) FROM items_g2 WHERE main_cat = 'All Electronics';COUNT(*) 4562
- compute
SELECT COUNT(*) FROM items_g1 WHERE main_cat = 'Office Products';COUNT(*) 13933
- compute
SELECT COUNT(*) FROM items_g3 WHERE main_cat = 'Musical Instruments';COUNT(*) 17162
- compute
SELECT COUNT(*) FROM fdbk_g2 WHERE LENGTH(body) > 500;COUNT(*) 9875
- compute
SELECT ROUND(100.0 * SUM(CASE WHEN f.rtg = 1 THEN 1 ELSE 0 END) / COUNT(*), 2) FROM fdbk_g2 f JOIN items_g2 i ON f.ref_id = i.ref_id WHERE i.main_cat = 'All Electronics' AND i.prc <= 40 AND LENGTH(f.body) > 500;ROUND(100.0 * SUM(CASE WHEN f.rtg = 1 THEN 1 ELSE 0 END) / COUNT(*), 2) NULL
Learner, run 0
- read schema
SELECT name, sql FROM sqlite_master WHERE type='table'name | sql items_g1 | CREATE TABLE items_g1 ( ref_id TEXT PRIMARY KEY, ttl TEXT, avg_rtg REAL, avg_rtg_30d R… fdbk_g1 | CREATE TABLE fdbk_g1 ( id INTEGER PRIMARY KEY AUTOINCREMENT, item_id TEXT, ref_id TEXT, … … 10 more lines
- compute
SELECT COUNT(*) FROM fdbk_g2 WHERE LENGTH(body) > 500 AND rtg = 1.0 AND ref_id IN (SELECT ref_id FROM items_g2 WHERE prc_usd <= 40.0);COUNT(*) 222
- compute
SELECT COUNT(*) FROM fdbk_g2 WHERE LENGTH(body) > 500 AND ref_id IN (SELECT ref_id FROM items_g2 WHERE prc_usd <= 40.0);COUNT(*) 2172
Question. Among electronics products that have zero reviews, what percentage have a price listed (non-NULL, non-zero)? Round to 2 decimal places.
Stateless baseline
- list
.tablesattrs_g1 attrs_g3 fdbk_g1 … 9 more lines
- list
PRAGMA table_info(items_g1);cid | name | type | notnull | dflt_value | pk 0 | ref_id | TEXT | 0 | NULL | 1 1 | ttl | TEXT | 0 | NULL | 0 … 8 more lines
- compute
SELECT main_cat, COUNT(*) FROM items_g1 GROUP BY main_cat LIMIT 10;main_cat | COUNT(*) NULL | 741 AMAZON FASHION | 224 All Beauty | 25 All Electronics | 1308 Amazon Fire TV | 1 … 5 more lines
- compute
SELECT main_cat, COUNT(*) FROM items_g2 GROUP BY main_cat LIMIT 10;main_cat | COUNT(*) NULL | 1304 AMAZON FASHION | 232 All Beauty | 25 All Electronics | 4562 Amazon Devices | 106 … 5 more lines
- compute
SELECT main_cat, COUNT(*) FROM items_g3 GROUP BY main_cat LIMIT 10;main_cat | COUNT(*) NULL | 321 AMAZON FASHION | 177 All Beauty | 16 All Electronics | 534 Amazon Home | 253 … 5 more lines
- compute
SELECT COUNT(*) FROM items_g1 WHERE main_cat = 'All Electronics' AND rtg_ct = 0;COUNT(*) 0
- compute
SELECT COUNT(*) FROM items_g2 WHERE main_cat = 'All Electronics' AND rtg_ct = 0;ERROR: no such column: rtg_ct
- sample
SELECT * FROM taxn_g1 LIMIT 10;ref_id | cat_lvl | cat_nm B001DZWCDO | 0 | Office Products B001DZWCDO | 1 | Office Electronics … 8 more lines
- list
PRAGMA table_info(items_g2);cid | name | type | notnull | dflt_value | pk 0 | ref_id | TEXT | 0 | NULL | 1 1 | ttl | TEXT | 0 | NULL | 0 … 7 more lines
- list
PRAGMA table_info(items_g3);cid | name | type | notnull | dflt_value | pk 0 | ref_id | TEXT | 0 | NULL | 1 1 | ttl | TEXT | 0 | NULL | 0 … 6 more lines
- sample
SELECT cat_nm FROM taxn_g2 WHERE cat_lvl = 0 LIMIT 5;cat_nm Electronics Electronics … 3 more lines
- list
PRAGMA table_info(fdbk_g2);cid | name | type | notnull | dflt_value | pk 0 | id | INTEGER | 0 | NULL | 1 1 | item_id | TEXT | 0 | NULL | 0 … 8 more lines
- compute
SELECT COUNT(*) FROM items_g2 WHERE main_cat = 'All Electronics';COUNT(*) 4562
- compute
SELECT COUNT(*) FROM items_g2 WHERE main_cat = 'All Electronics' AND avg_rtg IS NULL;COUNT(*) 0
- repeat
SELECT COUNT(*) FROM items_g2 WHERE main_cat = 'All Electronics' AND avg_rtg IS NULL;COUNT(*) 0
Learner, run 0
- read schema
SELECT name, sql FROM sqlite_master WHERE type='table'name | sql items_g1 | CREATE TABLE items_g1 ( ref_id TEXT PRIMARY KEY, ttl TEXT, avg_rtg REAL, avg_rtg_30d R… fdbk_g1 | CREATE TABLE fdbk_g1 ( id INTEGER PRIMARY KEY AUTOINCREMENT, item_id TEXT, ref_id TEXT, … … 10 more lines
- compute
SELECT COUNT(*) FROM items_g2 WHERE ref_id NOT IN (SELECT ref_id FROM fdbk_g2) AND prc IS NOT NULL AND prc > 0;COUNT(*) 3395
- compute
SELECT COUNT(*) FROM items_g2 WHERE ref_id NOT IN (SELECT ref_id FROM fdbk_g2);COUNT(*) 11345
4. What it learned
The figure draws every step of every question, in the order the run answered them, for the baseline and for run 0. Each row is a question. Each cell is one query, colored by what the query does.
Counted per block of ten questions. A sample-row query looks at a few raw rows of a table (for example SELECT … LIMIT 5) instead of computing an answer. A schema read asks the database for its table definitions. A repeat is a query identical to one already sent on the same question.
| per block of ten questions (positions 1–10, 11–20, 21–30, 31–40) | baseline | run 0 | run 1 | run 2 |
|---|---|---|---|---|
| questions with a schema read in the first two queries | 0 · 0 · 0 · 0 | 0 · 6 · 5 · 5 | 1 · 6 · 6 · 10 | 2 · 10 · 10 · 10 |
| sample-row queries | 36 · 40 · 30 · 44 | 27 · 6 · 3 · 1 | 33 · 3 · 2 · 0 | 30 · 1 · 7 · 0 |
| queries in total | 108 · 113 · 133 · 125 | 86 · 63 · 52 · 57 | 103 · 47 · 61 · 37 | 83 · 40 · 42 · 22 |
| repeats of an identical earlier query | 5 · 6 · 5 · 2 | 3 · 4 · 0 · 5 | 3 · 3 · 4 · 0 | 4 · 4 · 4 · 0 |
From the second block of ten questions on, every run sends about half as many queries as the baseline or fewer, and the baseline never improves. In the last block the three runs send 57, 37 and 22 queries against the baseline's 125. Looking at sample rows almost disappears after the first ten questions, from 27 to 33 per block to 0 to 7, and reading the schema at the start of a question goes from 0 to 2 of the first ten questions to 5 to 10 of each later ten.
A lesson and where it was used
Run 0's 24th question was about musical-instrument brands, and it got it wrong. It grouped products by the catalog's brand column, which the migration had left empty for instruments, found no brand with fifty products, tried the other product groups' tables, and answered Yamaha (correct: Fender).
Learner, run 0, question 24
- list
.tablesattrs_g1 attrs_g3_legacy fdbk_g1 … 10 more lines
- list
PRAGMA table_info(items_g3);cid | name | type | notnull | dflt_value | pk 0 | ref_id | TEXT | 0 | NULL | 1 1 | ttl | TEXT | 0 | NULL | 0 … 6 more lines
- compute
SELECT str_nm, COUNT(DISTINCT ref_id) AS n_products, AVG(prc) AS avg_price FROM items_g3 GROUP BY str_nm HAVING COUNT(DISTINCT ref_id) >= 50 ORDER BY avg_price DESC, str_nm ASC LIMIT 10str_nm | n_products | avg_price NULL | 20000 | 123.38623097244937
- compute
SELECT str_nm, COUNT(DISTINCT ref_id) AS n_products, AVG(prc) AS avg_price FROM items_g3 WHERE str_nm IS NOT NULL GROUP BY str_nm HAVING COUNT(DISTINCT ref_id) >= 50 ORDER BY avg_price DESC, str_nm ASC LIMIT 10(no results)
- compute
SELECT str_nm, COUNT(DISTINCT ref_id) AS n_products, AVG(prc) AS avg_price FROM items_g1 WHERE str_nm IS NOT NULL GROUP BY str_nm HAVING COUNT(DISTINCT ref_id) >= 50 ORDER BY avg_price DESC, str_nm ASC LIMIT 10str_nm | n_products | avg_price HP | 104 | 374.79033898305084 Canon | 66 | 227.96085714285715 Amazon Renewed | 66 | 187.9235294117647 Brother | 66 | 137.94026315789472 Epson | 65 | 123.12 … 5 more lines
- compute
SELECT str_nm, COUNT(DISTINCT ref_id) AS n_products, AVG(prc) AS avg_price FROM items_g2 WHERE str_nm IS NOT NULL GROUP BY str_nm HAVING COUNT(DISTINCT ref_id) >= 50 ORDER BY avg_price DESC, str_nm ASC LIMIT 10str_nm | n_products | avg_price Canon | 54 | 67578.72222222222 Lenovo | 108 | 66982.97674418605 ASUS | 82 | 49604.0 HP | 206 | 44155.68292682927 Dell | 73 | 42321.31428571429 … 5 more lines
After that question the teacher wrote this procedure (verbatim):
Musical-instrument brand ranking (brands with at least N distinct products, highest average list price, alphabetical tie-break). The items_g3.str_nm brand column gives no brand with N or more products, so find where brand is recorded (the g3 key-value attribute tables) before computing or answering. Do not switch to other categories' tables or answer with an uncomputed guess.
The recipe built 6 training examples that query the migration's new attribute table, 522 supervised tokens in total. Nine questions later run 0 was asked how many instrument products have brand information in the current attribute feed but not in the old export. After a table listing and one exploratory count it read the instrument schema, then counted the rows with the key Brand in the new attribute table: 1145, correct, in 4 queries.
Learner, run 0, question 33
- list
.tablesattrs_g1 attrs_g3_legacy fdbk_g1 … 10 more lines
- compute
SELECT COUNT(*) AS n FROM items_g3 WHERE category LIKE '%Musical%' OR category LIKE '%Instrument%' LIMIT 1ERROR: no such column: category
- read schema
SELECT name, sql FROM sqlite_master WHERE type='table' AND name LIKE '%g3%';name | sql items_g3 | CREATE TABLE items_g3 ( ref_id TEXT PRIMARY KEY, ttl TEXT, avg_rtg REAL, rtg_ct … fdbk_g3 | CREATE TABLE fdbk_g3 ( id INTEGER PRIMARY KEY AUTOINCREMENT, ref_id TEXT, uid TEXT, … … 2 more lines
- compute
SELECT COUNT(*) AS n FROM product_attributes_g3 WHERE attr_key = 'Brand' AND ref_id NOT IN (SELECT ref_id FROM attrs_g3_legacy WHERE attr_key = 'Brand')n 1145
Run 0 used that table in 5 later questions and answered all 5 correctly. Before question 24 it had not written the table's name in any query.
5. Learner 1.0 learns efficiently online from very few examples
A run trains once after each question, on examples built from that question alone. Over forty questions the three runs trained on 539, 428 and 240 examples, 45,096, 37,579 and 20,547 supervised tokens: on average 13.5, 10.7 and 6.0 examples and 1,127, 939 and 514 supervised tokens per question. The table lists names each run did not write in any query until a training example used them. Each name was already visible in the schema listings the learner had read. What the examples taught was to use it.
We never train on a question paired with its answer. Every training example is a step of the learner's own attempt at a question: the question, the queries it sent so far and the database's real results, with the next action as the target. Most examples, 989 of the 1,207 across the three runs, teach the next query. The other 218 end in submitting an answer, and that value follows from the query results shown in the same example. Training happens only after a question has been answered, and each of the forty questions is asked once. Every training example is in the replication repository.
| run | name | what it is | taught after question | examples | supervised tokens | later questions using it | correct |
|---|---|---|---|---|---|---|---|
| run 0 | review_year | year column added to musical-instrument reviews by the migration | 22 | 3 | 543 | 2 | 2 |
| run 0 | attrs_g3_legacy | the old musical-instrument attribute table, renamed by the migration | 24 | 6 | 522 | 4 | 4 |
| run 0 | product_attributes_g3 | musical-instrument attribute table added by the migration | 24 | 6 | 522 | 5 | 5 |
| run 1 | cat_lvl | taxonomy level column | 7 | 7 | 1,519 | 2 | 1 |
| run 1 | cat_nm | taxonomy category name column | 7 | 7 | 1,519 | 2 | 1 |
| run 1 | prc_usd | electronics price in dollars | 12 | 5 | 1,300 | 6 | 2 |
| run 1 | attrs_g1 | office product attribute table | 14 | 2 | 254 | 2 | 1 |
| run 1 | verified_status | office review verification label added by the migration | 23 | 1 | 104 | 2 | 2 |
| run 1 | review_year | year column added to musical-instrument reviews by the migration | 24 | 2 | 440 | 2 | 1 |
| run 1 | attrs_g3_legacy | the old musical-instrument attribute table, renamed by the migration | 25 | 4 | 752 | 1 | 1 |
| run 1 | product_attributes_g3 | musical-instrument attribute table added by the migration | 25 | 4 | 752 | 3 | 2 |
| run 2 | taxn_g1 | office product taxonomy table | 2 | 5 | 180 | 4 | 2 |
| run 2 | fdbk_g2 | electronics review table | 2 | 5 | 180 | 9 | 6 |
| run 2 | fdbk_g3 | musical-instrument review table | 2 | 5 | 180 | 7 | 3 |
| run 2 | item_id | second product identifier in office reviews | 4 | 5 | 1,320 | 1 | 1 |
| run 2 | cat_lvl | taxonomy level column | 4 | 5 | 1,320 | 2 | 0 |
| run 2 | cat_nm | taxonomy category name column | 4 | 5 | 1,320 | 2 | 0 |
Learning the rule, not memorizing the examples. Each training example is used for one update and never again. Its loss is recorded in that same step, before the update, so every point below is the model's error on an example it has not trained on yet. A falling line means the model gets better at examples it has never trained on: it is learning how to act on this database, not memorizing the examples it was shown. The examples written after question 30 come from question 30, which the model has not trained on, so a lower loss there is carried over from earlier questions. From the first ten questions to the last ten, the mean loss on a new example falls from 0.27 to 0.11, from 0.23 to 0.10 and from 0.30 to 0.13 nats per target token in runs 0, 1 and 2.
Each lesson takes only a handful of examples. After a question the teacher usually writes several examples for each procedure it teaches. The first is often the step the learner already took, so its loss is low. The second is the first that teaches something new, with a median loss of 0.37 nats per target token. By the eighth example the median is 0.07, all within the training that follows that one question.
6. After the migration
At question 21 the database is migrated: the musical-instrument attribute table is renamed and replaced by a new one, musical-instrument timestamps are split into year and month columns, office reviews gain a partially filled verification label, and electronics gain a second price column. The paper that introduces the benchmark (Asawa et al., 2026, arXiv:2606.05661) illustrates the failure this is meant to catch: an agent that reuses what it learned before the migration and queries tables that no longer exist.
In the sixty post-migration questions of the three runs, 46 read the schema in their first two queries, and 0 queries used the renamed table's old name.
The three runs answered 9, 7 and 8 of the 20 questions before the migration and 12, 8 and 8 of the 20 after it. The stateless baseline answered 4 and 3. No run did worse after the database changed. By product group, office questions went from 1 to 3 of 8 correct before the migration to 3 to 5 of 6 after it, and musical instruments from 1 to 2 of 3 to 3 to 5 of 7. Electronics went the other way, from 4 to 5 of 8 to 1 to 2 of 7, against 2 of 7 for the baseline. Electronics is the group to which the migration added a second price column.
7. Results
| lane | correct of 40 | reward (of 40) | gain over baseline | queries | supervised tokens trained | vs baseline: gained / lost | exact McNemar p |
|---|---|---|---|---|---|---|---|
| Run 0 | 21 | 12.13 | +10.27 | 258 | 45,096 | 15 / 1 | 0.001 |
| Run 1 | 15 | 7.27 | +5.40 | 248 | 37,579 | 10 / 2 | 0.039 |
| Run 2 | 16 | 10.67 | +8.80 | 187 | 20,547 | 12 / 3 | 0.035 |
| mean of the three runs | 17.33 ± 1.86 | 10.02 | +8.16 | 231 | 34,407 | ||
| Stateless baseline | 7 | 1.87 | 0 | 479 | 0 |
What thinking adds without learning. The baseline uses the same thinking-when-stuck rule, so the gain above is what training adds on top of thinking. Thinking alone moves little: the baseline answers 7 questions with a reward of 1.87, against 6 and 1.40 for the same model with thinking off (the baseline of the earlier, no-thinking recipe). The baseline with thinking answers 2 questions that the no-thinking baseline misses, and misses 1 that the no-thinking baseline answers.
"Gained" counts questions a run answers correctly and the baseline gets wrong. "lost" counts the reverse. Pooled over the three runs, the learner gains 37 questions and loses 6.
8. Against the published leaderboard
Learner 1.0 is the only system on this leaderboard that learns strictly through its weights. Its context is reset before every question, so no in-context learning is allowed between questions. Every other system carries its experience forward in its context, a notepad, a memory system or a coding agent's workspace.
The benchmark publishes results for twelve systems on this task, each pairing a frontier model with a way of carrying experience forward: in context, through a notepad, through a memory system, or through a coding agent's workspace. Our three runs average a reward of 10.02 and a gain of +8.16.
| system | runs | reward | reward rank | gain | normalized gain | gain rank |
|---|---|---|---|---|---|---|
| Claude Code · Sonnet 4.6 | 5 | 22.05 | 1 | +13.85 | 0.436 | 1 |
| Mem0 · GPT-5.4 | 5 | 17.24 | 2 | +12.91 | 0.362 | 2 |
| ICL · Claude Opus 4.7 | 5 | 15.65 | 3 | +9.59 | 0.283 | 4 |
| ICL · Gemini 3 Flash | 5 | 15.03 | 4 | +11.49 | 0.315 | 3 |
| ICL · Claude Sonnet 4.6 | 5 | 15.01 | 5 | +8.48 | 0.253 | 5 |
| ICL · GPT-5.4 | 5 | 13.88 | 6 | +8.35 | 0.242 | 6 |
| ICL Notepad · GPT-5.4 | 5 | 12.37 | 7 | +6.37 | 0.187 | 9 |
| ICL · Gemini 3.1 Pro Preview | 5 | 11.56 | 8 | +6.83 | 0.194 | 8 |
| ICL Notepad · Claude Sonnet 4.6 | 5 | 11.00 | 9 | +3.67 | 0.112 | 12 |
| Learner 1.0 (ours) | 3 | 10.02 | 10 | +8.16 | 0.214 | 7 |
| Codex · GPT-5.4 | 1 | 9.60 | 11 | +6.13 | 0.168 | 10 |
| ICL Notepad · Gemini 3.1 Pro Preview | 5 | 8.52 | 12 | +4.32 | 0.121 | 11 |
| ACE · GPT-5.4 | 5 | 7.85 | 13 | +2.39 | 0.069 | 13 |
On reward we rank 10 of 13. Gain is ranked as the leaderboard ranks it, normalized by each system's headroom: gain divided by 40 minus the system's own baseline, because a system whose baseline is already high has less room to gain. On that measure we rank 7 of 13. The same convention is used for the sales task.
9. Against the earlier recipe
The earlier version of this recipe (full write-up) was identical except that the learner never thought before acting. It ran on the same three question orders. The final recipe answers 52 questions correctly over the three orders against 43, and repeats an identical query 34 times against 184. Its total reward is 30.07 against 30.00: the same.
The split by whether the learner got stuck shows why the two totals match. A question counts as stuck if at any step the stuck rule of section 2 would have fired. For the earlier recipe the rule is applied afterwards to its recorded trajectories.
| questions, pooled over the three orders | final recipe: questions | correct | reward | earlier recipe: questions | correct | reward |
|---|---|---|---|---|---|---|
| the learner got stuck at some step | 55 | 25 | 8.80 | 58 | 12 | 4.80 |
| the learner never got stuck | 65 | 27 | 21.27 | 62 | 31 | 25.20 |
Thinking happens only when the learner is stuck, and there it helps: the final recipe answers 25 of 55 such questions against 12 of 58. Those answers come late, after many queries, so each earns little reward. On questions where the learner never got stuck no thinking happened, yet the final recipe answers 27 of 65 against 31 of 62. That difference cannot come from thinking. It comes from what each run's training taught. Run 1 accounts for it: on its never-stuck questions it answers 5 of 20 against 11 of 25 for the earlier recipe on the same order, while runs 0 and 2 answer 12 of 21 and 10 of 24 against 9 of 18 and 11 of 19. Runs 0 and 1 open most later questions with a table listing before the schema read (24 and 26 of questions 12–40), where the earlier runs opened with the schema read directly (3 and 0). Run 2 opens with the schema read. The extra listing costs one query per question, but it does not explain run 1's shortfall by itself, because run 0 opens the same way and does better than before. We do not know what in run 1's training caused it.
10. The recipe
After each question. The teacher receives the learner's trajectory on the question just answered, the benchmark's released feedback for it, the learner's trajectories on earlier questions as evidence, and the procedures it wrote before. It returns procedures in a fixed format: what the question needed, the states that lead to an answer, and the exact next query or answer at each state.
From procedures to examples. The recipe builds each example from a state the learner actually reached: the conversation up to that point, with the recorded query results, and the next action as the target. It keeps only examples grounded in recorded results, adds three fixed examples about reading the database's state, and caps any single procedure at a quarter of a question's examples. Every example is trained once.
Acting. Greedy decoding. Thinking before an action only when stuck, as defined in section 2: at most 512 tokens of reasoning, then at most 384 tokens for the action. The learner's context is 32,000 tokens.
11. Interpretation, and how to check it
Training on each answered question changed how the learner explores this database, and thinking when stuck recovered questions the earlier recipe lost. The benchmark score does not show the second effect, because those answers took many queries and because one run regressed on questions where it never got stuck.
To check: every trajectory of every lane, the feedback, the training examples and the scores are in the replication repository, and the figures regenerate from them.