Learner LabsLearner 1.0Foundation models that continually learn

Demonstrations

Facts, Recipe 2

We tried a second way of preparing and teaching facts, which we call Recipe 2. This page reruns the fact demonstrations with one fixed version of it, side by side with the recorded results of Recipe 1, and adds two measurements Recipe 1 does not have: questions that need facts from two or more separately taught lessons, and a document that was never used while the recipe was being developed. The recorded sessions on the facts page are unchanged and remain the record of Recipe 1. This update covers facts only: the skills demonstrations are unchanged, and thinking mode is not covered here.

Conclusions

  1. Recipe 2 answers more questions about a taught document. On the handbook it answers 15 of 16 against Recipe 1's 13 of 16. → Side by side
  2. Earlier lessons do not move. Across four lessons taught in sequence, no question that a lesson answered correctly at its own teach is lost at any later state, on the published quiz or on a 108-question panel read at every state. → Sequential retention
  3. A new learner takes a contradicting lesson as well as Recipe 1 does. Taught the Veyra facts as its only lesson, it answers 6 of the 8 facts the base model does not already hold, the same as Recipe 1, and all three Earth controls stay correct. → Side by side
  4. Questions that need several lessons at once fail when asked directly and partly succeed with a multi-call procedure. Letting the learner split the question, recall each part in its own call and combine its own recalled answers raises the result from 1 to 10 of 16 on a panel read once, at up to 8 calls per question. Decisions are the weak part. → Cross-lesson questions

What changed between the two recipes

The clearest change is the teaching material made from a document, and how many times it is shown. Recipe 2 also changes training settings that we do not publish, as with everything about the method.

Recipe 1 (the recorded sessions)Recipe 2 (this page)
what is taughtshort fact statements read out of the document by a rule pass and a language-model pass, each statement written in 12 phrasingseach fact expanded into 80 practice questions with the fact's exact value as the answer: 40 that ask for it directly, 16 that ask through a situation, 16 that reach it through another taught fact, and 8 that ask by category or description
how the material is writtenthe extraction instruction printed on the facts pagewritten by a language model, a few facts at a time, then checked question by question by a separate language-model reviewer. Anything it rejects is rewritten and re-checked, up to three rounds, and nothing is trained until every question has passed
how ofteneach fact written in 12 phrasings across 4 passes. One training pass over the document is the product default2 passes over the practice questions for a document, so 160 updates per fact
objectiveplain next-token predictionplain next-token prediction on the answer to each practice question, one question per update
not usedno reinforcement learning, no reward model, no preference tuning in either recipe. In the Recipe 2 sequence, each lesson is taught only on its own practice questions. Nothing from an earlier lesson is shown again.
where it ranthe deployed productour research training harness, one GPU per run

The results below compare the two recipes as wholes.

Side by side, demonstration by demonstration

Each row uses the same questions, the same grading rule and the same denominator for both recipes.

demonstrationRecipe 1Recipe 2notes
Teach a document: the 706-word handbook, 16 askings13/16 (policy 8/8, relationship 5/8), greedy15/16 (policy 7/8, relationship 8/8). Base model 0/16Four of the 16 questions appear word for word among Recipe 2's training questions. It answers all four. The other 12 are 11 of 12.
Teach a document: the handbook, 16 askings, sampled at temperature 0.913/16 (policy 7/8, relationship 6/8)not runOur research harness answers greedily only.
Teach a document: a new 866-word document, 16 askingsnot run14/16. Reworded questions 7/8. Base model 0/16Written for this test. The bar was the handbook's 13/16, and it is met. → An unused document
Teach in sequence: three topics, the published 4-question quizA 4, 4, 4 · B 4, 4 · C 3A 4, 4, 4, 4 · B 4, 4, 4 · C 4, 4Under a stricter whole-word rule both recipes read 4/3/3 at the end. The published grader counts the plurals "cinders" and "marks" as matches. Recipe 2 taught a fourth lesson, so it has one more reading of each earlier lesson. → Sequential retention
Override a belief: nine Veyra facts taught as a learner's only lesson6 of the 8 facts the base does not already hold, strict (7 of 8 lenient, the question-by-question page counts 7 of all 9)6 of the same 8, strict. Earth controls 3/3A new learner, taught nothing else, for 320 updates.
Does teaching erode base capabilities?no measurable change, 16,481 items150 items, paired with the base model: the same accuracy on every task (MMLU 90/100, ARC-Challenge 18/25, WinoGrande 21/25)148 of the 150 items got the same answer as the base model. On WinoGrande one was gained and one lost.
What a direct answer saves in prompt tokens (derived, no new run)median 52 times fewer than retrieval and 105 times fewer than putting the whole document in the prompt (68 asks)median 42 and 66 times fewer (32 asks: the handbook and the new document)A direct answer sends the question alone in both recipes, and the same retrieval and whole-document baselines are used. The ratios differ because the documents and questions differ. The multi-call procedure sends several prompts per question and is not in these figures.

Sequential retention: every lesson, after every later lesson

The claim that matters for incremental learning is not the final score. It is whether any later lesson moves an earlier one. So every lesson is read when it is taught and again after each later lesson.

Setup

Four lessons were taught into one learner, one after another: the Kestrel board (A), the Ondine protocol (B), Tallow billing (C), and a fourth lesson (D). The first three are the topics of the recorded teach-in-sequence session. Lesson A was taught in two passes over its practice questions, 640 updates. Lessons B, C and D were taught for 160, 320 and 320 updates. Lessons B, C and D therefore received fewer updates per fact than lesson A or a taught document. Each lesson is taught only on its own practice questions. Nothing from an earlier lesson is shown again.

The published quiz: four questions per earlier lesson

lessonafter Aafter Bafter Cafter D
A, Kestrel4/44/44/44/4
B, Ondine4/44/44/4
C, Tallow4/44/4

The bold figure in each row is the reading at the lesson's own teach. The reading of lesson A after A was taken from the saved state. These counts use the published grader. Under a stricter whole-word rule the end state reads A 4, B 3, C 3, because the published grader counts plurals such as "cinders" and "marks" as matches. That difference is present at every state and is the same for Recipe 1's recorded session.

A larger panel: 36 questions per lesson, read at every state

A panel of 108 questions, 36 per lesson, was read after every teaching step. Counts are under the stricter whole-word rule.

stateABC
after A3510
after B35251
after C352525
after D352525

No question that a lesson answered correctly at its own teach was answered wrongly at any later state: 0 lost for A, B and C, through all of lesson D. Within lesson C, the score moved from 27 after 160 updates to 25 after 320. That is inside lesson C's own teaching, not a later lesson moving it. Two further panels of 144 questions, written separately and read after lesson D only, scored 105 and 94.

Cross-lesson questions

Every demonstration so far asks one fact at a time. A person using what they learned combines it. Can the learner answer a question that needs facts from two, three or four lessons taught at different times?

Setup

The learner is the one from the sequence above, after all four lessons, with no further training. It was read on two panels of 16 questions. Each question needs two, three or four lessons. Some ask for several values, some compare, some calculate, and some ask for a decision from a stated rule. The lessons never mention each other, and no question asserts a relationship between them.

  • Development panel. Used while the answering procedure below was developed, on an earlier version of the recipe.
  • Fresh panel. Written by a separate author from the taught statements alone, without reading any model output, and sealed. It was read exactly once, for this version, under rules written down before the read. Before the read, while preparing files, we saw the wording of the first five of its questions, but no answer. We report those five separately.

A question is counted correct only if every value it asks for is stated correctly and, where it asks for a decision, the right decision is stated and the wrong one is not. There is no partial credit in the headline. The number of values right is shown beside it.

Two ways of answering

Direct. One call. The question is sent alone, exactly as in every other demonstration.

Multi-call. The same learner answers in steps, each a separate call: it writes its own list of sub-questions, each naming one fact. Each sub-question is sent on its own and answered from the weights, and a last call combines the learner's own recalled answers into the final answer. If the list joins several facts in one sub-question, the learner is asked once to split it again. For a comparison or a calculation, a small calculator computes the result from the recalled values and shows it in the last call. If the stated conclusion does not follow it, the learner is asked once to answer again. At most 8 calls per question. The final call is not answered from nothing: its prompt contains the answers the model itself recalled one step earlier. What never enters any prompt is the taught document, the expected answer, a list of which facts a question needs, or a plan written by us. Every fixed instruction is printed in the repository.

Results

panelway of answeringquestions fully correctvalues correctdecisions correctwrong values asserted
developmentdirect, one call0/164/461/528
developmentmulti-call13/1642/465/54
fresh, read oncedirect, one call1/165/452/939
fresh, read oncemulti-call10/1640/453/92

"Wrong values asserted" counts the wrong values in answers that say nowhere that a fact is unknown or missing. Counting every value not matched, stated or not, gives 41, 4, 39 and 4. On the fresh panel the multi-call answer recalls most values (40 of 45) and asserts far fewer wrong ones than the direct answer, but it gets only 3 of the 9 decisions right, and that is where most of its six failures are. Of the first five fresh questions, whose wording we had seen, it answered 4. Of the other eleven, 6.

What this shows. Facts taught in separate lessons can be brought together by the learner's own recall, and answering directly cannot do it.

A document never used in development

A recipe that works on the handbook might only work on the handbook. So we wrote a second document for this test and kept it out of development.

Setup

An 866-word operating handbook for an invented ropeway cooperative, with 34 facts, none of which the base model can know. Its 16 questions were written and fixed before it was taught: 8 that use the document's own wording and 8 reworded. The bar, fixed in advance, was the handbook's 13 of 16 for Recipe 1. The learner was taught this document alone, with the same settings as the handbook.

base modelafter teaching
all 16 questions0/1614/16
8 reworded questions0/87/8
8 questions in the document's wording0/87/8

Of the 8 questions in the document's wording, one appears word for word among the training questions and two nearly so. The other 5 are 5 of 5. The base model's 16 answers are all wrong, so none of the 14 comes from what it already knew. All 16 base answers also ran past the length limit, but the facts are invented, so the base model cannot know them. 5,438 of 5,440 scheduled updates were applied. The other 2 were withheld by the training process and are recorded.

The data

Every question, every expected answer, every served answer, every sub-question and recalled answer, and the grader for each table will be in the data repository under recipe-2/. Every Recipe 2 number on this page comes from one fixed version of the recipe. The repository lists, for each table, the run it came from.

How these answers were produced. As on the facts page: no system prompt, no examples, no retrieval, greedy decoding, questions written before the runs. The one exception is the multi-call procedure for cross-lesson questions, which uses fixed instructions, printed in full in the repository's recipe-2/composition/INSTRUCTIONS.md, with the learner's own recalled answers in the last step.