Learner LabsLearner 1.0Foundation models that continually learn

Demonstrations

We test two kinds of learning, and for both we measure on held-out data and publish every answer and full evals. Skills are procedures the model learns to carry out, and facts are statements it learns to recall and use it in its thinking process as well. We also run the model on a public continual-learning benchmark, where it learns from its own attempts at each task.

For why we ran these experiments and what we set out to prove, read the letter on the home page.

Benchmarks

Beyond teaching the model from data we prepare, the harder test is whether a model can learn from its own attempts at a task, the way a person gets better at a job by doing it. Continual Learning Bench is a public benchmark built for this. It tests whether a system gets better from a small number of attempts, learning online one task at a time, and whether it keeps what it learned as the task changes. Every system is compared with the same model answering the same questions with learning off. The difference is the benchmark’s learning metric. Learner 1.0 is the only system measured on this benchmark that learns by updating its weights. Every system on the public leaderboard carries its experience forward in its context, a notepad, a memory store or a coding agent’s workspace. We reset Learner 1.0’s context before every question, so anything it learned from earlier questions has to be carried in its weights. After each question, a teacher LLM reads the learner’s attempt and writes training examples from it, and Learner 1.0 trains on them one example at a time, once, with no replay and no optimizer resets.

Benchmark 1: Database exploration. Forty questions about one product database, answered in sequence, with the database migrated halfway through. Over three runs Learner 1.0 answered 21, 15 and 16 correctly, where the same model with learning off answered 7. On the leaderboard’s normalized gain it would place 7th of 13 systems, and it is the only one of the 13 that learns only through its weights. Read the database exploration report

The earlier recipe

Benchmark 2: Cohort studies. Twenty clinical studies answered in sequence, with training after each one. Over three runs Learner 1.0 was ahead of the same model with learning off in every run, with a mean normalized gain of +4.35%, and ahead on 12 of the 20 studies on average. On the leaderboard it would place 2nd of 13 systems by reward and 5th of 13 by normalized gain. Read the cohort studies report

Benchmark 3: Sales forecasting. Twelve consecutive years of product sales forecasts, with training after each year. Learner 1.0 scored 8.44 in total against 5.76 with learning off, and beat that baseline in 10 of the 11 years after the first. This is one run (n = 1). The database exploration and cohort studies results each use three runs. On the leaderboard’s own ranking, normalized reward, it would place 8th of 13 systems, and 10th of 13 by normalized gain. Read the sales forecasting report

Skills

Demo 1 : Ten skills, one after another. We taught one model ten unrelated skills in sequence with 63,282 single-example updates. For each skill the base model scores ~0 before training. After each new skill is trained, all previous skills are evaluated on a per skill 225 held out eval panel and a base eval panel is also measured at every checkpoint. In addition, comprehensive base evals are done at checkpoint 0, 5 and 10: a 16,481-item likelihood battery on stock lm-eval, 65 task-metric rows, covering MMLU across all 57 subjects, ARC-Challenge and WinoGrande. A 498-item sentinel covering four benchmarks: MMLU, ARC-Challenge, HellaSwag and WinoGrande (ARC 128, HellaSwag 128, MMLU 114, WinoGrande 128) is run at every single checkpoint. Every acquired skill is retained through all the sequential teaches and the base eval holds constant showing no loss of base capabilities. The one skill that moved by ~ 1% was the one the model never learned well in the first place. You can see the train loss, validation loss and accuracy details in the report. Read the ten-skill report

Demo 2 : Four domains, against LoRA. This was an earlier demonstration on a stream of four text domains. At the first evaluation, after 25 updates per domain, Learner 1.0 showed 82–125% of LoRA’s 300-update loss reduction, measured from each condition’s recorded base reference. As newer domains were trained previous related domains improved (Positive backward transfer). LoRA lost between 0.03 and 0.07 nats on every earlier domain. Learner 1.0 showed greater average loss reduction than LoRA while retaining earlier domains. Read the four-domain report

Demo 3: Two invented languages. We taught two made-up languages, 22,000 words each, back to back. Learner 1.0 kept the first after learning the second, with no words of one appearing in the other. A LoRA adapter given the same text in the same order lost the first language completely and got worse than an untrained base. Read the two-languages report

Facts

All five fact demonstrations are on one page, and it is best read from top to bottom. See the fact demonstrations

Demo 1: Teach a document. A 706-word fictional handbook about a company was taught, and the learner answered 15 of 16 questions about it with no document in the prompt.

Demo 2: Teach in sequence. Four unrelated topics were taught into one learner, one after another. No question a lesson answered correctly at its own teach was answered wrongly at any later stage, on the published quiz or on a 108-question panel read after every lesson.

Demo 3: Override a belief. Taught facts about an invented planet that contradict what the base model believes, a new learner served back 6 of the 8 facts its base did not hold, and all three Earth control questions stayed correct.

Results for an earlier variant of the fact recipe, Recipe 1, are also in the reports.

No Base degradation. The same 16,481-item held-out benchmark battery, covering MMLU across all 57 subjects, ARC-Challenge and WinoGrande, scores the base capabilities before and after teaching, at every checkpoint. In some cases we also measured free generation, GSM8K in the model’s own reasoning mode, and that holds too. For Recipe 2 we ran a smaller check, 150 items paired with the base model, and every task scored the same as the base.