Learner LabsLearner 1.0Foundation models that continually learn
~3.5 billion years ago first self-copying cells ~1.6 billion complex cells that hunt ~600 million the first nerve net ~325 million flight — the sky's first hunter ~25 million the first true cats ~7 million our line splits from chimps at least ~300,000 our species appears · stone tools 2026 sequential learning with measured retention
Learner Labs · introducing Learner 1.0

What will you build with an AI that can continually learn?

Evolution found a learning algorithm built on biological neurons that can continually learn without catastrophic forgetting. Learner 1.0, a ~28B dense foundation model, achieves the same on silicon. We believe this solves a ~40-year-old problem of online continual learning in the field. Sign up for early API access and read our demo reports.

A continual parametric learning agent that learns from its own experience on Continual Learning Bench and retains what it learns

0 0.1 0.2 0.3 0.4 0.5 1 5 10 15 20 25 30 35 40 question (position in the run) running average reward database migration run 0 0.303 run 1 0.182 run 2 0.267 learning off 0.047 Database exploration, 40 questions. Reward per question = 1 − queries/15 if the answer is correct, else 0. The running average is over the questions answered so far.

Agents today cannot update their weights online on incrementally new data due to catastrophic forgetting. The graph shows 3 replicates of an agent built on the Learner 1.0 model that learns from its own experience, with only a handful of examples, and improves its performance on Continual Learning Bench. On average the three runs trained on 13.5, 10.7 and 6.0 examples and 1,127, 939 and 514 supervised tokens per question.

See all demonstrations · Read the database exploration report

Motivation

A lot has been said about the need for models to be able to continually adapt to your context. If you haven't already, I would like you to experience it. If you use AI regularly and spend 10–20M tokens a day as a user, you can experience this need right now on your own laptop. As you use AI for your project or on your codebase, I need you to read the traces and count every time you see the model have a repeat “Aha!” moment. You will notice that to solve any task, today's model must rediscover all the rules and relationships in every fresh 1M context window, and spends tokens and time relearning these relationships, 10s or 100s of times every day.

The obvious benefit of a model that can continually adapt is that the time and token cost spent on relearning will dramatically decrease (read more here: our note on what re-reading costs). As the model becomes skilled in your data and context via internalized weight updates, you should be able to accomplish a lot more with a lot fewer tokens. The non-obvious benefit, and the bigger impact I believe, might be in your ability to tackle much more difficult challenges. The kind of challenges that often need very long horizon work spanning weeks, months or even years, and which may need some form of compounding that In Context Learning can simply not provide.

Our mission at Learner Labs is to help you tackle such goals and play a small yet meaningful role in this journey. With this, we want to introduce you to Learner 1.0, the world's first continually adapting and learning foundation model. We will be providing limited access to our API in the coming weeks as we work to ensure productive and safe release of this new capability. Unlike the usual model APIs, you will be able to use it to continually train it on your own data, with model instance and weights that are yours. More details towards the end.

A different prior

If you are with us till now, you are convinced that continual learning is very important. So, why wasn't this solved already? I have a theory. I believe it might be a matter of a different prior. As an engineer who has wrestled with biology for the last decade, understanding cell systems and memory within them, when I began working on deep learning I was surprised by the fact that for a field whose entire goal is to mimic a biological brain, very few discussions happen about appreciating the rules of biology, dissecting them and maybe questioning which abstractions can be coded up into architectural primitives. I also find an obsession with neurons that might be misleading and hiding the bigger picture. Let me give you a concrete example. Did you know you can convert any cell, for instance a neuron, into another cell type by just changing 4-5 molecules (eg Yamanaka factors, nobel prize 2012)? Visualize the JSON file for a cell - it has millions of entities, and yet all it took was 4-5 changes to change the entire cell identity and function. I think the commonalities between different cell types are more interesting than their differences.

It also may be that deep learning as a field is just progressing through its stages starting with proving universal function approximation and being able to learn deep representations, and then only recently earning the right to think deeply about generalization, learning efficiency and adaptation.

As a consequence, I believe most prior work so far are post hoc attempts to prevent forgetting in an architecture that was not designed for continual adaptation. Learner 1.0 is only possible due to new biologically grounded architectural and algorithmic primitives that we believe do for continual adaptation of neural nets what attention does for long sequence modeling.

Defining the problem precisely

Before we proceed further, let's define what continual learning of a model would need to do much more precisely so we have the right ruler to measure success, and I am talking about the kind where a single model instance's weights update based on a stream of data. Below are criteria, I believe must be simultaneously met to claim the problem solved.

  • Task label free. Real user data has no defined boundaries. The user can switch tasks and domains at any time so the model must be able to work completely label free.
  • Incremental updates. One example, one update, in the order the data arrives. In conventional model training pipeline, large scale batching with different data distributions interleaved help prevent forgetting during training. In contrast, to prove the point, our experiments train 1 example at a time and 5000-10000+ of examples of a single domain, and then do this sequentially for 10 domains.
  • Replay free. The cost of incremental training after a year's worth of data must remain low. If old data has to be shown again to retain it in the model, the cost grows with everything you have ever learned. Separate from that, I dont think you can even replay the training data for the pretrained knowledge while being economical.
  • Fixed capacity. The number of trainable parameters cannot grow with every new thing that is taught. The power of deep learning is in shared representations. There is an obvious upper bound governed by how many bits you can compress per trainable parameter in any architecture.
  • It must actually learn. It is easy to not forget if you do not learn much. Each new thing has to be acquired about as well as a strong weight-update method would acquire it.
  • It must retain. Earlier things have to survive later teachings, and we have to go back and measure them again after every later stage, not once at the end.
  • It must preserve the base. If starting with a pretrained base, the base capability must not degrade when learning new distributions. Most of the value is in what the pretrained model already knows. Learning your data cannot cost you that.
  • It should use pretrained priors efficiently when learning new things. This is a practical constraint. A learner that starts from zero on every new topic wastes what the base already knows. As you will see in our results, this should be observable as data efficiency, how many examples or optimizer steps did it take to get to the saturation loss (how quickly it learnt).
  • It must be scalable. Also a practical constraint. It has to have reasonable training and inference speed and costs.

Overview of our experimental setup and results

So how do we convince ourselves and everyone that this new architectural and algorithmic primitive in our Learner 1.0 solves online continual learning in a generalizable way? Eventual proof is wide deployment in the world. But, to start the conversations before we can safely and widely release this, we are doing 2 things:

  • We are releasing our own testing reports and demos, with what we believe are sufficient experimental details and methods to pressure test our evaluation strategy. The core mechanism for obvious reasons is not disclosed.
  • We will provide a limited release of our API, using which you can run your own tests on your own data or replicate our results. We are initially prioritizing researchers who can help us bring this capability safely to end users.

Our experimental setup

Our Learner 1.0 is an architecturally modified Qwen 3.6-27B model with a fixed number of trainable parameters added to it once, 1.14B or 2.28B depending on the experiment, and that number never grows no matter how much we teach. We also have internally tested the architecture in GPT family (GPT2), and Llama family (Llama 3.1 8B) and believe the architecture will scale to larger MoE style models relatively trivially provided we have the compute. Currently our training and inference are all run on a single gpu. Data always arrives as one stream. There are no optimizer resets between task distribution boundaries and the model is never told which task or domain it is looking at, or where one ends and the next begins. After every stage we go back and measure everything taught before it again, on held-out data in addition to doing comprehensive base evals. Where we compare, we compare against LoRA with a matched number of trainable parameters (eg LoRA rank 256), on the same data in the same order. In addition, there is no replay of any kind during training. In conclusion, all the criteria listed in previous section, are followed in the strictest sense. Further details to replicate our experiments using our API are provided in relevant reports. Github repo with full details of training data, evals and more is also shared.

Skills

In my understanding, the harshest way to pressure test a continually learning architecture is to incrementally train it with 1 example at a time (batch=1, 1 optimizer step per example), of a narrow task or domain, and do this for 1000s of examples of that 1 narrow domain sequentially, and then repeat this in sequence for many domains. In the skill experiments we do exactly this, with one pass over the data, and we never show an old example again.

  • Demo 1 : Ten skills, one after another. We taught one model ten unrelated skills in sequence with 63,282 single-example updates. For each skill the base model scores ~0 before training. After each new skill is trained, all previous skills are evaluated on a per skill 225 held out eval panel and a base eval panel is also measured at every checkpoint. In addition, comprehensive base evals are done at checkpoint 0, 5 and 10: a 16,481-item likelihood battery on stock lm-eval, 65 task-metric rows, covering MMLU across all 57 subjects, ARC-Challenge and WinoGrande. A 498-item sentinel covering four benchmarks: MMLU, ARC-Challenge, HellaSwag and WinoGrande (ARC 128, HellaSwag 128, MMLU 114, WinoGrande 128) is run at every single checkpoint. Every acquired skill is retained through all the sequential teaches and the base eval holds constant showing no loss of base capabilities. The one skill that moved by ~ 1% was the one the model never learned well in the first place. You can see the train loss, validation loss and accuracy details in the report. Read the ten-skill report
  • Demo 2 : Four domains, against LoRA. This was an earlier demonstration on a stream of four text domains. At the first evaluation, after 25 updates per domain, Learner 1.0 showed 82–125% of LoRA’s 300-update loss reduction, measured from each condition’s recorded base reference. As newer domains were trained previous related domains improved (Positive backward transfer). LoRA lost between 0.03 and 0.07 nats on every earlier domain. Learner 1.0 showed greater average loss reduction than LoRA while retaining earlier domains. Read the four-domain report
  • Demo 3: Two invented languages. We taught two made-up languages, 22,000 words each, back to back. Learner 1.0 kept the first after learning the second, with no words of one appearing in the other. A LoRA adapter given the same text in the same order lost the first language completely and got worse than an untrained base. Read the two-languages report

Benchmarks

The demonstrations above teach the model from data we prepared. A harder test is whether a model can learn from its own attempts at a task, the way a person gets better at a job by doing it. Continual Learning Bench is a public benchmark built for this. It tests whether a system gets better from a small number of attempts, learning online one task at a time, and whether it keeps what it learned as the task changes. Every system is compared with the same model answering the same questions with learning off. The difference is the benchmark’s learning metric. Learner 1.0 is the only system measured on this benchmark that learns by updating its weights. Every system on the public leaderboard carries its experience forward in its context, a notepad, a memory store or a coding agent’s workspace. We reset Learner 1.0’s context before every question, so anything it learned from earlier questions has to be carried in its weights. After each question, a teacher LLM reads the learner’s attempt and writes training examples from it, and Learner 1.0 trains on them one example at a time, once, with no replay and no optimizer resets.

  • Benchmark 1: Database exploration. Forty questions about one product database, answered in sequence, with the database migrated halfway through. Over three runs Learner 1.0 answered 21, 15 and 16 correctly, where the same model with learning off answered 7. On the leaderboard’s normalized gain it would place 7th of 13 systems, and it is the only one of the 13 that learns only through its weights. Read the database exploration report
  • Benchmark 2: Cohort studies. Twenty clinical studies answered in sequence, with training after each one. Over three runs Learner 1.0 was ahead of the same model with learning off in every run, with a mean normalized gain of +4.35%, and ahead on 12 of the 20 studies on average. On the leaderboard it would place 2nd of 13 systems by reward and 5th of 13 by normalized gain. Read the cohort studies report
  • Benchmark 3: Sales forecasting. Twelve consecutive years of product sales forecasts, with training after each year. Learner 1.0 scored 8.44 in total against 5.76 with learning off, and beat that baseline in 10 of the 11 years after the first. This is one run (n = 1). The database exploration and cohort studies results each use three runs. On the leaderboard’s own ranking, normalized reward, it would place 8th of 13 systems, and 10th of 13 by normalized gain. Read the sales forecasting report

See all Continual Learning Bench results

Facts

These experiments were run out of a curiosity that can the same architecture without any changes to the mechanism also store facts and recall them. These experiments should be treated as proof of capability. The results below use our second fact recipe, Recipe 2. I also believe that what facts you store into the weights is a very nuanced question with application specific answers. In these experiments, we try to override beliefs of the model by teaching it counterfactuals, made up company details or invented planet details and then ask it to recall and use the facts in its thinking. In Recipe 2, the facts are first extracted from a document by a rule pass and a language-model pass, turned into practice questions, and then stored in the weights over 2 passes. Besides the update into the weights, no other fact specific tensors or crutches are available to the model at inference. Then we simply ask questions to the trained model without any document, retrieval or system prompt. The model must intrinsically form its associations between the prompt and the facts and output the correct answer.

  • Demo 1: Teach a document. A 706-word fictional handbook about a company was taught, and the learner answered 15 of 16 questions about it with no document in the prompt. See the fact demonstrations
  • Demo 2: Teach in sequence. Four unrelated topics were taught into one learner, one after another. No question a lesson answered correctly at its own teach was answered wrongly at any later stage, on the published quiz or on a 108-question panel read after every lesson.
  • Demo 3: Override a belief. Taught facts about an invented planet that contradict what the base model believes, a new learner served back 6 of the 8 facts its base did not hold, and all three Earth control questions stayed correct.
  • No Base degradation. The same 16,481-item held-out benchmark battery, covering MMLU across all 57 subjects, ARC-Challenge and WinoGrande, scores the base capabilities before and after teaching, at every checkpoint. In some cases we also measured free generation, GSM8K in the model’s own reasoning mode, and that holds too. For Recipe 2 we ran a smaller check, 150 items paired with the base model, and every task scored the same as the base.

Results for an earlier variant of the fact recipe, Recipe 1, are also in the reports.

Every number above links to a full report with the method, the data and every recorded answer.

Our limited API release

Here is how it works. The weights are specific to each user (and training run). When you start a session you pick a checkpoint to load on the GPU, for serving or for training. Once training is complete, a new checkpoint is produced and saved to persistent storage, and the GPU is released. The overall experience is slow today due to limited GPU supply, and we plan to improve this as we add capacity. This is a research preview only setup with no batching of any kind between users. A single user always gets a dedicated GPU for their weights for inference and training, billed for the time they are using the GPU. The cost of training 1 Million tokens of skill data today is ~$6. We expect this to drop significantly with engineering optimizations.

The API we are releasing is in research preview only, as we must get the use cases and safety implications right before we can broadly make this accessible to everyone in the field.

We are prioritizing API access to researchers, with an initial focus on understanding the safety implications of this new architecture, and we would love to collaborate with individuals and organizations who can help us understand how we can safely launch this in the world.

Sign up for early API access

Looking forward

Learner 1.0 is initially launched on a dense model, but that doesn't limit its application to MoE style, much larger models, and we will work to launch a continually learning 1T+ parameter model in the near future.

On that same note, we hope this architecture will also help advance the field of robotics in the near future, where continual learning could potentially have the largest impact of any field.

AI is likely humanity's biggest accomplishment till date and our mission as a company is to contribute, even if it is only a small bit, to this collective effort.

Anurup Ganguli, Learner Labs

Learner Labs · Learner 1.0 · the world's first foundation model that continually learns