Training Data
Training Data

Memory and Continual Learning: Engram's Dan Biderman and Jessy Lin

June 24, 2026

AI Summary

5 min read

In the early days of large language models, a single Wikipedia article about Taylor Swift could balloon into 80 gigabytes of GPU memory just to hold its key-value cache—roughly the same footprint as the entire 70-billion-parameter Llama model that memorized the internet. That absurd asymmetry is the kind of problem that Dan Biderman and Jessy Lin, co-founders of the new lab Engram, are building toward. Their premise: the bottleneck for making AI more useful is no longer raw intelligence, but the ability to learn new, private, and evolving context deeply into model weights—not just retrieve it at test time.

Always training vs. context engineering

The standard approach today is context engineering: stuffing a huge prompt with documents, conversation history, and system instructions so the model can figure out what to do. Engram calls this "externalized memory"—like sticky notes or a search engine. It works, but it has limits. As users collectively generate tens of millions of tokens per day, the cost of reading and searching through all that context grows. More fundamentally, the model never really knows anything about your team, your priorities, or your way of doing things the way a long-time employee does. It has to rediscover everything from scratch on every query.

Continue reading the full summary in the app — free to try.

Read Full Summary →

Free • No credit card required

What you'll learn

  • 1 (01:09) **Always-Training Models vs. Context Engineering** - Dan and Jessy define the core premise: the bottleneck is learning new, evolving context, not raw intelligence.
  • 2 (03:29) **The Engram Architecture: Training Team-Level Models** - Dan describes how Engram turns raw workspace data into a model that deeply understands a team's context.
  • 3 (05:45) **Internalized vs. Externalized Knowledge** - The key challenge is deciding what to learn into weights and what to leave as external tools.
  • 4 (08:23) **Workspace vs. Individual Level Training** - They explain their strategic choice to start with teams and organizations.
  • 5 (08:57) **Memorization vs. Learning: A False Dichotomy** - The guests argue that separating "fact memorization" from "skill learning" is unnatural and often misguided.
  • 6 (13:21) **The "Dream" Analogy for Model Improvement** - Dan uses the concept of dreams to explain how models can experiment with their own affordances.
  • 7 (14:19) **Canonical Example: The Math Olympiad** - A motivating hypothetical that shows the power of training over pure context engineering.

+ Full timestamped outline available in the app

Show Notes

Dan Biderman and Jessy Lin, co-founders of Engram, are building a neolab around memory and continual learning, which they call two sides of the same coin. Their contrarian premise: instead of stuffing ever-larger prompts into the context window or bolting on RAG, bake a team's knowledge directly into the model's weights, so it knows your company the way an employee of several years does. 

The payoff: matching or beating frontier models while consuming up to 100x fewer tokens. Working with partners like Microsoft, Notion, and Harvey, the team draws on roots in computational neuroscience and state-space architectures to attack what they see as the real bottleneck in AI — not raw intelligence, but memory and continual learning. In contrast to the frontier labs' race toward one ever-bigger model and AGI, Dan and Jessy imagine a world where everyone has their own model — privately trained, always learning, and good at the things you actually care about. The real ChatGPT moment for memory, they argue, is the day your model feels like an intern that genuinely got smarter overnight.

Hosted by Sonya Huang and Shaun Maguire, Sequoia Capital



Training Data

More from this podcast

Training Data →