Research

AI Agent Memory Explained: Types, Architectures and What Actually Works

AI agent memory explained: working, episodic, semantic and procedural memory, the main architectures (MemGPT, Generative Agents, Mem0), why bigger context windows are not memory, and how to design it.

Contents

Floor-to-ceiling shelves of old books in a bookshop

Image: Iñaki del Olmo on Unsplash.

A language model remembers nothing between calls. Every time an agent runs, the model sees only what is in its context window at that moment. "Memory" is the engineering that decides what goes into that window, what is kept outside it, and what gets pulled back in when it matters.

Get it wrong and an agent either forgets the correction you made yesterday, or drags so much history into every call that it gets slower, more expensive and, measurably, worse. This guide covers the types of memory, the main research architectures, the evidence on what works, and a practical design.

The four types of agent memory

The most widely used taxonomy comes from CoALA (Cognitive Architectures for Language Agents), a 2023 framework from Princeton researchers that borrows from decades of cognitive-science work. It gives a language agent "modular memory components": short-term working memory plus several kinds of long-term memory (Sumers et al., CoALA).

Type What it holds Agent example
Working What the agent is using right now The current task, recent tool results, the plan
Episodic Experiences from earlier runs "Last Tuesday's report failed because the sales API timed out"
Semantic Facts about the world and the user "The company uses Postgres, not MySQL"
Procedural How to do things Instructions, skills, tool definitions, code

CoALA describes episodic memory as storing "experience from earlier decision cycles," which can be retrieved into working memory during planning, and semantic memory as storing "an agent's knowledge about the world and itself." Most production systems implement semantic and episodic memory explicitly, and treat procedural memory as the agent's prompts, skills and code.

Why a bigger context window is not memory

Context windows now reach millions of tokens, so it is tempting to put everything in. The evidence says don't.

Chroma evaluated 18 models and found that "model performance consistently degrades with increasing input length," even on deliberately simple tasks. On LongMemEval, a conversational-memory benchmark, models scored significantly higher when given a focused prompt of about 300 tokens containing only the relevant history than when given the full ~113,000-token conversation that also contained it (Chroma, *Context Rot*). The full version forces the model to do retrieval and reasoning at once, and it does both worse.

Distractors make it worse: "Even a single distractor reduces performance relative to the baseline." Replaying all history into every call brings in exactly those near-miss distractors.

Memory is selection, not storage. The job is to put the right 300 tokens in front of the model.

Three architectures worth knowing

MemGPT: memory as an operating system

MemGPT (UC Berkeley, 2023) borrows the idea of virtual memory from operating systems: a small, fast tier (the context window) and larger, slower tiers outside it, with the model itself moving information between them. This lets an agent analyse documents far larger than its window and sustain multi-session conversations that "remember, reflect, and evolve" over time (Packer et al., MemGPT).

Lesson: give the agent explicit tools to read and write its own memory, rather than hoping the right text lands in the prompt.

Generative Agents: a memory stream with reflection

Stanford's Generative Agents paper (2023) populated a small town with 25 agents that plan their days, remember interactions and coordinate, such as spreading invitations to a party. Their architecture stores "a complete record of the agent's experiences using natural language," synthesises those memories "over time into higher-level reflections," and retrieves them dynamically to plan. Ablations showed observation, planning and reflection "each contribute critically" (Park et al., Generative Agents).

Lesson: raw logs are not enough. Periodically distil episodes into short, general lessons.

Mem0: extract, consolidate, retrieve

Mem0 (2025) targets production: it extracts salient facts from conversations, consolidates them by updating or merging rather than piling up duplicates, and retrieves only what is relevant. On the LOCOMO benchmark, the authors report a 26% relative improvement on an LLM-as-judge metric over OpenAI's memory, and against a full-context baseline, 91% lower p95 latency and more than 90% lower token cost (Chhikara et al., Mem0). These are the authors' own measurements, so read them as a strong signal rather than an independent benchmark.

Lesson: consolidation (update, don't just append) is what keeps memory small enough to be useful.

Designing memory for a real agent

A practical design uses a few explicit layers:

  1. Working context (per run). Goal, instructions, the plan, recent tool results. Compress it when it grows. Anthropic's research agents save their plan to external memory because content beyond a 200,000-token window would be truncated, and they summarise completed phases before moving on (Anthropic, 2025).
  2. Typed long-term memory. Store facts, preferences and corrections as separate, short entries, not one growing blob. Types let you weight a correction more heavily than trivia.
  3. Retrieval by relevance. Fetch the handful of entries that matter for this task. Never replay everything.
  4. Consolidation. When a new fact contradicts an old one, update the old one.
  5. Reflection. After runs, distil what worked or failed into a lesson, and review it before it changes behaviour.
  6. Forgetting. Give entries a way to expire or be deleted. Stale memory is a quiet source of wrong answers.

Failure modes to design against

  • Memory poisoning. An agent that stores text from untrusted sources, like web pages or inbound email, can remember an attacker's instructions. Store conclusions, not raw untrusted text, and scope memory per agent.
  • Contradictions. Without consolidation, "the client prefers email" and "the client prefers phone" can both be retrieved.
  • Over-retrieval. Pulling in 50 "relevant" memories recreates the long-context problem you were avoiding.
  • Privacy creep. Memory is personal data. Know where it is stored and who can read it.

How AGNT implements memory

Disclosure: AGNT is our product. AGNT stores memory per agent, locally, as typed entries: fact, preference, correction, context, pattern, tool_insight, workflow_insight and prompt_guidance. Each entry carries a relevance score and is retrieved when it is relevant to the task, rather than replayed into every prompt. Insights the system derives from past runs are pending by default: nothing changes an agent until it is applied. The limitations are documented too. Memory is not a shared team brain, keyword recall is deliberately forgiving and poor at exact-phrase search, and if a step calls a hosted model, the content of that step goes to that provider. The design is described in the AGNT Memory whitepaper.

FAQ

What are the types of memory in AI agents?

Working memory for the current task, plus long-term episodic memory (past experiences), semantic memory (facts), and procedural memory (instructions, skills and code). The taxonomy comes from the CoALA framework.

Is RAG the same as agent memory?

RAG retrieves from a document collection, usually read-only. Agent memory is read and written by the agent as it works. They use similar retrieval machinery for different purposes.

Do longer context windows replace memory?

No. Chroma's tests across 18 models found performance degrades as input grows, and focused ~300-token prompts beat full ~113,000-token histories on the same questions.

How do AI agents remember users?

By writing short facts, preferences and corrections to an external store and retrieving the relevant ones at the start of later runs.

Sources

AI Agent MemoryAgent MemoryMemory SystemsContext EngineeringAgent Architecture