Skip to content

ResearchPaper

Organizational memory for interchangeable agents

How can an organization keep what it has learned when the model or agent doing the work is replaced? We propose a five-stage flow — Trace, Episode, Research, Review, Library — that turns working records into reviewed knowledge.

  • Memory
  • Agents
  • Organizations

Summary

In an organization where AI agents do much of the work, the worker keeps changing. Models are replaced every few months, and whether Claude, Codex, or Gemini handles a given task can change from one day to the next. The organization still needs to remember what it decided and what went wrong last time.

This paper explores keeping memory in the organization’s documents rather than inside any model. Agents don’t share a brain; instead, each one starts a task by reading documents the organization trusts. We propose a five-stage flow — Trace, Episode, Research, Review, Library — that turns raw working records into reviewed knowledge, one step at a time.

Problem

Agent memory today usually lives inside a single agent or a single product: a conversation history, a set of summarized memories, a product’s own storage. This has three weaknesses.

  • When you switch models or products, the memory either stays behind or has to be converted.
  • Records pile up, but nothing separates what is established from what was a passing guess.
  • The same fact gets written in several places, and the copies drift apart.

We run into this every day. The agent that acts as our secretary starts fresh in every session. If a new session doesn’t know about a correction our founder made in the previous one, the same mistake comes back. It is the familiar handover problem of a human team, except that it happens several times a day.

Approach

We design memory around where something is written, not who remembers it. Every agent reads from the same place when it starts and writes back to the same place when it finishes. Expertise belongs to places, too, not to agents: an agent takes on a role by reading the rules and guides kept in that place.

Records mature into knowledge through five stages:

  1. Trace — the raw record of what happened: conversations, commands, changes. Kept as is.
  2. Episode — one unit of work: what was asked, what was done, how it ended.
  3. Research — patterns and hypotheses drawn from many episodes.
  4. Review — a person or a separate agent checks each claim against the original records.
  5. Library — knowledge that has been checked. This is what an agent reads first on its next task.

Each stage is smaller and more trustworthy than the one before it. Anything in the Library can be traced back to the records it came from, and each fact is written in exactly one place.

This design sits at the center of Moss, the multi-agent system we have been building since June 2026. As of September 2026, the Moss repository has about 3,100 commits and roughly 160,000 lines of code, and our founder uses it for everyday work.

What we’ve seen so far

Two things have become clear from using it internally.

First, agents don’t much care how things are organized. They search and pull together what they need, so folder layout matters less than we expected. What trips them up is inconsistency and omission: two records that disagree, or a decision that was never written down. Both show up the same way: as a wrong answer given to a person.

Second, people need a way to see into the memory. An early prototype put almost everything in a database that could only be inspected through a text interface, and it became hard for anyone to tell whether the records were right. We are now trying a split by state — work in progress in git, settled decisions in a database, finished material in long-term storage — with a separate screen for what people need to see.

Open questions

  • Evaluation. How do we know the memory is working? One candidate test: when a person asks, can an agent return the right answer along with its source? We don’t yet have a reproducible way to run it.
  • The cost of review. If people check everything, people become the bottleneck. How much can be left to agents checking each other?
  • Forgetting. When and how should outdated knowledge be retired? We are leaning toward marking it as superseded rather than deleting it.
  • Memory inside models. Does research on giving models their own internal memory compete with this approach, or complement it? Our own experiments in that direction are on hold until we have the compute to run them.

Related

Version history

  1. v0.3Oct 4, 2026Added what we have learned from internal use and rewrote the open questions.
  2. v0.2Aug 27, 2026Organized the path from records to knowledge into five stages, from Trace to Library.
  3. v0.1Jul 24, 2026First draft: the problem, and the idea of keeping memory in documents rather than in models.

This document’s URL will not change. When citing it, please include the version number.

All research