Agents need documentation more than memory: what to write down
Answer first: most production agents do not need a memory system. They need documentation that is written for them, kept current, and loaded on purpose. A memory store accumulates whatever the agent happened to see and surfaces it by similarity. Documentation is curated by someone who knows what is true, versioned, and reviewable. For an agent inside a SaaS product, where the same workflows run thousands of times a day, curated docs beat accumulated memory on accuracy, cost, and debuggability. Memory still has a job, but it is narrower than the pitch decks suggest.
What people mean by agent memory, and why it disappoints
Agent memory usually means a store of past interactions, tool results, or extracted facts, retrieved at run time by similarity to the current task. The promise is that the agent gets better the more it runs. The reality for most teams is a store that grows without review, returns stale or contradictory snippets, and makes failures impossible to reproduce because the context on Tuesday was not the context on Monday.
Three problems show up repeatedly in production.
- Nobody owns the content. Facts are written by the agent itself, so an early mistake becomes a retrieved mistake, and there is no review step where a human would have caught it.
- Retrieval is by similarity, not by relevance to the step. The agent asks how to refund an order and gets five memories about refunds, two of which describe a policy that changed last quarter.
- It hides the real issue. When an agent fails on a workflow, the fix is usually a clearer instruction, a better tool description, or a missing business rule. A memory layer papers over that by sometimes getting it right, which makes the gap harder to see and fix.
We wrote earlier about how agent memory decays in production. The decay is the mechanism, and the missing documentation is the cause.
What documentation for an agent looks like
Agent documentation is a small set of files the agent reads at the start of a task or at the start of a step, written in plain language, owned by a person, and changed through review. It is the runbook a new hire would get, minus the small talk. In a typical SaaS product the set looks like this.
A domain glossary
Every product has words that mean something specific. In a billing product, "credit," "refund," and "adjustment" are three different ledger operations with three different approval rules. A model that does not have the glossary will use them interchangeably and call the wrong tool. A 40-line glossary fixes a whole class of failures that no amount of memory would.
Workflow procedures
For each workflow the agent handles, a short numbered procedure: the preconditions to check, the order of tool calls, what to do when a step fails, and when to stop and hand off to a human. This is the content that memory systems try to learn from examples, and it is far more reliable when written once by the person who owns the workflow. It is also where you encode the approval points where a human has to sign off.
Tool contracts
Each tool's description, its arguments, what it returns, its side effects, and its rate limits. Most of this already lives in the tool schema, but the parts that matter most, like "this tool is idempotent" or "never call this twice for the same order," usually do not. We covered the shape in how to design tools for production agents, and the documentation is the human-readable layer on top of the schema.
Policies and limits
What the agent may never do, the thresholds that require escalation, data handling rules per tenant tier. These change when legal or product changes them, and the change should be a reviewed edit to a file, not a hope that the memory store eventually notices.
Known failure modes
A short list of ways this workflow has gone wrong before and what to check first. This is the one place where something like memory earns its keep, except it is written by an engineer after an incident, not extracted automatically.
Why this works better for a product agent
The argument is not that documentation is more elegant. It is that a product agent runs the same small set of workflows at high volume, and that profile favors curation over accumulation.
Accuracy improves because the content is correct by construction. A reviewed procedure is right until someone changes it. A memory is right until the agent misreads something, which is a matter of time at volume.
Cost drops because you load a known, bounded set of text. A glossary plus a procedure plus the relevant tool contracts is often under 3k tokens. A similarity search over a memory store returns a variable amount of text, and teams routinely find they are paying for ten retrieved snippets to use one. The context engineering rules for production agents apply here: load what the step needs, not everything that is similar.
Debugging becomes possible. When a run fails, you can see exactly which docs were loaded and diff them against the last known-good run. With memory, the retrieved set differs per run and the failing combination may never recur.
Updates become a normal engineering change. Product changes a policy, someone edits the policy file, it goes through review, and the next run picks it up. There is no retraining, no re-embedding, no waiting for the memory to catch up.
Where memory still belongs
Documentation covers what is true about the product and the workflows. It does not cover what is true about this user, this account, or this session. That is the narrower, legitimate job of memory, and it is better thought of as state than as memory.
- Session state: what has already been done in this task, so the agent does not repeat a step after a retry.
- Account facts: this tenant's plan, their timezone, their configured approval threshold. These belong in the database you already have, fetched by a tool, not in a vector store.
- User preferences the user has stated explicitly, stored with a timestamp so they can be shown and corrected.
Notice that none of these require similarity search. They are lookups by key. The agent asks for the account's settings and gets the account's settings. That is the shape most memory needs take once you separate them from the documentation job.
The exception is open-ended personal assistants, where the set of tasks is unbounded and the user's history genuinely is the main source of truth. That is a different product with a different cost profile, and most SaaS agents are not that.
How to make the switch on an existing agent
If you already run an agent with a memory store, you do not have to rip it out. Start by reading what is in it. Export the top few hundred memories by retrieval frequency and sort them into three piles: facts about the product or workflow, facts about specific accounts, and noise. The first pile becomes documentation, written properly and reviewed. The second pile becomes database lookups through a tool. The third pile is deleted.
Then wire the docs in deliberately. Load the glossary and the policies on every run. Load the workflow procedure for the workflow the task maps to. Load tool contracts only for tools in that procedure, which keeps you clear of the tool sprawl problem where every tool description goes into every prompt. Version the doc set alongside the prompt so a failing run can be reproduced exactly.
Finally, put an owner on each file. The glossary belongs to product. The procedures belong to whoever owns the workflow. Tool contracts belong to the engineer who built the tool. A doc with no owner drifts exactly the way a memory store does, just more slowly.
Teams that bring us an agent that "works most of the time" are usually one glossary and three procedures away from one that works predictably. That rewrite is a bounded piece of work, and it is the kind of task our AI agent engineers take on in a week rather than a quarter.
Does documentation for agents replace RAG?
No. RAG retrieves from a large, changing corpus such as a knowledge base or customer documents. Agent documentation is a small, curated set that defines how the agent operates. Most products need both: the docs tell the agent how to behave, and RAG gives it the material to behave on.
How large should an agent's documentation be?
Small enough to load the relevant parts on every run without thinking about it. A glossary of 30 to 60 terms, procedures of 10 to 30 lines each, and tool contracts of a paragraph each is typical. If a file grows past a page, it is probably two files.
How do I keep agent documentation from going stale?
Treat it like code. Keep it in the repo next to the agent, require a review on changes, and add a check to your release process that any workflow change touches its procedure file. Record the doc version on every run so you can see which version a failure happened under.
Rather we just build it?
Book a free scoping call and we'll ship your production-safe AI feature this week.