Agent memory is not a substitute for readable docs

Agent memory is not a substitute for readable docs

4 min read

A Hacker News provocation gets at a real builder problem: many agent systems do not fail because they forgot a user preference. They fail because the work rules, constraints, examples, and handoffs were never written down in a form the agent can use.

TL;DR: Most agent failures blamed on missing memory are really failures of shared context, so start with documentation the agent can read, test, and update before adding a memory system.

What does documentation give an agent that memory does not?

The Hacker News thread titled “Agents don’t need memory, they need documentation” is a useful provocation because it flips the usual agent pitch around.

Memory sounds more human. Documentation sounds boring. But boring is often what production systems need.

A memory system usually means the agent can store facts from past interactions and retrieve them later. That can help with stable preferences, recurring workflows, and long-running projects. The catch is that memory is a write path. The agent writes something down, often with fuzzy importance scoring, then later treats the retrieved note as context. If that note is stale, incomplete, or too broad, the agent can confidently repeat old mistakes.

Documentation is different. It is explicit operating context. What is the product? What are the rules? What is the tone? What files matter? What are the known edge cases? What should never be changed without approval? What does a good output look like?

That is not memory. That is a manual.

For agents, a good manual beats a pile of remembered fragments. It gives the model a stable frame before it starts planning. It can be reviewed by humans. It can be versioned. It can be tested against tasks. Most important, it can be wrong in a visible way. Memory often hides its errors until the agent acts on them.

Where does agent memory still make sense?

I do not read the Hacker News title as saying memory is useless. I read it as a warning about sequencing.

Memory is useful after the system has a clear base layer. If an agent helps a user every day, it should remember that the user prefers concise drafts, works in a certain repo, or always wants test files included with code changes. If an agent runs customer support, it may need case history. If it manages a research project, it may need a running project log.

But memory should not carry the weight of process design.

If the agent needs to know how to triage a bug, write that down. If it needs to know which APIs are deprecated, write that down. If it needs to know how your team names branches, write that down. If it needs to know when to ask for approval, write that down.

Then let memory capture the small deltas: user preferences, recent decisions, unresolved questions, and state that changes between sessions.

a cluttered pile of loose notes beside a clean binder feeding into a small robot workspace

The practical pattern is documentation first, memory second, retrieval third. Give the agent a canonical place to look. Add memory only for things that are personal, recent, or too dynamic for the manual. Then use retrieval to pull the right slice at the right time.

How should builders structure docs for agents?

Human docs are often written for onboarding. Agent docs should be written for action.

That means fewer essays and more task-shaped context. A coding agent benefits from a repo map, install steps, test commands, style rules, common failure modes, and examples of accepted patches. A marketing agent benefits from positioning, banned claims, approved terminology, customer segments, proof points, and sample outputs. An operations agent benefits from escalation rules, tool permissions, data handling rules, and checklists.

The document should answer the questions the agent will otherwise guess.

It also needs ownership. If nobody maintains the docs, the agent will drift. If every team drops random notes into one giant file, the agent will drown. Treat agent docs like an API surface. Small, named, current, and tied to real tasks.

The evaluation loop matters too. Run the same task with and without the docs. See what improves. Watch for overfitting, where the agent follows a rule too literally and misses the actual goal. Watch for contradiction, where the docs say one thing and the codebase says another. That is not an agent problem. That is organizational debt made visible.

My practitioner’s take: before buying or building agent memory, write a one-page operating manual for the job you want the agent to do, then run ten real tasks against it. Add only the memories that would have changed the outcome. The catch most teams miss is that agent readiness is not a model feature. It is whether your work is legible enough for another worker, human or machine, to pick up without guessing.