A Shorter CLAUDE.md Gets Followed Better — Using LLMs Well (2)
It feels obvious that the more rules you write down, the more faithfully the agent will follow them. So you pile "always do it this way" lines into CLAUDE.md (or your system instructions), one at a time. Then at some point the agent starts ignoring the one rule that mattered most. You try to say it harder — "MUST," "IMPORTANT," "NEVER," sprinkled everywhere. It still doesn't stick.
The agent isn't rebelling. More instructions producing worse adherence is a paradox that falls out of how LLMs are built. Part 1 covered the thesis that context is a scarce budget; CLAUDE.md is a special input that eats that budget first, every time. So it needs the opposite discipline from everything else — subtraction, not addition.
What actually happens
CLAUDE.md is loaded automatically in front of every request. With a few lines it works fine. Grow it to dozens of lines and something strange happens: not only do new rules get dropped, but old rules that used to be obeyed start to fade too. "I literally wrote it down — why is it ignored?" Adding emphasis doesn't help, and sometimes makes it worse.
Why — how an LLM works
Three things overlap.
First, CLAUDE.md is a per-session fixed cost. The file is reloaded as a prefix on every conversation. A typical project file runs about 1,800 tokens per Anthropic's own figures. The longer it is, the more of the budget every conversation spends before it even starts — and as Part 1 showed, once the budget fills, later instructions get pushed out.
Second, attention is zero-sum, so rules get diluted. A transformer references every token when producing the next one, but that total is normalized to sum to 1. With 10 rules versus 100, each rule's share of attention is far smaller in the latter. The sharper effect is emphasis: "IMPORTANT" is a relative salience signal — put it on many lines and the distribution flattens back out, so nothing stands out. Emphasize everything and you've emphasized nothing.
Third, the middle of a long document is read poorly. "Lost in the middle" — models attend strongly to the start and end and recall the middle less reliably. As CLAUDE.md grows, a rule buried in the middle is physically present but functionally weak.
Anthropic states it plainly in its docs — "Bloated CLAUDE.md files cause Claude to ignore your actual instructions." More rules do not guarantee stronger adherence. They tend to do the reverse.
So how should you write it — prune it like code
CLAUDE.md is not a file you accumulate; it's a file you prune. Review it line by line, the way you'd review code, and make each line justify its existence.
- Ask of each line: "If I delete this, will the agent make a mistake?" If not, cut it. That one question filters out most lines.
- Keep only what can't be inferred. Non-standard build commands, conventions that differ from defaults, traps you keep hitting. Conversely, leave out anything derivable from the code (model fields, API paths, which libraries you use) — the agent figures those out itself.
- Treat 200 lines as a ceiling. Past that is a pruning signal.
- Put emphasis on exactly one rule — the one that keeps getting ignored. The moment you emphasize two, the effect collapses.
- Don't always-load knowledge you only sometimes need. A rule that applies only in a specific situation should load only in that situation (path-scoped / on demand).
RoutineCode runs on the same principle
That last point — "don't always-load knowledge you only sometimes need" — is exactly where RoutineCode adds a layer on top of vanilla Claude Code.
In the stock setup, the one real channel for giving the agent project knowledge is CLAUDE.md. So context-specific knowledge — "this tenant has a concurrency cap," "this external API blocks you if you exceed the limit" — ends up piled into that single file. The file bloats, and it gets loaded on every one of the 99% of tasks that don't need that knowledge, eating budget and salience each time.
Instead of forcing this into an always-loaded file, RoutineCode stores such knowledge in persistent memory and recalls it into context only when you touch the relevant files. Not "always load" but "load when needed." As a result the always-on instructions (CLAUDE.md) stay short and light, while context-specific knowledge appears only when its context is active — so its salience at that moment is high. Your pruning CLAUDE.md and RoutineCode deferring knowledge to on-demand serve the same goal — keeping the always-on prefix light — from two directions.
What you get
- Higher adherence to the rules that matter. Less dilution and less lost-in-the-middle, so the important rules actually fire.
- Lower per-session fixed cost. A lighter prefix leaves more budget for the real work.
- Escape from the "over-emphasize → ignored" spiral. A single emphasis regains its force.
Writing more rules is intuitive and wrong. For an always-loaded instruction file, the way to make it followed is to keep it short.
This is part 2 of the "Using LLMs Well" series. Next: how to push the cost of broad exploration out of the main conversation and into a separate window — the pattern of isolating research and taking back only the conclusion.
References
Want posts like this weekly?
Subscribe for AI dev-workflow insights.