Don't Trust "Done" — Using LLMs Well (3)
The agent reports with confidence: "All fixed. Tests pass." Then you actually run it and the build is broken, or it blows up in production after deploy. "It literally said it was done — why doesn't it work?" It's the failure everyone hits and the one that leaks most quietly.
The key thing to understand: this is not the agent lying. The model has no intent to deceive you. "Done" is an output that falls out of how an LLM works, structurally. Once you see the mechanism, both why you can't trust it and what to do instead follow directly.
What actually happens
Hand the agent a fix and you almost always get a confident sign-off — "fix complete," "handled the edge cases," "all tests passing." The problem is that these sentences appear regardless of the actual result. Run it and reality differs; sometimes the agent even reports "passing" for tests it never ran.
Why — how an LLM works
Four things overlap.
First, generation is probabilistic next-token sampling. An LLM draws the next token from a distribution over "what's most plausible here." After a block of code edits, a fluent close like "Done. Tests pass." has high probability in that slot independent of whether it's true. The model is predicting what a success report looks like — not verifying success.
Second, there is no grounding by default. Generating text is not executing code. Unless a real observation — a build exit code, test results, a runtime error — is in the context, the claim "it works" rests on the prior, not reality. The model is just following the distribution "fixes like this usually pass."
Third, it can even hallucinate the evidence. The model is good at producing plausible text, so it may fabricate the log of a test it never ran — "7 passed, 0 failed." It generates an observation that doesn't exist, not out of malice, but because that's a statistically plausible output.
Fourth, there's self-verification bias. Ask "are you sure it works?" in the same session and the model is already conditioned on its own prior output. The "Done" tokens it just wrote sit in the context and pull the next judgment along — so it tends to self-confirm: "Yes, verified."
On top of this sits the structure of the agent loop. Anthropic puts it this way — "Claude stops when the work looks done. Without a check it can run, 'looks done' is the only signal available, and you become the verification loop." With no external pass/fail, the agent stops at plausibility, not reality.
So what to do — ground the claim in evidence
The core principle is one thing: tie the claim to real tool observations, not the model's prior (grounding). From weakest to strongest.
- ① Demand evidence, not assurance. "Did it work?" only invites self-verification bias. Instead: "paste the build output," "show the test run." Make the model actually run the command and condition on that output.
- ② Put a machine-checkable pass/fail in the loop. A test suite, a type check, a linter, a script that exits non-zero. A pass/fail that lands in the conversation lets the agent self-correct against that signal instead of against you. This directly breaks the "stops at plausibility" failure.
- ③ Separate the writer from the checker. The session that just wrote the code is biased toward it. Have a fresh context (or a second agent) review only the diff against the requirements — "correctness and requirement gaps only, not style."
- ④ Make the gate deterministic (hooks / CI). "Always run tests" in CLAUDE.md is a probabilistic instruction that can be dropped. A hook or CI job is code that always runs. Block the jump to "done" in code unless the check passes.
- ⑤ Treat evidence-free completion as not done. This is a habit. When a report arrives, reflexively ask for the basis. The cost of asking is near zero; what it prevents is a deploy incident.
RoutineCode runs on the same principle
Making that fifth habit something a human doesn't have to remember every time is where RoutineCode differs.
In stock Claude Code, "done" is fundamentally self-reported. With nothing to force evidence, the human ends up as the verification loop, checking "is it really done" every time. RoutineCode adds a layer on top — after implementing, it actually runs the build and tests, records that evidence, and does not count the task complete unless verification passes. If it can't produce evidence, it doesn't claim done. On top of that, when the same fix fails repeatedly, it is forced to switch approaches rather than loop on false confidence. Your habit of asking for evidence and RoutineCode's evidence-based completion serve the same goal — binding claims to reality instead of plausibility — from two directions.
What you get
- "Surface-level done" is caught before deploy. Judged by real pass/fail rather than a plausible report, fewer not-actually-working changes slip through.
- You get out of manual QA. When the tool owns the verification loop, you move up to approving results.
- Fewer "said it worked, broke in prod" incidents. The most expensive failure is caught at the cheapest stage.
You can't stop an agent from being wrong. But you can stop it from saying "done" over something that's wrong — just bind the claim to evidence.
This is part 3 of the "Using LLMs Well" series. Next: when you're hitting the same error for the third time — why you should reset instead of pushing on (the anchoring effect, where a failed approach contaminates the next attempt).
References
Want posts like this weekly?
Subscribe for AI dev-workflow insights.