"Won't a better model just fix this?" Partially, only if you're measuring whether it gets one task right in isolation. Not on these three failure modes, they're architectural, not capability-limited.
A code-search tool saw page one and reported "no other usages." 12 of 14 were on page two.
# anything goes, fails quietly tool(query: str) → returns str (maybe empty)
tool( query: Enum[...] # typed on_error: Result | Failure # must handle idempotent: true # safe retry )
The failure that hurts isn't the tool erroring, it's the tool succeeding with wrong data. Give it a third state most schemas can't express: "succeeded, but low confidence."
"Isn't this just normal API design?" Yes. That's the point, an LLM is a less forgiving caller than a human, it can't read your docs' caveats.
A change flagged safe because retrieval never saw the caller two layers away.
You don't need a perfect graph. A shallow one-hop dependency graph on top of RAG catches most silent drops, for a fraction of the effort.
"What about GraphRAG, hybrid retrieval?" Real, and reasonable. The blind spot doesn't disappear, you're just choosing a different way to close it.
A validator that only checked whether the buggy stage had flagged itself successful.
The human checkpoint is only a bottleneck if it's everywhere. Gate the irreversible ~5% of steps; leave the rest autonomous.
"Doesn't a human checkpoint just re-add the bottleneck?" Only if you gate everything. Gate the 5% that's irreversible, not the throughput.