CYC26  ·  The Three Ways AI Agents Fail
Tools Memory Coordination
01 / 13
Commit Your Code 2026

The Three Ways
AI Agents Fail.

And how to fix yours.
Moeez Khan  ·  @moeez-khan
Views my own, not my employer's. No confidential information.
No agents were harmed making this talk. Several were mildly embarrassed.
Before we start

Has your agent ever once
said "I'm not sure"?

You've all probably felt this. Everyone working with AI agents likely has.
(If that's you, this is the right room.)
The moment it breaks
✓DEMO
✕REAL WORLD

It worked in the demo.

Then it met real code.
MIT · State of AI in Business 2025
95%
of companies' own GenAI pilots deliver no measurable return
only 5% reach real production
the model → the system around it  is where it breaks.
You might be thinking

"Won't a better model just fix this?" Partially, only if you're measuring whether it gets one task right in isolation. Not on these three failure modes, they're architectural, not capability-limited.

The Failure Map
1 · TOOLS 2 · MEMORY 3 · COORDINATION AGENT
Running example: Ledger, a legacy billing monolith an agent is modernizing.
By the time we're done here: you'll know exactly which of these three is breaking your agent, and the fix for each.
Layer 1 · Tools  /  the break
Ledger migration  ·  agent searches Ledger for interest logic

An agent is only as reliable
as the tools it calls.

agent search_codemisses half of Ledger wrong planno error thrown
True story

A code-search tool saw page one and reported "no other usages." 12 of 14 were on page two.

Layer 1 · Tools  /  the fix
Ledger migration  ·  fixing search_code
Paradigm: fail loud, not silent  ·  idempotency = an elevator button (press it 5x, one trip)
LOOSE
# anything goes, fails quietly
tool(query: str)
  → returns str (maybe empty)
STRICT
tool(
  query: Enum[...]          # typed
  on_error: Result | Failure  # must handle
  idempotent: true          # safe retry
)
Insight — confirm / replace

The failure that hurts isn't the tool erroring, it's the tool succeeding with wrong data. Give it a third state most schemas can't express: "succeeded, but low confidence."

You might be thinking

"Isn't this just normal API design?" Yes. That's the point, an LLM is a less forgiving caller than a human, it can't read your docs' caveats.

✓ Takeaway 1 — Reliable Tool Calls
(Somewhere, a tool schema is still returning "success" on a fire.)
Layer 2 · Memory  /  the break
Ledger migration  ·  what depends on Balance?

RAG retrieves snippets.
It misses how they relate.

Balance caller row config Overdraftrule missed · 3 hops
True story

A change flagged safe because retrieval never saw the caller two layers away.

Layer 2 · Memory  /  the fix
Ledger migration  ·  the same dependency, caught
Paradigm: retrieval vs. relationships  ·  RAG = index cards. Knowledge graph = a subway map.
Balance caught
Facts → RAG
Relationships → graph
wrong = a missed connection? → graph
Insight — confirm / replace

You don't need a perfect graph. A shallow one-hop dependency graph on top of RAG catches most silent drops, for a fraction of the effort.

You might be thinking

"What about GraphRAG, hybrid retrieval?" Real, and reasonable. The blind spot doesn't disappear, you're just choosing a different way to close it.

✓ Takeaway 2 — RAG vs. Knowledge Graph, know which, when
(RAG found three files that mention it. None of them were it.)
Layer 3 · Coordination  /  the break
Ledger migration  ·  three agents now, not one

More agents is not
more reliability.

analyzer migrator validator
40%
of agentic projects canceled by 2027, mostly missing risk controls (Gartner, Jun 2025)
True story

A validator that only checked whether the buggy stage had flagged itself successful.

Layer 3 · Coordination  /  the fix
Ledger migration  ·  the same three agents, contained
Paradigm: the bulkhead: watertight ship compartments, one floods, the ship floats.  +  least privilege.
analyzer migrator validator boundary human sign-off
Insight — confirm / replace

The human checkpoint is only a bottleneck if it's everywhere. Gate the irreversible ~5% of steps; leave the rest autonomous.

You might be thinking

"Doesn't a human checkpoint just re-add the bottleneck?" Only if you gate everything. Gate the 5% that's irreversible, not the throughput.

✓ Takeaway 3 — No-Cascade Coordination  ·  delegate the work, never the accountability
(Three agents walked into a pipeline. Only one walked out correct.)
What you take home
As promised: here's the map, and the fix for each.
Three-quarters are adopting agents. Few reach production. These three layers are the gap.
(Forrester, Jun 2026)
✓

The Failure Map

locate any agent bug
✓

Reliable Tool Calls

kill silent failures
✓

RAG vs. Knowledge Graph

right memory for the job
✓

No-Cascade Coordination

contain, don't spread
Take it with you

Thank you.

scan for everything
write-up · this deck · LinkedIn, all linked from there
Moeez Khan  ·  @moeez-khan  ·  find me in the hallway.
(If your agent is reading this slide: hi. Please be nice to production.)
← → / space  ·  F full-screen