Agents don’t always crash. Sometimes they slow down
By Rohan Kataria · Mar 26, 2026
Key points
Slow agents hold resources, back up queues and cascade failures downstream before anyone notices.
The router sends work to a pool instead of a single agent, and assigns it by health, not just availability.
When an agent's p95 latency crosses the threshold, the router shifts its traffic to a degraded pool.
The degraded pool drops RAG context and serves cached results, so the user still gets an answer.
Updates and sources
Sources I cited in the comments: