Agents don’t always crash. Sometimes they slow down

By Rohan Kataria · Mar 26, 2026

Key points

  • Slow agents hold resources, back up queues and cascade failures downstream before anyone notices.

  • The router sends work to a pool instead of a single agent, and assigns it by health, not just availability.

  • When an agent's p95 latency crosses the threshold, the router shifts its traffic to a degraded pool.

  • The degraded pool drops RAG context and serves cached results, so the user still gets an answer.

Updates and sources

Sources I cited in the comments: