Is Your AI Agent Losing Its Mind?
Measure AI agent reliability using
Toby Ord's exponential decay model
Swipe or use arrows to navigate →
🎯 5-minute task — 90% success
⚠️ 1-hour task — 50% success
💀 4-hour task — 6% success
This isn't a bug — it's mathematics.
University of Oxford · May 2025
arXiv:2505.05115 · tobyord.com
Agent success rates decay exponentially with task duration — just like radioactive decay.
Each agent has its own half-life.
Building on METR's empirical work (Kwa et al., 2025)
λ = hazard rate (chance of failing per minute)
T½ = ln(2) / λ = your agent's half-life
Double the task duration → Square the failure rate
| Reliability | Max Duration | Use Case |
|---|---|---|
| 50% success | T½ (baseline) | Coin flip |
| 80% success | T½ ÷ 3 | Routine work |
| 90% success | T½ ÷ 7 | Important tasks |
| 99% success | T½ ÷ 70 | Critical tasks |
| 99.9% success | T½ ÷ 700 | Mission-critical |
If T½ = 60 min → 99% reliable for only 51 seconds!
Kwa et al. (2025) — METR benchmark · 170 tasks
T₅₀ = 59 min
T₈₀ = 15 min
Capabilities double every
7 months
50% → 99% at same task length:
~4 years
Based on METR benchmarks (Kwa et al., 2025)
Your results may differ. Test YOUR agent.
Tasks = chain of subtasks. Fail any one → entire task fails.
6 subtasks × 95% = 73% overall
20 subtasks × 95% = 36% overall
50 subtasks × 95% = 7.7% overall
Key insight: Current AI agents don't recover from earlier mistakes — unlike humans.
🚀
Across 5 difficulty tiers
📊
Fit exponential model to your data
🎯
Exact T½ with verdict
Works with Claude · GPT · Amp · Gemini · Any agent
Tasks calibrated by human-equivalent completion time
Does the code compile/run?
Is the logic correct?
Are edge cases handled?
Is the solution complete?
No human needed — fully automated benchmarking
✅ Verdicts: Excellent · Good · Degrading · Critical · Broken
📄 Export results as Markdown report
Amp's design philosophy already does this
"Prefer a sequence of small, validated edits over one large change"
Agent hazard rates actually decline over time.
→ Agents get relatively better at later subtasks
→ The exponential model is pessimistic — reality is slightly better
But the half-life model remains the right baseline for planning.
Better to plan conservatively and be pleasantly surprised.
Embedded Software Laboratory
We build tools for software quality & AI assurance
SBOM generation & vulnerability scanning
AI agent reliability benchmarking
📧 sales@eswlab.com
1. Download the installer
2. Select your agent (Claude Code / API / OpenAI)
3. Click Run All — get your half-life in minutes
After trial: $50/year per seat
Contact: sales@eswlab.com
Questions?
📧 sales@eswlab.com
Based on:
Toby Ord, "Is there a Half-Life for the Success Rates of AI Agents?" (2025)
arXiv:2505.05115 · tobyord.com
Kwa et al., "Measuring AI Ability to Complete Long Tasks" (2025)
© 2025-2026 ESW Lab Ltd. All rights reserved.