1 / 17
ESW LAB LTD.

Agent Half-Life Tester

Is Your AI Agent Losing Its Mind?

Measure AI agent reliability using
Toby Ord's exponential decay model

Swipe or use arrows to navigate →

AI Agents Fail More
As Tasks Get Longer

🎯 5-minute task — 90% success

⚠️ 1-hour task — 50% success

💀 4-hour task — 6% success

This isn't a bug — it's mathematics.

Toby Ord's Discovery

University of Oxford · May 2025

arXiv:2505.05115 · tobyord.com

Key Finding

Agent success rates decay exponentially with task duration — just like radioactive decay.

Each agent has its own half-life.

Building on METR's empirical work (Kwa et al., 2025)

The Math — Simplified

S(t) = e−λt

λ = hazard rate (chance of failing per minute)

= ln(2) / λ = your agent's half-life

50%
0 minTask Duration →

Double the task duration → Square the failure rate

What Your Half-Life
Really Means

ReliabilityMax DurationUse Case
50% success (baseline)Coin flip
80% successT½ ÷ 3Routine work
90% successT½ ÷ 7Important tasks
99% successT½ ÷ 70Critical tasks
99.9% successT½ ÷ 700Mission-critical

If T½ = 60 min → 99% reliable for only 51 seconds!

What The Research Found

Kwa et al. (2025) — METR benchmark · 170 tasks

Claude 3.7 Sonnet

T₅₀ = 59 min

T₈₀ = 15 min

Doubling Time

Capabilities double every

7 months

Reliability Gap

50% → 99% at same task length:

~4 years

How Agents Compare

Based on METR benchmarks (Kwa et al., 2025)

Claude 3.7 Sonnet
59 min
Claude 3.5 Sonnet
~25 min
GPT-4o
~18 min
Gemini 1.5 Pro
~12 min
GPT-4o-mini
~8 min

Your results may differ. Test YOUR agent.

Why Agents Fail

Tasks = chain of subtasks. Fail any one → entire task fails.

95%
95%
95%
95%
95%
95%

6 subtasks × 95% = 73% overall

20 subtasks × 95% = 36% overall

50 subtasks × 95% = 7.7% overall

Key insight: Current AI agents don't recover from earlier mistakes — unlike humans.

Now YOU Can Measure This

Agent Half-Life Tester

🚀

26 Automated Tasks

Across 5 difficulty tiers

📊

Decay Analysis

Fit exponential model to your data

🎯

Your Half-Life

Exact T½ with verdict

Works with Claude · GPT · Amp · Gemini · Any agent

26 Tasks Across 5 Tiers

Tier 1(1-5 min)Fix syntax · Hello World · Regex
Tier 2(5-15 min)Algorithms · Debugging · REST APIs
Tier 3(15-60 min)Refactoring · CLI tools · State machines
Tier 4(1-4 hrs)Full APIs · Design patterns · Pipelines
Tier 5(4+ hrs)Full-stack · Architecture · Search engines

Tasks calibrated by human-equivalent completion time

Automated LLM-as-Judge

📝 Task Prompt
🤖 Agent Response
⚖️ Judge Evaluates
✅ PASS / ❌ FAIL

The Judge Asks:

Does the code compile/run?

Is the logic correct?

Are edge cases handled?

Is the solution complete?

No human needed — fully automated benchmarking

Your Results at a Glance

HALF-LIFE
42.5m
VERDICT
Good
SURVIVAL CURVE
📉 Exponential decay chart
TIER SUCCESS
📊 Bar chart by tier

✅ Verdicts: Excellent · Good · Degrading · Critical · Broken

📄 Export results as Markdown report

Agent Failing?
Here's How to Fix It

Amp's design philosophy already does this
"Prefer a sequence of small, validated edits over one large change"

Good News — Feb 2026 Update

Gus Hamilton's Analysis (Jan 2026)

Agent hazard rates actually decline over time.

→ Agents get relatively better at later subtasks

→ The exponential model is pessimistic — reality is slightly better

But the half-life model remains the right baseline for planning.

Better to plan conservatively and be pleasantly surprised.

ESW Lab

Embedded Software Laboratory

We build tools for software quality & AI assurance

SBOMator

SBOM generation & vulnerability scanning

Half-Life Tester

AI agent reliability benchmarking

📧 sales@eswlab.com

Try It Free for 7 Days

1. Download the installer

2. Select your agent (Claude Code / API / OpenAI)

3. Click Run All — get your half-life in minutes

⬇ Download v1.0.0 📖 User Guide

After trial: $50/year per seat

Contact: sales@eswlab.com

Thank You

Questions?

📧 sales@eswlab.com

Based on:

Toby Ord, "Is there a Half-Life for the Success Rates of AI Agents?" (2025)

arXiv:2505.05115 · tobyord.com

Kwa et al., "Measuring AI Ability to Complete Long Tasks" (2025)

arXiv:2503.14499 · metr.org

© 2025-2026 ESW Lab Ltd. All rights reserved.