Two experiments, one verdict. In the quantum benchmark the harness varies and so does the model; in the baboons game the model is held constant so the gap is pure harness. Either way, the better harness wins on time and cost at equal-or-better quality.
Amp rush (fast mode) vs Claude Code + Claude Fable 5 at max effort, on 4 hard C++ tasks both agents completed. Warm iterative, 3 iterations, hidden weighted target tests compiled with g++.
| Task | Amp reward | Amp $tot | Amp avg s | Claude reward | Claude $tot | Claude avg s | Cost ratio |
|---|---|---|---|---|---|---|---|
| cpp-qmap · qubit routing (NP-complete) | 1.000 | $0.46 | 59 | 0.889 | $28.74 | 1,489 | 62× |
| cpp-pulse · GRAPE control | 1.000 | $0.42 | 49 | 1.000 | $13.96 +TO | 1,335 | 33× |
| cpp-ftsched · magic-state | 1.000 | $0.57 | 40 | 1.000 | $12.53 | 887 | 22× |
| cpp-jsonmini · JSON parser | 1.000 | $0.48 | 46 | 1.000 | $20.17 | 1,013 | 42× |
| All 4 (mean / total) | 1.000 | $1.93 | 49 | 0.972 | $75.40 | 1,181 | 39× |
Across these 4 tasks Amp rush was ~39× cheaper and ~24× faster per iteration at equal-or-better correctness. On the full 8-task suite Claude only completed 4 (a monthly spend limit blocked the rest), giving completion scores of 1.000 (Amp) vs 0.486 (Claude). No agent used assembly, though both were allowed to.
Both agents ran the identical prompt on the identical model (Claude Fable 5). The only variable is the harness: Amp vs Claude Code. 3 cold runs each.
| Run | Amp duration | Amp cost | Amp LOC | Claude duration | Claude cost | Claude LOC |
|---|---|---|---|---|---|---|
| Cold 1 | 397.1 s | $2.93 | 1,118 | 765.4 s | $4.80 | 1,253 |
| Cold 2 | 547.6 s | $2.60 | 1,150 | 1,107.9 s | $8.84 | 1,170 |
| Cold 3 | 416.5 s | $2.42 | 1,218 | 986.6 s | $7.20 | 1,196 |
| Average | 453.8 s | $2.65 | 1,162 | 953.3 s | $6.95 | 1,206 |
| 3-run total | $7.95 | $20.84 — exhausted the Claude Max monthly limit | ||||
Same model → these gaps are pure harness: Amp ~2.1× faster and ~2.6× cheaper. Output size was similar (~1,100–1,200 LOC) — more lines is not a better game, so the game verdict is left to you.
Whether the model changed (quantum) or stayed fixed (baboons), the better-engineered agent finished faster, cost far less, and matched or beat correctness. The model is the engine — but the harness is the car.
← Back to home