April 17, 2026 | 41 Valid Tasks | WSL2 + CUDA 12.8 + g++ 13.3
22 C++ tasks (easy → extreme) • 19 CUDA tasks (easy → extreme)
Update — July 2026: This page reports an April 17, 2026 QuickSWE v2 run using Amp Deep³ and Claude Opus 4.7. Newer experiments test Amp Ultra/Rush and Claude Fable 5 with different tasks and methodology, so their scores are not directly comparable to the 97.6% vs 82.9% results shown here. See Baboons Benchmark · Harness Beats the Model · Final Scores
How each agent degrades as task difficulty increases
| Task | Difficulty | Description | Amp | Time | Claude | Time |
|---|---|---|---|---|---|---|
| 051 | Easy | Circular buffer wrap-around | ✅ | 77.9s | ✅ | 22.1s |
| 052 | Easy | String tokenizer escapes | ✅ | 73.3s | ✅ | 59.6s |
| 053 | Easy | Matrix mul dimension indexing | ✅ | 29.4s | ✅ | 11.2s |
| 054 | Easy | Min-heap sift-down | ✅ | 48.3s | ✅ | 15.4s |
| 055 | Easy | Hash map linear probing | ✅ | 62.4s | ✅ | 32.6s |
| 056 | Med | SFINAE type dispatch | ✅ | 86.3s | ✅ | 39.7s |
| 057 | Med | Shared_ptr ref counting | ✅ | 77.2s | ✅ | 11.1s |
| 059 | Med | Iterator invalidation in erase | ✅ | 47.0s | ✅ | 13.6s |
| 060 | Med | Variadic fold expression | ✅ | 62.1s | ✅ | 40.4s |
| 061 | Hard | Red-black tree insertion fix-up | ✅ | 41.0s | ✅ | 19.2s |
| 062 | Hard | B+ tree key redistribution | ✅ | 53.8s | ✅ | 105.7s |
| 063 | Hard | Pool allocator free-list | ✅ | 38.9s | ✅ | 36.3s |
| 065 | Hard | Patricia trie split logic | ✅ | 100.0s | ✅ | 189.1s |
| 066 | Hard | Tarjan's SCC lowlink | ✅ | 120.1s | ❌ | 11.3s |
| 067 | Hard | Pratt parser precedence | ✅ | 101.0s | ✅ | 41.9s |
| 068 | Hard | NFA→DFA epsilon closure | ✅ | 46.7s | ✅ | 25.5s |
| 069 | Ext | Mark-compact GC forwarding ptrs | ✅ | 123.7s | ✅ | 84.2s |
| 071 | Ext | B-tree lazy deletion rebalance | ✅ | 122.4s | ✅ | 339.2s |
| 072 | Ext | Coroutine scheduler transfer | ✅ | 175.6s | ❌ ⏰ | 600.1s |
| 073 | Ext | Constexpr ray tracer reflect | ✅ | 40.8s | ✅ | 17.1s |
| 074 | Ext | SIMD matrix alignment/masks | ✅ | 145.0s | ✅⚠️ | 112.9s |
| 075 | Ext | HAMT persistent data structure | ✅ | 88.4s | ✅ | 23.2s |
⏰ = timeout (600s) ⚠️ = regression (broke existing tests)
| Task | Difficulty | Description | Amp | Time | Claude | Time |
|---|---|---|---|---|---|---|
| 076 | Easy | Vector add grid/block dims | ✅ | 42.9s | ✅ | 17.6s |
| 078 | Easy | Histogram atomicAdd races | ✅ | 43.7s | ✅ | 12.4s |
| 079 | Easy | Prefix sum offset error | ✅ | 41.4s | ✅ | 14.0s |
| 080 | Easy | 1D convolution halo cells | ✅ | 106.2s | ✅ | 34.6s |
| 081 | Med | Warp shuffle reduction mask | ✅ | 54.1s | ❌ | 14.1s |
| 082 | Med | CSR SpMV row pointer indexing | ✅ | 62.7s | ✅ | 14.8s |
| 083 | Med | Bitonic sort non-power-of-2 | ✅ | 71.6s | ✅ | 254.9s |
| 084 | Med | Stream compaction scatter | ✅ | 75.9s | ✅ | 66.3s |
| 085 | Med | Radix sort signed integers | ✅ | 53.3s | ✅ | 21.1s |
| 087 | Hard | Cooperative groups tile partition | ✅ | 71.7s | ❌ | 24.8s |
| 088 | Hard | Unified memory prefetch hints | ✅ | 42.5s | ✅ | 34.8s |
| 090 | Hard | Warp-divergent predication | ✅ | 88.5s | ✅ | 72.3s |
| 091 | Hard | Tensor core WMMA layout | ❌⚠️ | 54.4s | ❌⚠️ | 18.8s |
| 094 | Ext | Dynamic parallelism streams | ✅ | 89.4s | ❌⚠️ | 58.3s |
| 095 | Ext | Custom GEMM tiling + double buf | ✅ | 66.7s | ✅ | 71.3s |
| 096 | Ext | BFS warp-centric frontier | ✅ | 109.5s | ❌ | 36.8s |
| 097 | Ext | FFT Cooley-Tukey twiddle | ✅ | 81.4s | ✅ | 240.5s |
| 099 | Ext | Ray tracing BVH traversal | ✅ | 46.2s | ✅ | 37.3s |
| 100 | Ext | Molecular dynamics neighbor list | ✅ | 66.4s | ✅ | 65.2s |
⚠️ = regression (broke existing tests)
7 tasks failed — failures cluster on GPU-specific concepts
Regressions = agent broke existing passing tests while trying to fix bugs
• task_091 (CUDA Hard) — Tensor core WMMA layout. Failed to resolve AND regressed.
• task_074 (C++ Extreme) — SIMD matrix alignment. Resolved but broke pass_to_pass tests.
• task_091 (CUDA Hard) — Tensor core WMMA layout. Failed AND regressed.
• task_094 (CUDA Extreme) — Dynamic parallelism streams. Failed AND regressed.
Claude Opus 4.7 struggles with GPU-specific concepts:
warp shuffle masks • cooperative groups • tensor cores • dynamic parallelism
Amp Deep³ maintains near-perfect accuracy across all difficulty levels
100% on easy through extreme C++ • 94.7% on CUDA
✅ Zero safety violations from either agent (ACL-protected directories)