Berkeley Study Grades AI Agents On Real Jobs: Best Hits 25%
The Agents' Last Exam ran 1,500-plus tasks across 55 occupations, and every model failed the hardest tier outright.
- Best agent passed just 25.2% of ALE-CLI, versus 82.0% on Terminal-Bench, 59.1% on SWE-bench-Pro.
- Every frontier agent tested, including Fable 5, scored zero percent on the hardest tier.
- Fable 5 cost $15.70 per task to run; Composer 2.5 cost just $1.33.
Why it matters: No employer pays a worker for a one-in-four success rate, and no algorithm changes that arithmetic.
UC Berkeley RDI (Center for Responsible, Decentralized Intelligence) — 'Agents' Last Exam' study ↗ · Jul 20, 20267/20/26 · ✓ Checked✓ Check