Sourced News

Atom Brief

Tue, Jul 21, 2026 · 0 stories · Confirmed · Subscribe
Tech

Berkeley Study Grades AI Agents On Real Jobs: Best Hits 25%

The Agents' Last Exam ran 1,500-plus tasks across 55 occupations, and every model failed the hardest tier outright.

  • Best agent passed just 25.2% of ALE-CLI, versus 82.0% on Terminal-Bench, 59.1% on SWE-bench-Pro.
  • Every frontier agent tested, including Fable 5, scored zero percent on the hardest tier.
  • Fable 5 cost $15.70 per task to run; Composer 2.5 cost just $1.33.

Why it matters: No employer pays a worker for a one-in-four success rate, and no algorithm changes that arithmetic.

UC Berkeley RDI (Center for Responsible, Decentralized Intelligence) — 'Agents' Last Exam' study ↗ · Jul 20, 20267/20/26 · ✓ Checked✓ Check