News, Fast

Atom Brief

Sep 15, 2026 · Archive
Tech

IBM Research Cuts AI Agents' Repeat-Task Failure Gap Nearly In Half

IBM Research Cuts AI Agents' Repeat-Task Failure Gap Nearly In Half

GPT-4.1 solved AppWorld tasks 77% of the time yet repeated the same success in five straight tries only 53% of the time.

  • The Consistency Analyzer resamples each agent decision point five times to catch coin-flip steps.
  • Pass^5 consistency jumped from 53.0% to 69.0%, a 16-point gain on AppWorld.
  • Medium-difficulty tasks rose 22.9 points; related tasks improved 13.0 points from the same guidelines.

Why it matters: Succeeding once and failing on retry means the skill was never there, only luck.

Hugging Face ↗ · Sep 15, 20269/15/26