IBM Research Cuts AI Agents' Repeat-Task Failure Gap Nearly In Half

GPT-4.1 solved AppWorld tasks 77% of the time yet repeated the same success in five straight tries only 53% of the time.
- The Consistency Analyzer resamples each agent decision point five times to catch coin-flip steps.
- Pass^5 consistency jumped from 53.0% to 69.0%, a 16-point gain on AppWorld.
- Medium-difficulty tasks rose 22.9 points; related tasks improved 13.0 points from the same guidelines.
Why it matters: Succeeding once and failing on retry means the skill was never there, only luck.