people · May 16, 2026
AgentLens Framework Reveals Lucky Pass Problem in SWE-Agent Evaluation
Share the canonical public link.
Microsoft researchers Priyam Sahoo, Gaurav Mittal, Xiaomin Li, Shengjie Ma, Benjamin Steenhoek, Pingping Lin, and Yu Hu released AgentLens on May 13, 2026. The framework analyzes SWE-agent trajectories on SWE-Bench Verified and identifies that 10.7% of passing trajectories qualify as Lucky Passes due to regression cycles or blind retries. AgentLens-Bench contains 1,815 annotated trajectories with quality scores and 47 Prefix Tree Acceptor references. AgentLens changes model rankings by up to five positions when quality score replaces pass rate alone.