ProgramBench paper lists mini-swe-agent for test generation and baselines
Below validation threshold — auto-passed without scoring
The ProgramBench paper from May 5, 2026, describes using mini-swe-agent to extend test generation pipelines and reproduce baselines on software reconstruction tasks. Authors include researchers from Meta FAIR, Meta TBD, Stanford University, and Harvard University such as John Yang, Kilian Lieret, and Ofir Press. The benchmark requires agents to architect and implement codebases matching reference executables given only a program and documentation. Mini-swe-agent supports evaluation across frontier models on this challenging long-horizon task.