Skip to main content

System status

Coverage is stale.

Collection is paused. Latest public event: Aug 21, 2026 (11 days ago).

← Intel index

This coverage is stale.

Last updated May 22, 2026 (about 3 months ago).

people · May 22, 2026

ProgramBench paper lists mini-swe-agent for test generation and baselines

Share the canonical public link.

Share as image

The ProgramBench paper from May 5, 2026, describes using mini-swe-agent to extend test generation pipelines and reproduce baselines on software reconstruction tasks. Authors include researchers from Meta FAIR, Meta TBD, Stanford University, and Harvard University such as John Yang, Kilian Lieret, and Ofir Press. The benchmark requires agents to architect and implement codebases matching reference executables given only a program and documentation. Mini-swe-agent supports evaluation across frontier models on this challenging long-horizon task.

Below validation threshold — auto-passed without scoring

Supporting evidence