product · May 22, 2026
Fireworks AI Publishes Agent Execution Tax Benchmark with Nottecore on 720 Tasks
Share the canonical public link.
Fireworks AI published a Notte × Fireworks AI benchmark report on May 20, 2026, detailing results from 720 browser agent tasks run across frontier models including Kimi K2.5, GLM-5, MiniMax M2.5, and Gemini 2.5 Flash. The report defines Agent Execution Tax as the ratio of wasted inference calls to productive ones and measures structured output reliability in multi-step agent loops with metrics such as retry rates, latency, token overhead, and reliability-adjusted accuracy. Kimi K2.5 recorded zero parse retries across 852 calls with 2.1s p50 latency while Gemini 2.5 Flash showed an 18.6% retry rate and 22.9% execution tax; GLM-5 achieved 57.1% task accuracy and MiniMax M2.5 delivered the lowest cost per successful task at $0.062. Nottecore and Fireworks AI conducted the benchmark to highlight execution reliability differences in production agent systems.