Skip to main content

System status

Coverage is stale.

Collection is paused. Latest public event: Aug 21, 2026 (10 days ago).

← Intel index

This coverage is stale.

Last updated May 22, 2026 (about 3 months ago).

product · May 22, 2026

Fireworks AI Publishes Agent Execution Tax Benchmark with Nottecore on 720 Tasks

Share the canonical public link.

Share as image

Fireworks AI published a Notte × Fireworks AI benchmark report on May 20, 2026, detailing results from 720 browser agent tasks run across frontier models including Kimi K2.5, GLM-5, MiniMax M2.5, and Gemini 2.5 Flash. The report defines Agent Execution Tax as the ratio of wasted inference calls to productive ones and measures structured output reliability in multi-step agent loops with metrics such as retry rates, latency, token overhead, and reliability-adjusted accuracy. Kimi K2.5 recorded zero parse retries across 852 calls with 2.1s p50 latency while Gemini 2.5 Flash showed an 18.6% retry rate and 22.9% execution tax; GLM-5 achieved 57.1% task accuracy and MiniMax M2.5 delivered the lowest cost per successful task at $0.062. Nottecore and Fireworks AI conducted the benchmark to highlight execution reliability differences in production agent systems.

Below validation threshold — auto-passed without scoring

Supporting evidence