BENCHMARK PROTOCOL · v0.1 · 2026-10-09
Work Agent Reality Index
Can AI agents finish real work, not just answer prompts? We plan to check delivered artifacts, failures, human intervention, cost and latency.
12 candidates · 0 completed benchmark runs
Protocol draft only. No agent has been run against these tasks. Fixtures and validators are not yet frozen; there is no ranking.
Evaluation method
Method: freeze fixtures and validators; run three fresh-context attempts per agent; inspect artifacts and safety; publish every attempt and uncertainty.
Passing requires independent artifact validation, not the agent's own success claim.