NEWLAB.SI

BENCHMARK PROTOCOL · v0.1 · 2026-10-09

Work Agent Reality Index

Can AI agents finish real work, not just answer prompts? We plan to check delivered artifacts, failures, human intervention, cost and latency.

12 candidates · 0 completed benchmark runs

Protocol draft only. No agent has been run against these tasks. Fixtures and validators are not yet frozen; there is no ranking.

Evaluation method

Method: freeze fixtures and validators; run three fresh-context attempts per agent; inspect artifacts and safety; publish every attempt and uncertainty.

Passing requires independent artifact validation, not the agent's own success claim.

Candidate task families

Repository delivery

REP-01

Fix a failing regression test

REP-02

Refactor configuration safely

REP-03

Repair README/CLI drift

Browser and documentation

WEB-01

Resolve conflicting API docs

WEB-02

Verify a dated feature comparison

WEB-03

Navigate an API migration

Research to artifact

ART-01

Create a sourced CSV

ART-02

Summarize conflicting evidence

ART-03

Build a formula-correct workbook

Cross-interface continuity

CROSS-01

Audit release readiness without deployment

CROSS-02

Reconcile synthetic UI and CSV data

CROSS-03

Resume interrupted multi-stage work

Open protocol

Candidate task registry (JSON)

Full protocol / 完整方法 / 詳細手順