AI Models
New benchmark scores AI agents on real jobs, not puzzles
An independent consortium launched an evaluation suite built from genuine workplace tasks — bookkeeping, support queues, code maintenance — with pay-rate baselines.
By Priya Sharma, Research Editor — BOSTON
BOSTON — An independent research consortium on Monday launched what it calls the first economically grounded benchmark for AI agents: a suite of tasks lifted from real workplaces — reconciling ledgers, resolving support tickets, maintaining ageing codebases — each scored against the time and cost of the professionals who currently do them.
Frontier agents completed just under half the tasks at professional quality, the consortium reported, but with wide variance: near-parity on structured digital work, steep drop-offs where tasks required chasing missing context across tools and people.
The suite will refresh quarterly with held-out tasks, and its maintainers pledged to publish results for every major model — including the failures.
Enable JavaScript to read the full story on Neural Daily News.