THE ACTORS
SECTION 02
ISSUE 001
Economic Workloads Replace Toy Agent Demos
Projection: GPT-5.6 advances practical coding, computer use, knowledge work, and token efficiency; METR’s current time-horizon tracker measures how reliable task completion changes with task duration, while RE-Bench remains a fixed methodological precedent. Together they move agent evaluation toward economically legible work under real time and cost constraints.
Why this idea is here
What the evidence establishes.
OpenAI documents GPT-5.6 gains in coding, computer use, tool-heavy work, and useful work per token; METR maintains a current task-completion time-horizon tracker. RE-Bench supplies a dated time-budget methodology, not a current leaderboard claim.
Source ledger
Read the sources.
- S01GPT-5.6: Frontier intelligence that scales with your ambition
official model release / published 2026-07-09 / retrieved 2026-07-10
- S02RE-Bench
peer-reviewed conference paper / published 2025-07-01 / retrieved 2026-07-09
- S03Task-completion time horizons of frontier AI models
current independent evaluation tracker / published 2026-05-08 / retrieved 2026-07-10