Release date: 2026-10-04. Major release. State at release: calibrating.
Why
Calibration of the 2.0.0 corpus showed 27 of 29 tasks solved by every configuration. Historical runs for every model and effort level clustered between 19 and 23 passes with no stable ordering. The benchmark no longer separated models. See benchmark-design-research.md and task-selection-3.0.md.
What changed
- **Corpus.** Twenty-three saturated tasks moved to
archive/2.0/and are
recorded as deprecated with reason saturated. Six 2.0 tasks remain (rust-no-network-build, devshell-tooling-contract, module-service-options, string-escaping-systemd, module-system-boundaries, xdg-portal-merge). Twenty-four new tasks were added from community-reported failures and verified language-semantics traps.
- **Prompts.** All prompts are symptom-driven. Source links and implementation
recipes were removed. See prompt-rewrite-3.0.md.
- **Evaluators.** A pinned nixpkgs
libis vendored undervendor/nixpkgs-lib
and exposed to evaluators as NIXBENCH_VENDOR_LIB. Module tasks grade through lib.evalModules; shell output is executed instead of regex-matched; each criterion evaluates in its own process so one failure cannot zero the rest. See evaluator-audit-3.0.md.
- **Scoring.** Task pass rate is the headline. At most one optional
maintainability or formatting criterion per task. New saturated empirical band at pooled solve rate 0.9 or above.
- **Reporting.**
bench.py study-matrixemits per-configuration pass rate with
Wilson intervals, pass@k and pass^k, per-task solve rates, failure sets, and discrimination. The site consumes that file directly.
- **Agents.** Fresh-home bubblewrap adapters for Codex CLI and Claude Code
strip user skills, settings, plugins, and MCP servers.
Lifecycle
Every task in this release is calibrating. Activation requires at least three configurations with three valid observations each, reviewed accepted-alternatives and evaluator-dispute status, and a saturation check. Published 3.0.0 results before activation are calibration evidence, not leaderboard claims.
Compatibility
2.0.0 trials remain bound to their digest and are shown only on the history page. No 2.0.0 result is rescored or pooled with 3.0.0.