Skip to content

Documentation

NixBench 3.0.0

Release date: 2026-10-04. Major release. State at release: `calibrating`.

Maintained in docs/releases/3.0.0.md

Documentation

Release date: 2026-10-04. Major release. State at release: calibrating.

Why

Calibration of the 2.0.0 corpus showed 27 of 29 tasks solved by every configuration. Historical runs for every model and effort level clustered between 19 and 23 passes with no stable ordering. The benchmark no longer separated models. See benchmark-design-research.md and task-selection-3.0.md.

What changed

  • **Corpus.** Twenty-three saturated tasks moved to archive/2.0/ and are

recorded as deprecated with reason saturated. Six 2.0 tasks remain (rust-no-network-build, devshell-tooling-contract, module-service-options, string-escaping-systemd, module-system-boundaries, xdg-portal-merge). Twenty-four new tasks were added from community-reported failures and verified language-semantics traps.

  • **Prompts.** All prompts are symptom-driven. Source links and implementation

recipes were removed. See prompt-rewrite-3.0.md.

  • **Evaluators.** A pinned nixpkgs lib is vendored under vendor/nixpkgs-lib

and exposed to evaluators as NIXBENCH_VENDOR_LIB. Module tasks grade through lib.evalModules; shell output is executed instead of regex-matched; each criterion evaluates in its own process so one failure cannot zero the rest. See evaluator-audit-3.0.md.

  • **Scoring.** Task pass rate is the headline. At most one optional

maintainability or formatting criterion per task. New saturated empirical band at pooled solve rate 0.9 or above.

  • **Reporting.** bench.py study-matrix emits per-configuration pass rate with

Wilson intervals, pass@k and pass^k, per-task solve rates, failure sets, and discrimination. The site consumes that file directly.

  • **Agents.** Fresh-home bubblewrap adapters for Codex CLI and Claude Code

strip user skills, settings, plugins, and MCP servers.

Lifecycle

Every task in this release is calibrating. Activation requires at least three configurations with three valid observations each, reviewed accepted-alternatives and evaluator-dispute status, and a saturation check. Published 3.0.0 results before activation are calibration evidence, not leaderboard claims.

Compatibility

2.0.0 trials remain bound to their digest and are shown only on the history page. No 2.0.0 result is rescored or pooled with 3.0.0.