Skip to content

Documentation

NixBench 3.1.0

Release date: 2026-10-04. Major release. State at release: `calibrating`.

Maintained in docs/releases/3.1.0.md

Documentation

Release date: 2026-10-04. Major release. State at release: calibrating.

Why

The 3.0.0 calibration (2026-10-04-nixbench-3-calibration.md) showed 24 of 30 tasks saturated: single-file idiom tasks that frontier models solve from memory. The only tasks they missed were graded by executing real behaviour. 3.1 changes the oracle, not the prompt style.

What changed

  • **Real nixpkgs as the oracle.** A full nixpkgs tree pinned at revision

774debe7a0d1b496e35677ad955a1011c6ff74f3 (2026-10-02) is referenced by content hash in vendor/nixpkgs/pin.json, exposed to evaluators as NIXBENCH_NIXPKGS, and bound to <nixpkgs> on the agent's NIX_PATH inside the sandbox. Evaluators grade by evaluating complete NixOS configurations (config.system.build.toplevel) or real package expressions through the real package set. No fake package sets, no builds.

  • **Scale and discovery.** Every new task is a multi-file repository (14 to 40

files; two or more hosts, sites, or target platforms). The constraint the agent must honour lives in the repository, not in the prompt. Evaluators compare against an evaluator-owned baseline copy of the starter and vary hosts, profiles, and inputs.

  • **Fresh material.** Nine tasks are built on nixpkgs changes that a model

trained before mid-2026 is unlikely to know, found in the pinned tree (most through its 26.05 release notes) and verified by paired old/new evaluation (fresh-nixpkgs-changes-2026-10.md). Only some of those changes carry dated post-cutoff evidence. The task list is in research-derived-tasks.md.

  • **Corpus.** Fifteen new tasks. The 24 saturated 3.0 tasks moved to

archive/3.0/ with deprecation records. Six 3.0 tasks remain, including the repaired editor-project-state-isolation.

  • **Evaluator timeout.** Full-system tasks declare 120 seconds; observed

runtimes are 4 to 29 seconds.

Lifecycle

Every task is calibrating. Activation requires the calibration minimums in benchmark-governance.md and a saturation review.

Compatibility

3.0.0 and 2.0.0 trials stay bound to their digests and are never pooled with 3.1.0 results.