Skip to content

Documentation

2026-10-04 NixBench 3.1 calibration

This study collects the first calibration evidence for corpus 3.1.0. Five

Maintained in docs/runs/2026-10-04-nixbench-3.1-calibration.md

Documentation

This study collects the first calibration evidence for corpus 3.1.0. Five configurations run the full 21-task corpus three times each in fresh-home bubblewrap sandboxes, with the same configurations and protocol as the 3.0 calibration. Every 3.1 task is calibrating, so these results are calibration evidence, not leaderboard claims. They feed calibration-report and the activation review described in Benchmark governance.

Study design

KeyAgentModelEffortSeries
gpt-6.1-solCodex CLIgpt-6.1-solmediumgpt61Sol
gpt-6-lunaCodex CLIgpt-6-lunamediumgpt6Luna
gpt-6-astraCodex CLIgpt-6-astramediumgpt6Astra
claude-opus-5-5Claude Codeclaude-opus-5-5defaultclaudeOpus55
claude-sonnet-5-5Claude Codeclaude-sonnet-5-5defaultclaudeSonnet55

Agent versions are recorded in each study's metadata and will be copied here with the results.

  • Trials: 3 full-corpus trials per configuration, 15 in total. Three

configurations with three valid observations per task is the minimum for activation review.

  • Corpus: nixbench-public 3.1.0, 21 tasks, digest

7add4fb2607fb150643d2f59e0f0c614f1d41848ba7b4ca7e1771895d7afc8e1 (python3 bench.py corpus-id --json). Release commit 83a2ac7e4eb44c1b883b89c6dd4e46263896d5d6.

  • Pinned nixpkgs: revision 774debe7a0d1b496e35677ad955a1011c6ff74f3, NAR

hash sha256-nQyFkMR78WP6PiXKkDW2SjqzLf/spQqjRuxRGSbcV8k= (PIN.md).

  • Isolation: linux-bwrap-fresh-home-v1, through the

codex-json-fresh-home-bwrap and claude-json-bwrap adapters. Each task gets a fresh home with login material only: no user skills, AGENTS.md or CLAUDE.md, settings, plugins, MCP servers, or session history. Network is enabled for the model connection. There is no Nix daemon socket, and NIX_PATH binds <nixpkgs> to the pinned tree in the read-only store.

  • Tools: web search and web fetch are disabled for both agents. Codex runs

with --disable apps. Claude Code runs with --setting-sources user and --strict-mcp-config, and with no --effort flag, so each model uses its default effort.

  • Agent timeout: 300 seconds per task for every configuration. Evaluator

timeouts come from task metadata: 120 seconds for the fifteen full-system tasks, 60 seconds for the six tasks carried over from 3.0.

  • Protocols: protocols/fresh-home-2026-10/bwrap/<key>.toml. Wrapper prompt:

protocols/agent-wrapper.txt.

with --trials 3 --isolation bwrap. One worker per configuration runs concurrently; trials run sequentially inside each worker.

The exact agent commands, flags, and the namespace policy are documented in Running Agents. A trial that stops part-way is recorded as an excluded attempt and rerun in full. Task cells from different runs are never stitched into one trial.

Model IDs for Codex configurations are recorded with model_identity_evidence = "vendor-api-direct". Claude Code configurations use router-alias: the recorded ID names the requested alias, not an independent attestation of the upstream model.

Analysis plan

When all five configurations have three valid trials:

python3 bench.py study-matrix --studies-dir results/studies \
  --corpus-digest 7add4fb2607fb150643d2f59e0f0c614f1d41848ba7b4ca7e1771895d7afc8e1 \
  --json
python3 bench.py calibration-report --studies-dir results/studies

Report per configuration the task pass rate with its Wilson interval, pass@k and pass^k, timeouts, and invalid attempts. Report per task the pooled solve rate, empirical difficulty, and failure set. List every task labelled saturated (pooled solve rate of 0.9 or more) for the activation review. Discrimination needs 20 valid observations across two configurations; with 15 observations per task it is expected to be unavailable at the default threshold.

Questions this run should answer:

  • Does grading against the real pinned nixpkgs move the corpus out of

saturation? The 3.0 run had 24 of 30 tasks at or above 0.9.

  • Do the nine fresh-material tasks

(research-derived-tasks.md) separate configurations more than the six designed full-system tasks?

  • Does the repaired editor-project-state-isolation evaluator lose its

negative discrimination?

  • Are any evaluator measurements invalid because the pinned tree was missing

on the runner (evaluator exit 2)?

Results

All five configurations completed three valid trials of twenty-one tasks: 315 valid task-trial observations, zero timeouts, zero invalid or incomplete attempts. Corpus digest 7add4fb2607fb150643d2f59e0f0c614f1d41848ba7b4ca7e1771895d7afc8e1. Study summaries are archived under evidence/2026-10-04-nixbench-3.1-calibration/ and the site reads site/src/data/study-matrix.json.

Configurations

ConfigurationPassedPass rateWilson 95%pass^3TimeoutsAgent s/taskStudy
GPT-6 Astra via Codex CLI60/6395.2%87–98%95%095 s20261004T091312Z-4dffeb21
Claude Sonnet 5.5 via Claude Code60/6395.2%87–98%95%035 s20261004T091312Z-aab13662
Claude Opus 5.5 via Claude Code59/6393.7%85–98%90%058 s20261004T091312Z-94c28f0e
GPT-6.1 Sol via Codex CLI58/6392.1%83–97%90%0111 s20261004T091312Z-4ebc15c0
GPT-6 Luna via Codex CLI30/6347.6%36–60%33%073 s20261004T091312Z-0c4a08a9

The top four intervals overlap; they are not ranked against each other. GPT-6 Luna is separated from all four by a wide margin. Intervals are descriptive because cells of the same task are not independent.

Tasks

Pooled solve rate over all 15 observations per task. Discrimination is the leave-one-task-out point-biserial at the exploratory threshold of 15 observations (the default is 20, so the committed site data shows it as unavailable).

TaskPooledBandDiscriminationConfigurations below 1.00
cgit-export-policy-agreement1.00saturatedn/anone
debug-infinite-recursion-fleet1.00saturatedn/anone
devshell-tooling-contract0.80easy0.98GPT-6 Luna 0.00
dovecot-structured-settings-versions0.80easy0.96GPT-6 Luna 0.00
editor-project-state-isolation1.00saturatedn/anone
fleet-shared-profile-conflict0.67mixed0.69GPT-6 Luna 0.00, GPT-6.1 Sol 0.33
home-assistant-declarative-dashboards1.00saturatedn/anone
home-manager-nixos-module-boundary-fleet0.93saturated0.67GPT-6 Luna 0.67
initrd-hook-systemd-stage1-migration0.00hardn/aGPT-6 Luna 0.00, Claude Opus 5.5 0.00, GPT-6 Astra 0.00, GPT-6.1 Sol 0.00, Claude Sonnet 5.5 0.00
lazy-type-error-boundary0.80easy0.98GPT-6 Luna 0.00
module-generated-jobs-fixed-point0.80easy0.96GPT-6 Luna 0.00
nginx-acme-virtualhost-migration1.00saturatedn/anone
overlay-finalattrs-reoverride0.87easy0.79GPT-6 Luna 0.33
python-package-set-override-fleet0.87easy0.63GPT-6 Luna 0.33
rust-no-network-build0.87easy0.61GPT-6 Luna 0.33
rust-workspace-vendored-deps1.00saturatedn/anone
stalwart-state-version-migration0.80easy0.98GPT-6 Luna 0.00
stdenv-cross-and-host-tools0.87easy0.33GPT-6 Luna 0.67, Claude Opus 5.5 0.67
systemd-unit-ordering-and-paths1.00saturatedn/anone
users-groups-impermanence-boundary0.93saturated0.65GPT-6 Luna 0.67
yggdrasil-credential-key-owner0.80easy0.97GPT-6 Luna 0.00

Findings

  • **3.1 separates the field more than 3.0 did.** The spread between the

strongest and weakest configuration grew from 17 points on 3.0 to 48 points here, and pass^3 for the top configurations dropped from 97% to 90–95%. Twelve of twenty-one tasks are unsaturated, against six of thirty in 3.0.

  • **One task defeated every configuration.** initrd-hook-systemd-stage1-migration

scored 90 of 100 on every one of its fifteen attempts: all migrated the hook to a systemd stage-1 unit correctly, and none changed the root filesystem from /dev/disk/by-label to the mapped LUKS device. That requirement is the first bullet of the pinned 26.05 release notes, which every agent had on NIX_PATH. No agent read them. This is the fresh-material effect the round was designed to test, and the task is retained unchanged.

  • **The repaired editor task is now saturated.** After the coreutils fix,

editor-project-state-isolation solves at 1.00 everywhere, confirming the 3.0 dispute was an evaluator gap.

  • **Still too easy for the strongest models.** Nine tasks are saturated and the

top four configurations fail only one to three tasks each. Luna provides most of the discrimination. The next round should push further on scale and on constraints that only appear in post-cutoff material.

  • **Cost axis.** Claude Code used 35–58 agent seconds per task; Codex used

73–111 seconds at the same or lower pass rates.

  • **Sandbox and attestation held.** Every observation is valid and attested;

agents evaluated full NixOS systems against the pinned tree inside bubblewrap.

Decisions

  • The release stays calibrating. No task is activated from this study.
  • All calibration records are recorded with decision pending.

Next steps

  1. Review the nine saturated tasks for retirement or for a coupled constraint

that raises them out of the ceiling.

  1. Author the next round from the remaining candidates in

fresh-nixpkgs-changes-2026-10.md, favouring changes that produce no evaluation error, as the initrd task does.

  1. Rerun with five trials per configuration before any activation decision.