Skip to content

Documentation

2026-10-04 NixBench 3.2 calibration

This study collects the first calibration evidence for corpus 3.2.0. Five

Maintained in docs/runs/2026-10-04-nixbench-3.2-calibration.md

Documentation

This study collects the first calibration evidence for corpus 3.2.0. Five configurations run the full 21-task corpus three times each in fresh-home bubblewrap sandboxes, with the same configurations and protocol as the 3.1 calibration. Every 3.2 task is calibrating, so these results are calibration evidence, not leaderboard claims. They feed calibration-report and the activation review described in Benchmark governance.

Study design

KeyAgentModelEffortSeries
gpt-6.1-solCodex CLIgpt-6.1-solmediumgpt61Sol
gpt-6-lunaCodex CLIgpt-6-lunamediumgpt6Luna
gpt-6-astraCodex CLIgpt-6-astramediumgpt6Astra
claude-opus-5-5Claude Codeclaude-opus-5-5defaultclaudeOpus55
claude-sonnet-5-5Claude Codeclaude-sonnet-5-5defaultclaudeSonnet55

Agent versions are recorded in each study's metadata and will be copied here with the results.

  • Trials: 3 full-corpus trials per configuration, 15 in total. Three

configurations with three valid observations per task is the minimum for activation review.

  • Corpus: nixbench-public 3.2.0, 21 tasks, digest

f783a51cb0ab545778cf74f03dbe97c4e2a6025b0b7ed293eed6e35c224ba380 (python3 bench.py corpus-id --json at commit 3922f8985c3137f449e4eddc8f4500b8ab1b2b2d; corpus commit c1831d75b56dc82eb60ecb8f1a9d3dea66e8aec0). Confirm the digest of the studies before analysis; any later change under tasks/ or contracts/ produces a different corpus identity.

  • Pinned nixpkgs: revision 774debe7a0d1b496e35677ad955a1011c6ff74f3, NAR

hash sha256-nQyFkMR78WP6PiXKkDW2SjqzLf/spQqjRuxRGSbcV8k= (PIN.md).

  • Isolation: linux-bwrap-fresh-home-v1, through the

codex-json-fresh-home-bwrap and claude-json-bwrap adapters. Each task gets a fresh home with login material only: no user skills, AGENTS.md or CLAUDE.md, settings, plugins, MCP servers, or session history. Network is enabled for the model connection. There is no Nix daemon socket, and NIX_PATH binds <nixpkgs> to the pinned tree in the read-only store. Agents can evaluate but cannot build or boot VMs.

  • Tools: web search and web fetch are disabled for both agents. Codex runs

with --disable apps. Claude Code runs with --setting-sources user and --strict-mcp-config, and with no --effort flag, so each model uses its default effort.

  • Agent timeout: 300 seconds per task for every configuration. Evaluator

timeouts come from task metadata: 600 seconds for the three VM-track tasks, 120 seconds for the thirteen other full-system tasks, 60 seconds for the five tasks that need only lib.

  • VM track: boot-hook-marker-survives-reboot,

firewall-nftables-forwarding-silently-dropped, and state-directory-ownership-after-hardening build and boot NixOS VM tests on the evaluator host. That host needs /dev/kvm and a Nix daemon with binary-cache access; otherwise those evaluators exit 2 and the observations are invalid. Five workers running concurrently can start up to ten VM test builds at once (baseline and candidate per task).

  • Protocols: protocols/fresh-home-2026-10/bwrap/<key>.toml. Wrapper prompt:

protocols/agent-wrapper.txt.

with --trials 3 --isolation bwrap. One worker per configuration runs concurrently; trials run sequentially inside each worker.

The exact agent commands, flags, and the namespace policy are documented in Running Agents. A trial that stops part-way is recorded as an excluded attempt and rerun in full. Task cells from different runs are never stitched into one trial.

Model IDs for Codex configurations are recorded with model_identity_evidence = "vendor-api-direct". Claude Code configurations use router-alias: the recorded ID names the requested alias, not an independent attestation of the upstream model.

Analysis plan

When all five configurations have three valid trials:

python3 bench.py study-matrix --studies-dir results/studies \
  --corpus-digest f783a51cb0ab545778cf74f03dbe97c4e2a6025b0b7ed293eed6e35c224ba380 \
  --json
python3 bench.py calibration-report --studies-dir results/studies

Report per configuration the task pass rate with its Wilson interval, pass@k and pass^k, timeouts, and invalid attempts. Report per task the pooled solve rate, empirical difficulty, and failure set. List every task labelled saturated (pooled solve rate of 0.9 or more) for the activation review. Discrimination needs 20 valid observations across two configurations; with 15 observations per task it is expected to be unavailable at the default threshold.

Questions this run should answer:

  • Do silent-failure tasks, whose starters evaluate cleanly, move the top

configurations off the ceiling? The 3.1 run had nine of twenty-one tasks saturated and the top four configurations tied at 92 to 95 percent.

  • Do the six eval-track silent-failure tasks and the three VM-track tasks

separate configurations differently?

  • Does initrd-hook-systemd-stage1-migration, which every configuration

failed in 3.1, stay unsolved?

  • Are any VM-track measurements invalid (evaluator exit 2) because of KVM,

binary-cache, or load problems on the runner, and do VM-track evaluator runtimes stay below 80 percent of their 600-second timeout under five concurrent workers?

Results

All five configurations completed three valid trials of twenty-one tasks: 315 valid task-trial observations, zero timeouts, zero invalid or incomplete attempts, including the three VM-graded tasks. Corpus digest f783a51cb0ab545778cf74f03dbe97c4e2a6025b0b7ed293eed6e35c224ba380. Study summaries are archived under evidence/2026-10-04-nixbench-3.2-calibration/ and the site reads site/src/data/study-matrix.json.

Configurations

ConfigurationPassedPass rateWilson 95%pass^3TimeoutsAgent s/taskStudy
GPT-6 Astra via Codex CLI52/6382.5%71–90%76%0100 s20261004T145807Z-c450ef1c
GPT-6.1 Sol via Codex CLI52/6382.5%71–90%76%0121 s20261004T145808Z-5f75e815
Claude Opus 5.5 via Claude Code51/6381.0%70–89%81%061 s20261004T145808Z-54bfd030
Claude Sonnet 5.5 via Claude Code47/6374.6%63–84%62%034 s20261004T145807Z-6afd814c
GPT-6 Luna via Codex CLI20/6331.7%22–44%19%069 s20261004T145807Z-6a2a2703

Astra, Sol, and Opus 5.5 overlap. Sonnet 5.5 sits below Opus 5.5 by four observations with overlapping intervals; the gap is suggestive, not established. GPT-6 Luna is separated from all four. Intervals are descriptive because cells of the same task are not independent.

Tasks

Pooled solve rate over all 15 observations per task. Discrimination is the leave-one-task-out point-biserial at the exploratory threshold of 15 observations.

TaskPooledBandDiscriminationConfigurations below 1.00
boot-hook-marker-survives-reboot0.80easy0.57GPT-6 Luna 0.33, Claude Sonnet 5.5 0.67
devshell-tooling-contract0.67mixed0.67GPT-6 Luna 0.00, GPT-6 Astra 0.67, GPT-6.1 Sol 0.67
dovecot-structured-settings-versions0.80easy0.97GPT-6 Luna 0.00
esphome-persistence-wrong-state-path0.87easy-0.10Claude Sonnet 5.5 0.33
firewall-nftables-forwarding-silently-dropped1.00saturatedn/anone
fleet-shared-profile-conflict0.67mixed0.71GPT-6 Luna 0.00, Claude Sonnet 5.5 0.33
hardened-profile-import-no-op0.13hard-0.49GPT-6 Luna 0.33, GPT-6 Astra 0.00, Claude Sonnet 5.5 0.33, GPT-6.1 Sol 0.00, Claude Opus 5.5 0.00
initrd-hook-systemd-stage1-migration0.00hardn/aGPT-6 Luna 0.00, GPT-6 Astra 0.00, Claude Sonnet 5.5 0.00, GPT-6.1 Sol 0.00, Claude Opus 5.5 0.00
lazy-type-error-boundary0.73mixed0.84GPT-6 Luna 0.00, Claude Sonnet 5.5 0.67
module-generated-jobs-fixed-point0.80easy0.97GPT-6 Luna 0.00
overlay-finalattrs-reoverride0.87easy0.65GPT-6 Luna 0.33
python-package-set-override-fleet0.80easy0.98GPT-6 Luna 0.00
rust-no-network-build0.80easy0.22GPT-6 Luna 0.67, GPT-6 Astra 0.67, GPT-6.1 Sol 0.67
service-requires-removed-network-setup0.80easy-0.30Claude Opus 5.5 0.00
stalwart-state-version-migration0.80easy0.98GPT-6 Luna 0.00
state-directory-ownership-after-hardening1.00saturatedn/anone
stdenv-cross-and-host-tools0.80easy0.98GPT-6 Luna 0.00
tor-onion-socket-outside-chroot0.00hardn/aGPT-6 Luna 0.00, GPT-6 Astra 0.00, Claude Sonnet 5.5 0.00, GPT-6.1 Sol 0.00, Claude Opus 5.5 0.00
vsftpd-local-login-without-pam0.93saturated0.64GPT-6 Luna 0.67
wireless-eap-key-unreadable-after-hardening0.73mixed0.54GPT-6 Luna 0.33, Claude Sonnet 5.5 0.33
yggdrasil-credential-key-owner0.80easy0.97GPT-6 Luna 0.00

Findings

  • **The ceiling is gone.** The strongest configurations dropped from 95% on

3.1 to 82.5%, pass^3 from 95% to 76–81%, and only three of twenty-one tasks are saturated. Luna fell to 31.7%. Silent-failure tasks did what evaluation errors could not: agents cannot loop on nix-instantiate when the starter already evaluates cleanly.

  • **Three evaluator disputes, all rejecting valid alternatives.** Auditing

every task with zero or negative discrimination found that tor-onion-socket-outside-chroot demands Tor be ordered after the backend (not required by the prompt or by runtime when the socket directory comes from tmpfiles), hardened-profile-import-no-op rejects candidates that apply module locking and SMT disabling fleet-wide although its own baseline document presents those as baseline controls with a recorded exception, and service-requires-removed-network-setup only recognises a two-word Match User block and rejects the stricter Match User ... LocalAddress form. All three are quarantined. Their pooled solve rates (0.00, 0.13, 0.80) understate the agents; repair and recalibration are required before any activation. The repairs are small (drop or justify the ordering criterion; compare module-locking against the documented baseline rather than the starter; parse Match criteria as key-value pairs).

  • **initrd-hook-systemd-stage1-migration held again** at 0 of 15 for the same

reason as on 3.1: every attempt migrated the hook and none set the LUKS root to the mapped device as the pinned release notes require. It remains the clearest fresh-material discriminator in the corpus.

  • **The VM track works operationally.** Three VM-graded tasks produced 45

valid observations with no infrastructure failures. Two of them saturated (state-directory-ownership-after-hardening, firewall-nftables-forwarding-silently-dropped); the reboot-marker task separated Sonnet and Luna. VM grading is not harder by itself; it is harder only when the runtime requirement is not inferable from the evaluation errors the agent can see.

  • **Cost axis.** Claude Code used 34–61 agent seconds per task; Codex used

69–121 seconds. Sonnet is the cheapest configuration by a factor of three at a pass rate eight points below Astra.

Decisions

  • The release stays calibrating. No task is activated.
  • tor-onion-socket-outside-chroot, hardened-profile-import-no-op, and

service-requires-removed-network-setup move to quarantined with the dispute recorded in the calibration registry.

  • All other records are pending.

Next steps

  1. Repair the three quarantined evaluators as described above, add passing

fixtures built from the rejected trial diffs, and recalibrate them.

  1. Retire or couple the three saturated tasks.
  2. Audit every task with pooled solve rate below 0.25 by re-scoring strong

models' diffs before each calibration; two of the three disputes here would have been invisible from aggregate numbers alone.

  1. Run five trials per configuration before an activation decision.