This study collects the first calibration evidence for corpus 3.2.0. Five configurations run the full 21-task corpus three times each in fresh-home bubblewrap sandboxes, with the same configurations and protocol as the 3.1 calibration. Every 3.2 task is calibrating, so these results are calibration evidence, not leaderboard claims. They feed calibration-report and the activation review described in Benchmark governance.
Study design
| Key | Agent | Model | Effort | Series |
|---|---|---|---|---|
gpt-6.1-sol | Codex CLI | gpt-6.1-sol | medium | gpt61Sol |
gpt-6-luna | Codex CLI | gpt-6-luna | medium | gpt6Luna |
gpt-6-astra | Codex CLI | gpt-6-astra | medium | gpt6Astra |
claude-opus-5-5 | Claude Code | claude-opus-5-5 | default | claudeOpus55 |
claude-sonnet-5-5 | Claude Code | claude-sonnet-5-5 | default | claudeSonnet55 |
Agent versions are recorded in each study's metadata and will be copied here with the results.
- Trials: 3 full-corpus trials per configuration, 15 in total. Three
configurations with three valid observations per task is the minimum for activation review.
- Corpus:
nixbench-public3.2.0, 21 tasks, digest
f783a51cb0ab545778cf74f03dbe97c4e2a6025b0b7ed293eed6e35c224ba380 (python3 bench.py corpus-id --json at commit 3922f8985c3137f449e4eddc8f4500b8ab1b2b2d; corpus commit c1831d75b56dc82eb60ecb8f1a9d3dea66e8aec0). Confirm the digest of the studies before analysis; any later change under tasks/ or contracts/ produces a different corpus identity.
- Pinned nixpkgs: revision
774debe7a0d1b496e35677ad955a1011c6ff74f3, NAR
hash sha256-nQyFkMR78WP6PiXKkDW2SjqzLf/spQqjRuxRGSbcV8k= (PIN.md).
- Isolation:
linux-bwrap-fresh-home-v1, through the
codex-json-fresh-home-bwrap and claude-json-bwrap adapters. Each task gets a fresh home with login material only: no user skills, AGENTS.md or CLAUDE.md, settings, plugins, MCP servers, or session history. Network is enabled for the model connection. There is no Nix daemon socket, and NIX_PATH binds <nixpkgs> to the pinned tree in the read-only store. Agents can evaluate but cannot build or boot VMs.
- Tools: web search and web fetch are disabled for both agents. Codex runs
with --disable apps. Claude Code runs with --setting-sources user and --strict-mcp-config, and with no --effort flag, so each model uses its default effort.
- Agent timeout: 300 seconds per task for every configuration. Evaluator
timeouts come from task metadata: 600 seconds for the three VM-track tasks, 120 seconds for the thirteen other full-system tasks, 60 seconds for the five tasks that need only lib.
- VM track:
boot-hook-marker-survives-reboot,
firewall-nftables-forwarding-silently-dropped, and state-directory-ownership-after-hardening build and boot NixOS VM tests on the evaluator host. That host needs /dev/kvm and a Nix daemon with binary-cache access; otherwise those evaluators exit 2 and the observations are invalid. Five workers running concurrently can start up to ten VM test builds at once (baseline and candidate per task).
- Protocols:
protocols/fresh-home-2026-10/bwrap/<key>.toml. Wrapper prompt:
protocols/agent-wrapper.txt.
- Runner:
scripts/run_study_matrix.py
with --trials 3 --isolation bwrap. One worker per configuration runs concurrently; trials run sequentially inside each worker.
The exact agent commands, flags, and the namespace policy are documented in Running Agents. A trial that stops part-way is recorded as an excluded attempt and rerun in full. Task cells from different runs are never stitched into one trial.
Model IDs for Codex configurations are recorded with model_identity_evidence = "vendor-api-direct". Claude Code configurations use router-alias: the recorded ID names the requested alias, not an independent attestation of the upstream model.
Analysis plan
When all five configurations have three valid trials:
python3 bench.py study-matrix --studies-dir results/studies \
--corpus-digest f783a51cb0ab545778cf74f03dbe97c4e2a6025b0b7ed293eed6e35c224ba380 \
--json
python3 bench.py calibration-report --studies-dir results/studies
Report per configuration the task pass rate with its Wilson interval, pass@k and pass^k, timeouts, and invalid attempts. Report per task the pooled solve rate, empirical difficulty, and failure set. List every task labelled saturated (pooled solve rate of 0.9 or more) for the activation review. Discrimination needs 20 valid observations across two configurations; with 15 observations per task it is expected to be unavailable at the default threshold.
Questions this run should answer:
- Do silent-failure tasks, whose starters evaluate cleanly, move the top
configurations off the ceiling? The 3.1 run had nine of twenty-one tasks saturated and the top four configurations tied at 92 to 95 percent.
- Do the six eval-track silent-failure tasks and the three VM-track tasks
separate configurations differently?
- Does
initrd-hook-systemd-stage1-migration, which every configuration
failed in 3.1, stay unsolved?
- Are any VM-track measurements invalid (evaluator exit
2) because of KVM,
binary-cache, or load problems on the runner, and do VM-track evaluator runtimes stay below 80 percent of their 600-second timeout under five concurrent workers?
Results
All five configurations completed three valid trials of twenty-one tasks: 315 valid task-trial observations, zero timeouts, zero invalid or incomplete attempts, including the three VM-graded tasks. Corpus digest f783a51cb0ab545778cf74f03dbe97c4e2a6025b0b7ed293eed6e35c224ba380. Study summaries are archived under evidence/2026-10-04-nixbench-3.2-calibration/ and the site reads site/src/data/study-matrix.json.
Configurations
| Configuration | Passed | Pass rate | Wilson 95% | pass^3 | Timeouts | Agent s/task | Study |
|---|---|---|---|---|---|---|---|
| GPT-6 Astra via Codex CLI | 52/63 | 82.5% | 71–90% | 76% | 0 | 100 s | 20261004T145807Z-c450ef1c |
| GPT-6.1 Sol via Codex CLI | 52/63 | 82.5% | 71–90% | 76% | 0 | 121 s | 20261004T145808Z-5f75e815 |
| Claude Opus 5.5 via Claude Code | 51/63 | 81.0% | 70–89% | 81% | 0 | 61 s | 20261004T145808Z-54bfd030 |
| Claude Sonnet 5.5 via Claude Code | 47/63 | 74.6% | 63–84% | 62% | 0 | 34 s | 20261004T145807Z-6afd814c |
| GPT-6 Luna via Codex CLI | 20/63 | 31.7% | 22–44% | 19% | 0 | 69 s | 20261004T145807Z-6a2a2703 |
Astra, Sol, and Opus 5.5 overlap. Sonnet 5.5 sits below Opus 5.5 by four observations with overlapping intervals; the gap is suggestive, not established. GPT-6 Luna is separated from all four. Intervals are descriptive because cells of the same task are not independent.
Tasks
Pooled solve rate over all 15 observations per task. Discrimination is the leave-one-task-out point-biserial at the exploratory threshold of 15 observations.
| Task | Pooled | Band | Discrimination | Configurations below 1.00 |
|---|---|---|---|---|
boot-hook-marker-survives-reboot | 0.80 | easy | 0.57 | GPT-6 Luna 0.33, Claude Sonnet 5.5 0.67 |
devshell-tooling-contract | 0.67 | mixed | 0.67 | GPT-6 Luna 0.00, GPT-6 Astra 0.67, GPT-6.1 Sol 0.67 |
dovecot-structured-settings-versions | 0.80 | easy | 0.97 | GPT-6 Luna 0.00 |
esphome-persistence-wrong-state-path | 0.87 | easy | -0.10 | Claude Sonnet 5.5 0.33 |
firewall-nftables-forwarding-silently-dropped | 1.00 | saturated | n/a | none |
fleet-shared-profile-conflict | 0.67 | mixed | 0.71 | GPT-6 Luna 0.00, Claude Sonnet 5.5 0.33 |
hardened-profile-import-no-op | 0.13 | hard | -0.49 | GPT-6 Luna 0.33, GPT-6 Astra 0.00, Claude Sonnet 5.5 0.33, GPT-6.1 Sol 0.00, Claude Opus 5.5 0.00 |
initrd-hook-systemd-stage1-migration | 0.00 | hard | n/a | GPT-6 Luna 0.00, GPT-6 Astra 0.00, Claude Sonnet 5.5 0.00, GPT-6.1 Sol 0.00, Claude Opus 5.5 0.00 |
lazy-type-error-boundary | 0.73 | mixed | 0.84 | GPT-6 Luna 0.00, Claude Sonnet 5.5 0.67 |
module-generated-jobs-fixed-point | 0.80 | easy | 0.97 | GPT-6 Luna 0.00 |
overlay-finalattrs-reoverride | 0.87 | easy | 0.65 | GPT-6 Luna 0.33 |
python-package-set-override-fleet | 0.80 | easy | 0.98 | GPT-6 Luna 0.00 |
rust-no-network-build | 0.80 | easy | 0.22 | GPT-6 Luna 0.67, GPT-6 Astra 0.67, GPT-6.1 Sol 0.67 |
service-requires-removed-network-setup | 0.80 | easy | -0.30 | Claude Opus 5.5 0.00 |
stalwart-state-version-migration | 0.80 | easy | 0.98 | GPT-6 Luna 0.00 |
state-directory-ownership-after-hardening | 1.00 | saturated | n/a | none |
stdenv-cross-and-host-tools | 0.80 | easy | 0.98 | GPT-6 Luna 0.00 |
tor-onion-socket-outside-chroot | 0.00 | hard | n/a | GPT-6 Luna 0.00, GPT-6 Astra 0.00, Claude Sonnet 5.5 0.00, GPT-6.1 Sol 0.00, Claude Opus 5.5 0.00 |
vsftpd-local-login-without-pam | 0.93 | saturated | 0.64 | GPT-6 Luna 0.67 |
wireless-eap-key-unreadable-after-hardening | 0.73 | mixed | 0.54 | GPT-6 Luna 0.33, Claude Sonnet 5.5 0.33 |
yggdrasil-credential-key-owner | 0.80 | easy | 0.97 | GPT-6 Luna 0.00 |
Findings
- **The ceiling is gone.** The strongest configurations dropped from 95% on
3.1 to 82.5%, pass^3 from 95% to 76–81%, and only three of twenty-one tasks are saturated. Luna fell to 31.7%. Silent-failure tasks did what evaluation errors could not: agents cannot loop on nix-instantiate when the starter already evaluates cleanly.
- **Three evaluator disputes, all rejecting valid alternatives.** Auditing
every task with zero or negative discrimination found that tor-onion-socket-outside-chroot demands Tor be ordered after the backend (not required by the prompt or by runtime when the socket directory comes from tmpfiles), hardened-profile-import-no-op rejects candidates that apply module locking and SMT disabling fleet-wide although its own baseline document presents those as baseline controls with a recorded exception, and service-requires-removed-network-setup only recognises a two-word Match User block and rejects the stricter Match User ... LocalAddress form. All three are quarantined. Their pooled solve rates (0.00, 0.13, 0.80) understate the agents; repair and recalibration are required before any activation. The repairs are small (drop or justify the ordering criterion; compare module-locking against the documented baseline rather than the starter; parse Match criteria as key-value pairs).
- **
initrd-hook-systemd-stage1-migrationheld again** at 0 of 15 for the same
reason as on 3.1: every attempt migrated the hook and none set the LUKS root to the mapped device as the pinned release notes require. It remains the clearest fresh-material discriminator in the corpus.
- **The VM track works operationally.** Three VM-graded tasks produced 45
valid observations with no infrastructure failures. Two of them saturated (state-directory-ownership-after-hardening, firewall-nftables-forwarding-silently-dropped); the reboot-marker task separated Sonnet and Luna. VM grading is not harder by itself; it is harder only when the runtime requirement is not inferable from the evaluation errors the agent can see.
- **Cost axis.** Claude Code used 34–61 agent seconds per task; Codex used
69–121 seconds. Sonnet is the cheapest configuration by a factor of three at a pass rate eight points below Astra.
Decisions
- The release stays
calibrating. No task is activated. tor-onion-socket-outside-chroot,hardened-profile-import-no-op, and
service-requires-removed-network-setup move to quarantined with the dispute recorded in the calibration registry.
- All other records are
pending.
Next steps
- Repair the three quarantined evaluators as described above, add passing
fixtures built from the rejected trial diffs, and recalibrate them.
- Retire or couple the three saturated tasks.
- Audit every task with pooled solve rate below 0.25 by re-scoring strong
models' diffs before each calibration; two of the three disputes here would have been invisible from aggregate numbers alone.
- Run five trials per configuration before an activation decision.