This study collects the first calibration evidence for corpus 3.0.0. Five configurations run the full 30-task corpus three times each in fresh-home bubblewrap sandboxes. Every 3.0 task is calibrating, so these results are calibration evidence, not leaderboard claims. They feed calibration-report and the activation review described in Benchmark governance.
Study design
| Key | Agent | Model | Effort | Series |
|---|---|---|---|---|
gpt-6.1-sol | Codex CLI 0.160.0 | gpt-6.1-sol | medium | gpt61Sol |
gpt-6-luna | Codex CLI 0.160.0 | gpt-6-luna | medium | gpt6Luna |
gpt-6-astra | Codex CLI 0.160.0 | gpt-6-astra | medium | gpt6Astra |
claude-opus-5-5 | Claude Code 2.1.288 | claude-opus-5-5 | default | claudeOpus55 |
claude-sonnet-5-5 | Claude Code 2.1.288 | claude-sonnet-5-5 | default | claudeSonnet55 |
- Trials: 3 full-corpus trials per configuration, 15 in total. Three
configurations with three valid observations per task is the minimum for activation review.
- Corpus:
nixbench-public3.0.0, 30 tasks, digest
0d4e05063591d17d598b9620543b57f03224db96271d403f03a7b56c25da9049 (python3 bench.py corpus-id --json). Git revision a20499aba85381fd518ab7a8c3688ed990bd8366.
- Isolation:
linux-bwrap-fresh-home-v1, through the
codex-json-fresh-home-bwrap and claude-json-bwrap adapters. Each task gets a fresh home with login material only: no user skills, AGENTS.md or CLAUDE.md, settings, plugins, MCP servers, or session history. Network is enabled for the model connection. There is no Nix daemon socket, and the vendored nixpkgs lib is on NIX_PATH as <nixpkgs/lib>.
- Tools: web search and web fetch are disabled for both agents. Codex runs
with --disable apps. Claude Code runs with --setting-sources user and --strict-mcp-config, and with no --effort flag, so each model uses its default effort.
- Agent timeout: 300 seconds per task for every configuration. Evaluator
timeouts come from task metadata (60 seconds).
- Protocols:
protocols/fresh-home-2026-10/bwrap/<key>.toml. Wrapper prompt:
protocols/agent-wrapper.txt.
- Runner:
scripts/run_study_matrix.py
with --trials 3 --isolation bwrap. One worker per configuration runs concurrently; trials run sequentially inside each worker. The matrix started at 2026-10-03T22:48:05Z.
The exact agent commands, flags, and the namespace policy are documented in Running Agents. A trial that stops part-way is recorded as an excluded attempt and rerun in full. Task cells from different runs are never stitched into one trial.
Model IDs for Codex configurations are recorded with model_identity_evidence = "vendor-api-direct". Claude Code configurations use router-alias: the recorded ID names the requested alias, not an independent attestation of the upstream model.
Analysis plan
When all five configurations have three valid trials:
python3 bench.py study-matrix --studies-dir results/studies \
--corpus-digest 0d4e05063591d17d598b9620543b57f03224db96271d403f03a7b56c25da9049 \
--json
python3 bench.py calibration-report --studies-dir results/studies
Report per configuration the task pass rate with its Wilson interval, pass@k and pass^k, timeouts, and invalid attempts. Report per task the pooled solve rate, empirical difficulty, and failure set. List every task labelled saturated (pooled solve rate of 0.9 or more) for the activation review. Discrimination needs 20 valid observations across two configurations; with 15 observations per task it is expected to be unavailable.
Results
All five configurations completed three valid trials of thirty tasks: 450 valid task-trial observations, zero timeouts, zero invalid or incomplete attempts. Corpus digest 0d4e05063591d17d598b9620543b57f03224db96271d403f03a7b56c25da9049. Study summaries are archived under evidence/2026-10-04-nixbench-3-calibration/ and the site reads site/src/data/study-matrix.json.
Configurations
| Configuration | Passed | Pass rate | Wilson 95% | pass^3 | Timeouts | Agent s/task | Study |
|---|---|---|---|---|---|---|---|
| Claude Opus 5.5 via Claude Code | 89/90 | 98.9% | 94–100% | 97% | 0 | 25 s | 20261003T224805Z-b039ace8 |
| GPT-6 Astra via Codex CLI | 88/90 | 97.8% | 92–99% | 97% | 0 | 81 s | 20261003T224805Z-9e44c874 |
| Claude Sonnet 5.5 via Claude Code | 88/90 | 97.8% | 92–99% | 93% | 0 | 16 s | 20261003T224805Z-c01f0b12 |
| GPT-6.1 Sol via Codex CLI | 87/90 | 96.7% | 91–99% | 97% | 0 | 77 s | 20261003T224805Z-0f369581 |
| GPT-6 Luna via Codex CLI | 74/90 | 82.2% | 73–89% | 73% | 0 | 47 s | 20261003T224805Z-ab247b82 |
The top four intervals overlap; they are not ranked against each other. GPT-6 Luna is separated from all four. Cells of the same task are not independent, so the intervals are descriptive.
Tasks
Pooled solve rate over all 15 observations per task. Discrimination is the leave-one-task-out point-biserial computed with the exploratory threshold of 15 observations (the default is 20, so the committed site data shows it as unavailable).
| Task | Pooled | Band | Discrimination | Configurations below 1.00 |
|---|---|---|---|---|
argument-default-forwarding | 1.00 | saturated | n/a | none |
debug-freeform-config-cycle | 1.00 | saturated | n/a | none |
debug-import-argument-cycle | 1.00 | saturated | n/a | none |
devshell-tooling-contract | 0.87 | easy | 0.70 | GPT-6 Luna 0.33 |
editor-project-state-isolation | 0.60 | mixed | -0.44 | GPT-6 Astra 0.33, Claude Opus 5.5 0.67, GPT-6.1 Sol 0.00 |
fetcher-fixed-output-identity | 0.93 | saturated | 0.68 | GPT-6 Luna 0.67 |
flake-nested-follows-identity | 1.00 | saturated | n/a | none |
follows-preserve-python-compatibility | 1.00 | saturated | n/a | none |
lazy-recursive-update-closure | 1.00 | saturated | n/a | none |
lazy-selected-report-validation | 1.00 | saturated | n/a | none |
lazy-type-error-boundary | 0.73 | mixed | 0.73 | Claude Sonnet 5.5 0.67, GPT-6 Luna 0.00 |
mkforce-across-option-reexport | 1.00 | saturated | n/a | none |
module-deferred-schema-default | 1.00 | saturated | n/a | none |
module-generated-jobs-fixed-point | 0.80 | easy | 0.53 | GPT-6 Luna 0.00 |
module-migration-assertion-gate | 0.93 | saturated | 0.14 | GPT-6 Luna 0.67 |
module-priority-submodule-apply | 1.00 | saturated | n/a | none |
module-service-options | 1.00 | saturated | n/a | none |
module-system-boundaries | 1.00 | saturated | n/a | none |
overlay-composed-final-prev | 1.00 | saturated | n/a | none |
overlay-finalattrs-reoverride | 0.87 | easy | 0.76 | GPT-6 Luna 0.33 |
package-offline-test-selection | 0.93 | saturated | -0.19 | Claude Sonnet 5.5 0.67 |
purity-explicit-release-inputs | 1.00 | saturated | n/a | none |
python-build-backend-false-lead | 1.00 | saturated | n/a | none |
runtime-tools-outside-devshell | 0.93 | saturated | 0.58 | GPT-6 Luna 0.67 |
rust-no-network-build | 0.80 | easy | 0.75 | GPT-6 Luna 0.00 |
scope-override-transitive-dependencies | 1.00 | saturated | n/a | none |
source-filter-traversal-stability | 1.00 | saturated | n/a | none |
string-command-dependency-context | 1.00 | saturated | n/a | none |
string-escaping-systemd | 1.00 | saturated | n/a | none |
xdg-portal-merge | 1.00 | saturated | n/a | none |
Findings
- **The corpus is still too easy for frontier models.** Twenty-four of thirty
tasks are saturated (pooled solve rate at or above 0.9). Only six tasks carry any discrimination, and five of those discriminate only because GPT-6 Luna fails them. The top four configurations differ by at most three observations out of ninety.
- **Evaluator dispute:
editor-project-state-isolation.** Its discrimination
is negative: the strongest configurations fail it while Luna passes. The evaluator's fake coreutils set (tests/probe.py, BASIC_TOOLS) did not include pwd, so any wrapper calling ${coreutils}/bin/pwd failed every launch. Re-scoring the 15 recorded trials against a repaired probe shows six were wrongly failed: five all-false trials (Opus 5.5 once, GPT-6.1 Sol twice, GPT-6 Astra twice) and one partial failure where Sol piped pwd -P into sha256sum without pipefail. All six pass the repaired evaluator. The task was quarantined for this release; the repair (complete coreutils set plus two new passing fixtures) landed afterwards and the task returns to calibrating for 3.1 under its new digest. package-offline-test-selection also has slightly negative discrimination and one Sonnet failure that should be reviewed before activation.
- **Sandbox and attestation held.** Every observation is valid and attested.
Each fresh home contained only login material. Codex authenticated directly with its auth.json; Claude Code went through the local router.
- **Cost axis.** Claude Code configurations used 16–25 agent seconds per task;
Codex configurations used 47–81 seconds at the same pass rates.
Decisions
- The release stays
calibrating. No task is activated from this study. editor-project-state-isolationmoves toquarantined.- All other calibration records are recorded with decision
pending.
Next steps
- Retire or rework the 24 saturated tasks; keep them in the archive for
regression use.
- Author the next round from the reserve pool in
task-candidates-designed.md and task-candidates-community.md, prioritising tasks that the strongest models failed here: laziness and type error boundaries, offline test selection, runtime tool wrapping, and fixed-output identity.
- Repair the quarantined evaluator and recalibrate.
- Rerun with five trials per configuration once the corpus has at least ten
tasks in the mixed band.