Skip to content

Documentation

Running agents

NixBench accepts any agent that can be invoked as a shell command. Generic

Maintained in docs/running-agents.md

Documentation

NixBench accepts any agent that can be invoked as a shell command. Generic commands are suitable for local runs. Publishable protocols also require a trusted completion adapter.

The harness copies a task starter into a temporary directory and runs the agent command with that temporary directory as the current working directory. The agent should read NIXBENCH_PROMPT.md, edit local files, and exit.

The agent environment includes NIXBENCH_TASK_ID, NIXBENCH_WORKDIR, and NIXBENCH_PROMPT. It does not include the original task directory, hidden evaluator path, reference solution path, or score file path. When a complete protocol requires attestation, only the registered adapter receives NIXBENCH_AGENT_STATUS_FILE. It removes that path and every other harness-private path from the model-controlled child environment.

Generic Pattern

python3 bench.py run-all \
  --protocol-file protocols/example.toml \
  --wrapper-prompt-file protocols/agent-wrapper.txt \
  --agent-timeout-seconds 240 \
  --agent-adapter codex-json \
  --agent-cmd 'your-agent-command'

For one task:

python3 bench.py run package-stdenv-cli \
  --agent-timeout-seconds 240 \
  --agent-cmd 'your-agent-command'

protocols/example.toml is a template. Copy it and set its system, timeout, model, harness, effort, network, isolation, tool policy, and registered adapter to the values used by the run. The timeout, system, and adapter must match the command-line values. The harness hashes the wrapper and agent command. For a registered adapter it records both the legacy entry-point executable digest and a canonical digest of the adapter's declared local trust bundle. It stores the hashes, not the raw command, in publication metadata.

Repeated studies

A publishable comparison should repeat the entire corpus. --trials writes every trial as a normal run and also writes a study summary containing the mean, observed range, standard deviation, and Student's t 95% confidence interval:

python3 bench.py run-all \
  --trials 5 \
  --protocol-file protocols/example.toml \
  --wrapper-prompt-file protocols/agent-wrapper.txt \
  --model gpt-5.6-sol \
  --series gpt56Sol \
  --effort high \
  --kind codex \
  --marker SH \
  --label "GPT-5.6 Sol via Codex CLI" \
  --agent-version "$(codex --version)" \
  --network unknown \
  --agent-timeout-seconds 240 \
  --agent-adapter codex-json \
  --agent-cmd 'your-agent-command'

Each attempt remains available at results/<run-id>/summary.json. The combined study is checkpointed after every attempt at results/studies/<study-id>/summary.json, so a later infrastructure or quota failure does not discard earlier evidence. Only complete, valid attempts enter the trials list and estimates. A single trial intentionally has no confidence interval because one observation cannot estimate uncertainty.

For a resumable configuration matrix, query completed evidence with the same protocol, wrapper, command, timeout, and corpus before scheduling another trial:

python3 bench.py --results-dir results study-count \
  --protocol-file protocols/example.toml \
  --wrapper-prompt-file protocols/agent-wrapper.txt \
  --agent-timeout-seconds 240 \
  --agent-adapter codex-json \
  --agent-cmd 'your-agent-command'

scripts/run-current-studies.sh uses this checkpoint to visit each current model/effort configuration once per round and resume until every configuration reaches the requested trial count.

To produce the checked JSON consumed by the website after running one or more fully labelled studies:

python3 bench.py --results-dir results export-site \
  --release-manifest corpus/releases/2.0.0.json \
  --task-count 29 \
  --minimum-trials 5 \
  --expected-configurations 14 \
  --merge-existing \
  --output site/src/data/benchmark-trials.json

export-site is a presentation transform over checked publication evidence, not an independent validator. Whenever any selected study uses the current protocol, --release-manifest is mandatory. The exporter invokes the canonical schema-3 current-study validator and publication-check policy before grouping or reporting any study. It fails without modifying the output when the release, corpus task matrix, primitive observations, controlled protocol identity, completion evidence, or publication policy does not match. It also retains the site-specific minimum-trial and expected-current-configuration gates. Zero-trial attempt ledgers are skipped so an infrastructure failure does not hide valid sibling studies or block resumption. Private-heldout and retired studies remain ineligible for direct public-site export. Historical studies require the explicit --allow-legacy-protocol compatibility flag and never count toward --expected-configurations.

When the local results archive contains only newly collected studies, merge those checked rows into the existing site dataset instead of replacing prior evidence:

python3 bench.py --results-dir results export-site \
  --release-manifest corpus/releases/2.0.0.json \
  --task-count 29 \
  --minimum-trials 1 \
  --expected-configurations 2 \
  --merge-existing \
  --output site/src/data/benchmark-trials.json

The trial and configuration gates apply to the studies being imported; --merge-existing then replaces matching row IDs and preserves unrelated checked rows. Keep the original run and study summaries in the results archive so an imported row remains independently auditable.

Codex

Example:

python3 bench.py run-all \
  --protocol-file protocols/example.toml \
  --wrapper-prompt-file protocols/agent-wrapper.txt \
  --agent-timeout-seconds 240 \
  --agent-adapter codex-json \
  --agent-cmd 'codex exec --json --ephemeral --skip-git-repo-check --sandbox workspace-write'

Notes:

  • --ephemeral avoids persistent session noise.
  • --skip-git-repo-check is useful because task workdirs are temporary copies.
  • --sandbox workspace-write allows editing local starter files.
  • --json provides native events for the trusted adapter. The adapter records

thread.started, turn.completed, native error events, and the launcher exit code. It does not search logs for success phrases.

The codex-json adapter runs Codex with the operator's own HOME, so the agent sees the operator's ~/.codex skills, AGENTS.md, plugins, and MCP servers. Use a fresh-home adapter, described next, for clean-profile comparisons.

Fresh-home Codex and Claude Code

The fresh-home adapters give every task a new HOME that contains only login material copied from the operator's home. The agent sees no user skills, AGENTS.md or CLAUDE.md, plugins, MCP servers, hooks, prompt templates, or session history. One launcher, scripts/fresh-home-agent-adapter.py, serves all four adapters:

AdapterAgentIsolation profileWhere the agent runs
codex-json-fresh-home-bwrapCodex CLIlinux-bwrap-fresh-home-v1bubblewrap; fresh home at /home/agent
claude-json-bwrapClaude Codelinux-bwrap-fresh-home-v1bubblewrap; fresh home at /home/agent
codex-json-fresh-homeCodex CLIlocal-workspace-fresh-homehost, with fresh HOME and TMPDIR
claude-jsonClaude Codelocal-workspace-fresh-homehost, with fresh HOME and TMPDIR

Prefer the bubblewrap adapters. They have been tested end to end with real Codex and Claude Code runs. The local adapters are a fallback for hosts without unprivileged bubblewrap. They clean the profile, but they are not a filesystem boundary: the agent can still read the host home if it tries.

The fresh home is created under /var/tmp/nixbench-home-<random>/home with mode 0700 and removed when the launcher exits. If the runner kills a timed-out launcher, the next launcher deletes leftover homes whose owner process has exited. The home contains:

  • Codex: a copy of ~/.codex/auth.json, plus a config.toml holding only

login keys (cli_auth_credentials_store, forced_login_method, forced_chatgpt_workspace_id, chatgpt_base_url, preferred_auth_method) when the host config sets them.

  • Claude Code: ~/.claude/.credentials.json filtered to its claudeAiOauth

entry (MCP OAuth tokens are dropped), plus an empty settings.json.

Local runs receive a minimal environment: PATH, locale, USER/LOGNAME, NIX_PATH, certificate variables, the public NIXBENCH_* task variables, the fresh HOME, and a per-task TMPDIR. Host agent variables such as CLAUDE_EFFORT, CLAUDE_CODE_*, CODEX_HOME, XDG_*, ANTHROPIC_*, and OPENAI_* are removed. The per-task TMPDIR also moves Claude Code's cross-session messaging socket away from /tmp/cc-socks, so an agent cannot reach the operator's live Claude sessions or a sibling benchmark run.

Required agent flags

The launcher validates the command shape. Codex must use exec --json. Claude Code must use -p --output-format stream-json --verbose. The prompt is passed to Codex as its final argument and to Claude Code on stdin. The variadic --disallowedTools list can therefore be the last option.

codex exec --json --ephemeral --skip-git-repo-check --sandbox workspace-write \
  --disable apps -c 'web_search="disabled"' \
  -m gpt-6-astra -c 'model_reasoning_effort="medium"'

claude -p --output-format stream-json --verbose --model claude-sonnet-5-5 \
  --setting-sources user --strict-mcp-config --no-session-persistence \
  --permission-mode bypassPermissions --disallowedTools WebFetch WebSearch
  • --disable apps turns off the ChatGPT-account app connectors that Codex

otherwise exposes over MCP. Codex's bundled system skills stay enabled because they ship with Codex.

  • --setting-sources user loads only the empty fresh settings.json, and

--strict-mcp-config without --mcp-config loads no MCP servers. Claude Code's built-in skills and built-in plugins remain, as in a stock install. No --effort is passed, so Claude Code uses each model's default effort.

  • Web search and fetch are disabled for both agents, so neither reads outside

the workspace through a hosted tool.

Completion attestation

For Codex, thread.started marks the session as started. turn.completed marks completion. turn.failed and non-retry error events are transport errors. Reconnecting... n/m error events are not fatal by themselves. A recovered stream still has to reach turn.completed, and an unrecovered one ends in turn.failed.

For Claude Code, the system/init event marks the session as started. The launcher rejects the run with a transport error when that event lists any MCP server or a plugin whose path is not builtin. It also rejects a model that differs from --model. Completion requires a single result event with subtype = "success" and is_error = false.

A non-zero agent exit is always a transport error. Bubblewrap runs record the namespace preflight evidence linux-bwrap-fresh-home-v1:generic-surfaces-absent,neutral-mounts,workspace-writable,nix-daemon-absent,auth-only-home in the status file and in the study's isolation_preflight metadata.

The linux-bwrap-fresh-home-v1 namespace

launchers/linux-bwrap-fresh-home-v1.toml declares the policy. The bundle digest covers the launcher, nixbench/agent_home.py, and that file. The profile follows linux-bwrap-v1 with these differences:

  • The fresh home is bind-mounted read-write at /home/agent. /home itself

is a tmpfs, so the host home is absent. The preflight checks that /home/agent comes from a nixbench-home-* staging path and that /home contains nothing else.

  • /run/current-system is a symlink into the read-only store, and PATH is

/run/nixbench:/run/current-system/sw/bin:/usr/bin:/bin. Agents therefore have the system tools, including nix-instantiate.

  • Nix-store executables, such as Home Manager wrapper scripts, run in place.

Other executables are staged read-only at /run/nixbench/agent.

  • /etc contains only synthetic passwd and group files and the CA bundle.

When the network is enabled, it also contains resolv.conf and hosts.

  • With network_policy = "enabled", the namespace shares the host network.

The agent can reach loopback services such as a local model router.

There is no Nix daemon socket. Nix falls back to a private chroot store under /home/agent/.local/share/nix/root. Pure --eval works. Builds must substitute or build into that empty store. Local runs use the host daemon instead, so their timings are not comparable with bubblewrap runs.

Credentials that the agent executable reads from outside the home must be named explicitly. NIXBENCH_AGENT_AUTH_FILES takes a colon-separated list of absolute files under /run, /etc, or /var. Each file is bound read-only at its host path. NIXBENCH_AGENT_AUTH_ENV takes a comma-separated list of environment variable names to copy into the cleared environment. Neither is part of the configuration identity. The adapter logs the names it forwards, never the values.

On a host where codex and claude are wrappers that read /run/secrets/cliproxyapi-local-api-key, run with NIXBENCH_AGENT_AUTH_FILES=/run/secrets/cliproxyapi-local-api-key, or pass --auth-file to the matrix script, which detects such references.

Concurrent study matrix

scripts/run_study_matrix.py runs these five configurations, one worker process each, with trials sequential inside each worker:

KeyAgentModelEffortSeries
gpt-6.1-solCodex CLIgpt-6.1-solmediumgpt61Sol
gpt-6-lunaCodex CLIgpt-6-lunamediumgpt6Luna
gpt-6-astraCodex CLIgpt-6-astramediumgpt6Astra
claude-opus-5-5Claude Codeclaude-opus-5-5defaultclaudeOpus55
claude-sonnet-5-5Claude Codeclaude-sonnet-5-5defaultclaudeSonnet55

The protocols are in protocols/fresh-home-2026-10/{bwrap,local}/<key>.toml. They use a 300-second agent timeout.

# Inspect commands and completed trial counts first.
python3 scripts/run_study_matrix.py --dry-run \
  --auth-file /run/secrets/cliproxyapi-local-api-key
python3 scripts/run_study_matrix.py --status \
  --auth-file /run/secrets/cliproxyapi-local-api-key

# Run all five configurations, three trials each, in bubblewrap.
python3 scripts/run_study_matrix.py \
  --trials 3 --tasks-dir tasks --results-dir results \
  --isolation bwrap \
  --auth-file /run/secrets/cliproxyapi-local-api-key

Each worker invokes bench.py run-all --trials <remaining>, so a clean run writes one study per configuration. Before each launch, it asks bench.py study-count how many valid trials exist for the exact configuration identity. Rerunning the same command therefore resumes. The resume unit is a full-corpus trial. A trial that stopped part-way is recorded as an excluded attempt and rerun from the first task. Task cells from different runs are never stitched into one trial. After --max-failed-attempts infrastructure aborts (default 3), a worker stops instead of spending more quota.

Runs, studies, and temporary directories use unique random names, so concurrent workers do not collide. A per-configuration lock in results/study-logs/ keeps two matrix invocations from running the same configuration at once. Logs are written to results/study-logs/matrix-<isolation>-<key>.log. Use --only <key> to run a subset, and --isolation local for the fallback adapters.

The configuration identity includes the agent command, with the absolute codex or claude path, and the adapter bundle digest. Editing the launcher, nixbench/agent_home.py, or the launcher policy during a matrix starts a new configuration identity. Its trial count starts at zero.

Each task receives a copy of the host login. A token refreshed inside a run is not written back to the host, so refresh the host login before a long matrix. The matrix script warns when ~/.codex/auth.json is more than six days old.

Agent prompt contract

Good benchmark prompts for agents should include:

  • Read NIXBENCH_PROMPT.md.
  • Modify only files in the current directory.
  • Do not inspect hidden evaluator files.
  • Run local checks if useful.
  • Exit when done.

Avoid telling the agent the hidden test path.

Timeouts

The agent timeout is controlled separately from task evaluator timeout:

python3 bench.py run-all \
  --protocol-file protocols/example.toml \
  --wrapper-prompt-file protocols/agent-wrapper.txt \
  --agent-timeout-seconds 240 \
  --agent-adapter codex-json \
  --agent-cmd '...'

Task evaluator timeouts are set in each task's metadata.toml:

timeout_seconds = 60

After each agent or evaluator command finishes, the harness terminates any remaining processes in that command's process group; it does the same immediately when a timeout expires. This prevents ordinary background children from continuing into the evaluator or later tasks. Commands that deliberately detach into a separate session still require external sandboxing; the harness is not a container or VM boundary.

Held-out isolation

Public development runs may use the provisional codex-json adapter. A private held-out publication must use a protocol with:

isolation_profile = "linux-bwrap-v1"
agent_adapter = "codex-json-bwrap"
network_policy = "disabled" # or "enabled", when the protocol requires it

Select the same adapter on the command line. The trusted launcher constructs the bubblewrap process itself. A raw command cannot claim this profile. The launcher mounts only the copied workspace read-write, creates a fresh home and /tmp, mounts required system paths read-only, and omits the repository, corpus, evaluator, reference, results, host home, and Nix daemon socket.

The trusted codex-json-bwrap bundle is exactly:

  • scripts/bwrap-codex-agent.py, the outer launcher and attestation owner;
  • nixbench/isolation.py, which constructs the namespace and preflight; and
  • launchers/linux-bwrap-v1.toml, which declares the reviewed policy.

Changing any member changes the adapter bundle digest and therefore the configuration identity. Held-out publication requires the study metadata, current adapter registration, and schema-3 release manifest to agree on that digest. The bundle does not cover provider-controlled remote code, model weights, or the external Codex/model executable.

Before starting Codex, a generic in-namespace preflight checks that the Nix daemon socket is absent and /workspace is writable. The launcher uses a fixed PATH and does not put forbidden host paths in the model process's arguments or environment. The cleared inner command runs as PID 1 so the outer environment is not visible through /proc/1/environ. The outer launcher owns the attestation path and records the preflight and Codex JSON completion state. After Codex exits, the runner rejects workspace symlinks that resolve outside /workspace before it runs the evaluator. publication-check rejects the study if this evidence is missing or failed.

Held-out workspaces use /tmp/nixbench-isolated-<random>/work on the host. The staging path contains no task, corpus, run, model, or configuration identity. The preflight verifies that neutral source shape against /proc/self/mountinfo before starting the model command. An agent executable outside the workspace is copied to /tmp/nixbench-isolated-agent-<random>/agent before its read-only bind, so its original host path is not exposed either.

Keeping workdirs

Use --keep-workdir when debugging a run:

python3 bench.py run lang-attrsets-normalize \
  --keep-workdir \
  --agent-cmd '...'

The resulting result.json will include the workdir path. Without --keep-workdir, temporary task workdirs are removed after evaluation.

Reading results

The fastest way to inspect a run:

python3 - <<'PY'
import json
from pathlib import Path

summary = json.loads(Path("results/<run-id>/summary.json").read_text())
print(f"{summary['passed']}/{summary['passed'] + summary['failed']} passed")
for task in summary["tasks"]:
    print(task["task_id"], task["passed"], task["score"])
PY

Then inspect failed tasks:

sed -n '1,160p' results/<run-id>/<task-id>/check.log
sed -n '1,220p' results/<run-id>/<task-id>/diff.patch