NixBench

tasks
29
areas
9
evaluators
29
current trials
70

Can AI coding agents write Nix that actually passes?

Objective repository-repair tasks scored by hidden shell evaluators—not by whether the output merely looks plausible.

Current release: 29-task corpus · one hidden evaluator per task · uncertainty shown when repeat trials exist

Compare the signal first. Inspect the scatter second.

Ordered effort paths lead the view. Select a model for uncertainty and effort labels, or reveal every trial to inspect the underlying variation.

Corpus
Evidence
Y-axis
3 models · 14 configurations · 70 trials14 means shown14/14 replicated

Configuration means, with uncertainty

Mean tasks passed against mean agent seconds per task. Paths connect ordered effort configurations; select a model to reveal its 95% Student's t intervals and effort labels. The task axis focuses on 18–29 to make the observed differences legible. Individual trials are hidden in this summary view.

Focused: 1829 tasks↑ more tasks← less time
95% CI / observed
GPT-5.6 Sol via Codex CLI29-task corpus · 5 recorded trialsmediumn=524.0 / 29
22.5–25.5observed 222538.5s18m 36s / corpus0
GPT-5.6 Sol via Codex CLI29-task corpus · 5 recorded trialsxhighn=524.0 / 29
23.1–24.9observed 232560.7s29m 20s / corpus0
GPT-5.6 Sol via Codex CLI29-task corpus · 5 recorded trialslown=523.6 / 29
22.2–25.0observed 222527.4s13m 14s / corpus0
GPT-5.6 Sol via Codex CLI29-task corpus · 5 recorded trialshighn=523.6 / 29
22.9–24.3observed 232447.4s22m 53s / corpus0
GPT-5.6 Luna via Codex CLI29-task corpus · 5 recorded trialsmediumn=523.2 / 29
22.2–24.2observed 222430.5s14m 45s / corpus0
GPT-5.6 Sol via Codex CLI29-task corpus · 5 recorded trialsmaxn=523.2 / 29
20.1–26.3observed 192581.4s39m 19s / corpus3
GPT-5.6 Luna via Codex CLI29-task corpus · 5 recorded trialslown=523.0 / 29
22.1–23.9observed 222426.6s12m 51s / corpus1
GPT-5.6 Terra via Codex CLI29-task corpus · 5 recorded trialshighn=522.8 / 29
20.8–24.8observed 202438.4s18m 34s / corpus0
GPT-5.6 Terra via Codex CLI29-task corpus · 5 recorded trialsxhighn=522.8 / 29
21.8–23.8observed 222448.8s23m 35s / corpus0
GPT-5.6 Luna via Codex CLI29-task corpus · 5 recorded trialshighn=522.6 / 29
21.2–24.0observed 212444.4s21m 29s / corpus0
GPT-5.6 Luna via Codex CLI29-task corpus · 5 recorded trialsxhighn=522.4 / 29
21.7–23.1observed 222353.9s26m 2s / corpus0
GPT-5.6 Luna via Codex CLI29-task corpus · 5 recorded trialsmaxn=522.4 / 29
20.7–24.1observed 212475.0s36m 16s / corpus4
GPT-5.6 Terra via Codex CLI29-task corpus · 5 recorded trialslown=522.0 / 29
21.1–22.9observed 212325.4s12m 16s / corpus0
GPT-5.6 Terra via Codex CLI29-task corpus · 5 recorded trialsmediumn=521.8 / 29
20.2–23.4observed 202328.4s13m 44s / corpus0
  1. GPT-5.6 Sol via Codex CLI29-task corpusmedium
    Evidence
    n=5
    Mean tasks
    24.0/29
    95% CI
    22.5–25.5
    Seconds / task
    38.5s
    Observed 2225 tasks · 0 timeouts
  2. GPT-5.6 Sol via Codex CLI29-task corpusxhigh
    Evidence
    n=5
    Mean tasks
    24.0/29
    95% CI
    23.1–24.9
    Seconds / task
    60.7s
    Observed 2325 tasks · 0 timeouts
  3. GPT-5.6 Sol via Codex CLI29-task corpuslow
    Evidence
    n=5
    Mean tasks
    23.6/29
    95% CI
    22.2–25.0
    Seconds / task
    27.4s
    Observed 2225 tasks · 0 timeouts
  4. GPT-5.6 Sol via Codex CLI29-task corpushigh
    Evidence
    n=5
    Mean tasks
    23.6/29
    95% CI
    22.9–24.3
    Seconds / task
    47.4s
    Observed 2324 tasks · 0 timeouts
  5. GPT-5.6 Luna via Codex CLI29-task corpusmedium
    Evidence
    n=5
    Mean tasks
    23.2/29
    95% CI
    22.2–24.2
    Seconds / task
    30.5s
    Observed 2224 tasks · 0 timeouts
  6. GPT-5.6 Sol via Codex CLI29-task corpusmax
    Evidence
    n=5
    Mean tasks
    23.2/29
    95% CI
    20.1–26.3
    Seconds / task
    81.4s
    Observed 1925 tasks · 3 timeouts
  7. GPT-5.6 Luna via Codex CLI29-task corpuslow
    Evidence
    n=5
    Mean tasks
    23.0/29
    95% CI
    22.1–23.9
    Seconds / task
    26.6s
    Observed 2224 tasks · 1 timeouts
  8. GPT-5.6 Terra via Codex CLI29-task corpushigh
    Evidence
    n=5
    Mean tasks
    22.8/29
    95% CI
    20.8–24.8
    Seconds / task
    38.4s
    Observed 2024 tasks · 0 timeouts
  9. GPT-5.6 Terra via Codex CLI29-task corpusxhigh
    Evidence
    n=5
    Mean tasks
    22.8/29
    95% CI
    21.8–23.8
    Seconds / task
    48.8s
    Observed 2224 tasks · 0 timeouts
  10. GPT-5.6 Luna via Codex CLI29-task corpushigh
    Evidence
    n=5
    Mean tasks
    22.6/29
    95% CI
    21.2–24.0
    Seconds / task
    44.4s
    Observed 2124 tasks · 0 timeouts
  11. GPT-5.6 Luna via Codex CLI29-task corpusxhigh
    Evidence
    n=5
    Mean tasks
    22.4/29
    95% CI
    21.7–23.1
    Seconds / task
    53.9s
    Observed 2223 tasks · 0 timeouts
  12. GPT-5.6 Luna via Codex CLI29-task corpusmax
    Evidence
    n=5
    Mean tasks
    22.4/29
    95% CI
    20.7–24.1
    Seconds / task
    75.0s
    Observed 2124 tasks · 4 timeouts
  13. GPT-5.6 Terra via Codex CLI29-task corpuslow
    Evidence
    n=5
    Mean tasks
    22.0/29
    95% CI
    21.1–22.9
    Seconds / task
    25.4s
    Observed 2123 tasks · 0 timeouts
  14. GPT-5.6 Terra via Codex CLI29-task corpusmedium
    Evidence
    n=5
    Mean tasks
    21.8/29
    95% CI
    20.2–23.4
    Seconds / task
    28.4s
    Observed 2023 tasks · 0 timeouts

Corpora are intentionally separated and time is normalized per task. The focused y-axis is explicitly labelled; Full scale restores the zero baseline. Lines show configuration order from lower to higher effort; they do not imply continuous scaling or monotonic treatment. See the reproducibility method. Raw run IDs are shown in trial tooltips.

Current trial environment: codex-cli 0.144.1 · host main-pc · corpus 6c4205b2e2cf · network unknown.

Plausible Nix often fails at evaluation time.

The benchmark gives agents a copied starter tree, a prompt, and no access to the hidden evaluator. It rewards final worktree behavior, not a fluent explanation of what the code should do.

Original repair tasks

Tasks are written for this corpus rather than lifted from merged patches, which keeps the answer out of the visible prompt.

Nix-specific failure surfaces

The corpus covers flakes, modules, overlays, derivations, fetchers, Home Manager, shell escaping, and package contracts.

Hand-written checks

Each task has a shell evaluator that checks behavior with small fake package sets and libraries instead of relying on LLM judging.

Diff-backed runs

Every run records logs, timings, pass state, score JSON, and the final diff so failures can be inspected after the benchmark ends.

29 small repositories, one hidden evaluator each.

All 29 tasks

Respect NixOS, Home Manager, and nix-darwin boundaries

Keep module outputs separated instead of leaking options across systems.

Patch Python CUDA package inputs

Repair Python/CUDA packaging without falling back to generic Linux path guesses.

Compose module paths from arguments

Build paths with Nix values while avoiding string interpolation traps.

Debug network symptoms without false leads

Explain the observed NixOS service behavior without chasing a plausible but wrong network diagnosis.

Manage home files declaratively

Use Home Manager file and XDG options rather than imperative setup.

Pin a GitHub source fetcher

Preserve the fixed-output fetcher contract with a commit pin and SRI hash.

7

easy tasks for syntax, lookup, stale options, and small contracts

16

medium repairs across flakes, containers, issue reports, overlays, packaging, and shell integration

6

hard tasks for modules, overlays, portals, Rust purity, and Python/CUDA package inputs

The agent edits a worktree. The evaluator scores the result.

A simple, inspectable protocol separates generation from evaluation.

Read the protocol
  1. copy

    Starter files and the prompt enter a clean temporary workdir.

  2. edit

    The agent reads NIXBENCH_PROMPT.md and modifies only local files.

  3. check

    A hidden shell evaluator scores the final tree after the agent exits.

  4. record

    Logs, timing, score JSON, and the final diff are written under results/.

Add evidence, not another claim.

Use the open harness, preserve the run artifacts, and add a comparable row to the benchmark.

Open the run guide