Skip to content

Documentation

Benchmark Design

NixBench is designed around agentic repair, not snippet generation.

Maintained in docs/benchmark-design.md

Documentation

NixBench is designed around agentic repair, not snippet generation.

A model is given an editable worktree and a task prompt. It must inspect the files, make changes, and leave the worktree in a state that passes hidden checks. This gives the benchmark room to measure practical behavior: reading requirements, editing the right file, avoiding irrelevant churn, running local checks, and handling Nix evaluation errors.

Goals

NixBench aims to measure:

  • Whether an agent can produce Nix code that evaluates.
  • Whether it can distinguish common Nix contexts: flakes, modules, overlays, derivations, shells, and fetchers.
  • Whether it can diagnose from the observed symptom instead of the most familiar fix.
  • Whether its fix survives the module system, overlay fixed points, and laziness as nixpkgs implements them.
  • Whether it can preserve existing metadata, inputs, patches, and options.
  • Whether it can fix broken code rather than replacing it blindly.

Non-Goals

NixBench does not currently try to measure:

  • Large-scale nixpkgs contribution quality.
  • Long-running real package builds.
  • Human taste in Nix style.
  • End-to-end NixOS deployment behavior beyond what a short NixOS VM test

shows (the 3.2 VM track boots hosts, but does not deploy to real hardware).

  • Cross-model leaderboard fairness across hardware and network conditions.

Those can be added later, but the initial corpus focuses on fast deterministic evaluation.

Real Evaluation Is The Oracle (corpus 3.1)

Corpus 2.0 graded with a fake lib and source-text matching. Corpus 3.0 moved to a pinned, vendored nixpkgs lib and real lib.evalModules, but kept small fake builders and package sets, on the theory that a fake stdenv.mkDerivation was enough to inspect a package's structure. The 3.0 calibration showed what that cost: 24 of 30 tasks were saturated. A task that fits a fake package set is a single-file idiom task, and frontier models solve those from memory. The only tasks they missed were the ones whose evaluators executed real behaviour. Fake builders were the reason the corpus saturated, and one of them (an incomplete fake coreutils set) also produced the 3.0 evaluator dispute.

Corpus 3.1 therefore grades against the real thing:

  • **Pinned nixpkgs, evaluated for real.** A complete nixpkgs tree is pinned

by store path and NAR hash in vendor/nixpkgs/pin.json (PIN.md). The runner exports it to evaluators as NIXBENCH_NIXPKGS after checking the hash; the fresh-home adapters bind it to <nixpkgs> for the agent. Evaluators evaluate complete NixOS configurations (config.system.build.toplevel) or real package expressions through the real package set, and assert on resolved option values, generated unit and configuration text, derivation attributes and inputs, and assertions and warnings. A configuration that the pinned NixOS rejects is rejected by the evaluator in the same way.

  • **Repositories, not snippets.** Each new task is a multi-file repository

(14 to 40 files) with at least two hosts, profiles, sites, or targets. The constraint that makes the obvious fix wrong is written somewhere in the repository, so the agent has to read it.

  • **Evaluator-owned baselines.** "Unchanged" is computed from

tests/baseline/, the evaluator's copy of the starter, not from a snapshot of the reference. Contract runs overlay candidates onto starter/, so the evaluator cannot use that directory as the original.

  • **Metamorphic variations.** Evaluators vary what the task contract says

may vary: another hostname, user set, storage layout, persistent root, or transport, often by layering modules on the candidate or calling extendModules. Answers hard-coded to the starter fail.

  • **Fresh material.** Nine tasks rest on nixpkgs changes a model trained

before mid-2026 is unlikely to know, and that the agent can discover in the pinned tree (research-derived-tasks.md).

The 3.0 conventions still hold for every task:

  • **One process per criterion.** Each criterion is evaluated in its own Nix

process, with candidate stderr sent to a scratch file. A candidate that aborts one criterion (an uncatchable type error, a missing attribute, a failed assertion) fails that criterion only.

  • **Execute, don't match.** Shell that the candidate produces is run and its

effects are checked. Source-text checks remain only for explicit prompt prohibitions.

  • **Symptom-driven prompts.** Prompts state what the user observes and the

goal, not the helper or option path that fixes it (prompt-rewrite-3.0.md).

  • **At least two passing alternatives.** Every task keeps at least two

passing contract fixtures that differ materially from the reference, beside a targeted rejecting fixture for every required criterion. release-check enforces one alternative; two is the authoring rule.

The tasks carried over from 3.0 (six in 3.1, five in 3.2) still grade with the vendored lib (PIN.md), and some of them with small fake packages. New tasks do not use fake package sets.

Why agent-side builds stay out of scope

Agents never build. Inside the sandbox they have no Nix daemon socket, only the pinned tree on NIX_PATH, so they can evaluate and instantiate but not realise derivations or boot VMs. Asking the agent to build would add minutes per task, depend on binary caches or the network, and fail for reasons unrelated to the candidate (a flaky upstream test, a missing substitute, a full disk). Evaluation already lets the agent check what most tasks grade: whether the configuration is accepted, which units, files, and inputs it produces, and which derivation would be built.

Eval-track evaluators also never realise derivations. What evaluation cannot show (that a package build succeeds, that a host boots, that a service starts) is outside the claims of those tasks, and several prompts say so explicitly. Since 3.2, three VM-track evaluators do build and boot the candidate's hosts on the evaluator side (see below); the agent still does not. Real package-build tasks may come back later behind a separate slow profile.

Evaluation-Only Grading Saturated (3.2)

Grading against the real pinned nixpkgs widened the spread between configurations in the 3.1 calibration, but nine of twenty-one tasks were still saturated and the top four configurations were tied. The cause was the grading signal itself. When every requirement shows up as an evaluation error, an assertion, or a warning, a frontier agent can treat nix-instantiate as the oracle: run it, fix what it reports, repeat until clean. The evaluator then checks the same thing the agent already checked. The only task every configuration failed, initrd-hook-systemd-stage1-migration, had one requirement that evaluation never reports. Every agent fixed the evaluation error, but none changed the root filesystem device, which evaluates fine and only fails at boot; the requirement is in the pinned release notes, which no agent read.

Corpus 3.2 builds every new task from that pattern:

  • **Silent-failure tasks.** The starter evaluates with no error and no

warning, and the defect only appears at runtime. The requirement is documented in the pinned tree or the repository, and the prompt quotes the runtime symptom (a journal excerpt, a refused login, a failed VM test) rather than an evaluation error. A clean evaluation is no longer evidence of a correct answer, so the agent has to read release notes, option descriptions, module source, and the resolved configuration of hosts it was not asked about. Each task is paired with a sibling host or profile that the natural overbroad fix would change. The candidates were admitted only if their old form evaluated with exit 0, empty stderr, and no warnings (silent-failure-changes-2026-10.md).

  • **The VM track.** Some runtime requirements have no faithful static

projection: ownership seen from inside a sandbox, whether packets are forwarded, what survives a reboot. Three tasks are graded by building and booting NixOS VM tests (testers.runNixOSTest from the pinned tree) with the candidate's hosts as nodes. The concerns that kept builds out of scope are handled inside the evaluator: it boots its own copy of the starter in parallel with the candidate and exits 2 (an invalid measurement, not a candidate failure) if that baseline does not fail in exactly the known way or if KVM is missing. A test script records one boolean per criterion in a verdict file, so one runtime failure does not zero the others. These tasks declare a 600-second timeout and need KVM and binary-cache access on the evaluator host. The conventions are in task-format.md.

Both tracks keep the 3.1 rules: real pinned nixpkgs, multi-file repositories, evaluator-owned baselines, metamorphic variations, and one independent result per criterion.

Why Hidden Evaluators Matter

If the public prompt says "include mainProgram", a weak model can satisfy the visible text with a string search strategy. A hidden evaluator can check semantic shape instead:

assert pkg.meta.mainProgram == "tinygrep";

Hidden evaluators also catch overfitting. For example, a prompt may include one package set, while the evaluator uses another package set with disabled packages, missing fields, or different system lists.

Task Difficulty

Suggested difficulty levels:

  • Easy: direct requirements, no subtle Nix semantics.
  • Medium: multiple constraints, some idiom or repository reading required.
  • Hard: the module system, overlays, recursive values, or purity constraints,

usually with a constraint that has to be discovered in the repository.

Difficulty should reflect the expected reasoning load, not the number of lines changed.

Corpus Health Checks

Every task should satisfy two checks:

python3 bench.py validate --solution reference
python3 bench.py validate --solution starter

The reference should pass. The starter should usually fail. If a starter passes, the task is not measuring anything useful. The validate command exits successfully only when every reference passes at full score or every starter is cleanly rejected with evaluator exit code 1 below full score, depending on the selected mode. Evaluator timeouts, invalid score files, and infrastructure exit codes are never counted as healthy starter failures. The command rejects an empty task selection rather than reporting a vacuous success.

Those two checks are only corpus smoke tests. Evaluator contract tests should additionally prove that known-invalid mutations fail and valid alternative implementations pass. This prevents hidden checks from becoming either a reference-solution snapshot or a loose shape check that can be gamed with hard-coded values.

The corpus-health command records these checks in versioned JSON keyed by the corpus digest:

python3 bench.py corpus-health \
  --studies-dir results/studies \
  --output results/corpus-health.json

It runs each reference and contract fixture twice to record evaluator determinism and runtime. The artifact also includes starter rejection, pass/reject fixture counts, criterion coverage, observed pass and timeout rates, invalid-measurement rates, and per-task, per-configuration Wilson intervals. A pooled empirical pass rate may be shown as a descriptive rate, but it has no Wilson interval and is not labeled configuration stability. Historical aggregate-only studies do not contribute invented task observations.

The discrimination statistic is the point-biserial correlation between a task's binary outcome and the sum of the same trial's normalized task scores with that task removed. Thresholds are applied after trials without a usable leave-one-task-out score are removed. The default threshold requires at least 20 valid observations across at least two correctness configurations. Reports return null with a reason below either threshold or when either input has zero variance. Author-assigned difficulty and the empirical solve-rate band remain separate fields; health reporting never changes the author label. The bands are low below 25%, mixed from 25% through less than 75%, and high at 75% or above. Calibration and study-matrix translate the pooled solve rate into an empirical difficulty of hard, mixed, or easy. A pooled solve rate of 90% or more is labelled saturated instead, which marks a retirement candidate.

Category and difficulty results describe the fixed tasks in that group. Groups with fewer than five tasks are descriptive only. NixBench does not treat the corpus as a random sample of all Nix work.

Publishable whole-corpus, category, difficulty, and task strata include the timeout count and timeout rate alongside their valid-observation denominator.

Retiring Saturated Tasks

A task that nearly every configuration solves no longer separates models. When calibration-report or study-matrix labels a task saturated (pooled solve rate of 0.9 or more), reviewers either harden it, keep it active with a recorded reason, or retire it. Retired public tasks are recorded in corpus/task-deprecations.toml and move out of tasks/ into a versioned archive with their contract fixtures. Corpus 3.0 moved 23 saturated 2.0 tasks to archive/2.0/, corpus 3.1 moved 24 saturated 3.0 tasks to archive/3.0/, and corpus 3.2 moved 9 saturated 3.1 tasks to archive/3.1/. The archive is outside the corpus digest and never enters active scoring, but its evaluators still run in the contract tests as regression guards. Moving tasks changes the corpus digest, so it ships in a new release (benchmark-governance.md).

Benchmark Integrity

For fair runs:

  • Do not expose tests/check.sh to the agent workdir.
  • Do not expose the original task directory, reference solution, or score file path to the agent.
  • Use the same timeout for every comparable agent.
  • Record the exact agent command.
  • Keep summary.json, result.json, agent.log, check.log, and diff.patch.
  • Pin task corpus versions when comparing results over time. The pinned

nixpkgs, the vendored lib, and the evaluator support files are outside tasks/ and not hashed into the corpus digest, so record the Git revision too; it fixes vendor/nixpkgs/pin.json.

  • Run full-system tasks only against the pinned tree. The runner withholds

NIXBENCH_NIXPKGS when the store path is missing or its NAR hash differs, and the evaluator then exits 2, which marks the measurement invalid (evaluator-error) instead of failing the candidate.