Each benchmark task lives in tasks/<task-id>/.
tasks/<task-id>/
metadata.toml
prompt.md
starter/
reference/
tests/
check.sh
baseline/ # full-system tasks: evaluator-owned copy of starter/
vm.nix # VM-track tasks: evaluator-owned NixOS VM test
vm-script.py # VM-track tasks: the test's Python test script
metadata.toml
Required fields:
id = "package-stdenv-cli"
name = "Package A stdenv CLI"
category = "packages"
difficulty = "medium"
timeout_seconds = 60
max_score = 100
systems = ["any"]
evaluator = "tests/check.sh"
[[criteria]]
id = "package-source"
points = 25
required = true
failure_class = "impurity"
systems = ["any"] means the task is system-independent. Use Nix system names such as x86_64-linux or aarch64-darwin for system-specific tasks.
Full-system tasks, which evaluate complete NixOS configurations or real package expressions against the pinned nixpkgs, declare timeout_seconds = 120. Observed 3.1 evaluator runtimes are 4 to 29 seconds; the release gate requires the runtime to stay below 80 percent of the timeout. Tasks that only need lib keep timeout_seconds = 60. VM-track tasks (see VM track) declare timeout_seconds = 600.
The loader has no field for the grading track. Since 3.2, a task records it as a comment line in metadata.toml, # track = "eval" or # track = "vm". Tasks without the comment are graded by evaluation. The comment is part of the file, so it is hashed into the corpus digest like any other byte.
All listed fields are mandatory. Task ids must be lowercase hyphenated slugs, difficulty must be easy, medium, or hard, and category must use the controlled vocabulary in Metadata guidance below. Corpus loading rejects an unknown category. Timeouts and maximum scores must be positive, and systems must be a non-empty list without duplicates. The evaluator must be a relative path that resolves inside the task directory. The prompt and evaluator must be files; starter/, reference/, and tests/ must be directories, and required paths may not escape the task through symlinks. Duplicate task ids are rejected when loading a corpus.
corpus/category-vocabulary.toml is the release-controlled category list. Changing it requires release review and a new checked gate digest.
Corpus 3.2 tasks declare 4 to 6 objective criteria. Criterion IDs are unique lowercase slugs, points are positive and sum exactly to max_score, and every public prompt requirement maps to a criterion. Most criteria are required. A task may mark at most one objective, secondary maintainability or formatting criterion required = false; no 3.2 task does. See scoring.md for failure classes and compatibility rules.
The corpus manifest lives beside tasks/ in corpus.toml. python3 bench.py corpus-id --json computes a SHA-256 digest over the manifest identity fields and all benchmark-owned task content. Paths, file kinds, executable bits, symlink descriptors, and byte lengths are part of the canonical stream. Corpus identity loading rejects every symlink inside editable tasks/<task-id>/starter/ and contracts/<task-id>/<case-id>/candidate/ trees, including links whose targets remain in the same editable tree. Symlinks in trusted reference and evaluator content remain supported when their targets stay inside the corpus root.
Changing task metadata, a prompt, starter, reference, evaluator, rubric, or contract fixture changes the corpus digest. A semantic change may also require a corpus version bump. Git revision is recorded separately as provenance and does not define corpus identity.
prompt.md
This is the prompt copied into the workdir as NIXBENCH_PROMPT.md. It is symptom-driven: it describes what the user observes (the error text, the failing command, or the wrong behaviour), the expected behaviour, the files to edit, the interface the evaluator calls, and the hard constraints. It does not reveal hidden evaluator details or name the helper, option path, or attribute shape that fixes the problem. Arbitrary interface names, such as output attributes and marker values, stay explicit. Prompts contain no links to the source threads that motivated a task; those live in research-derived-tasks.md.
Starter reproductions that need nixpkgs lib (demo.nix, repro.nix, or a local lib.nix) import it from $NIXBENCH_VENDOR_LIB when set and from <nixpkgs/lib> otherwise.
Full-system starters evaluate against <nixpkgs>. The fresh-home agent adapters set NIX_PATH=nixpkgs=<store path> to the pinned tree in vendor/nixpkgs/pin.json, so <nixpkgs>, <nixpkgs/lib>, and <nixpkgs/nixos> resolve to exactly what the evaluator uses. The starter includes a README (or justfile) showing how the repository is normally evaluated, for example nix-instantiate default.nix -A <host>.config.system.build.toplevel, so the agent can reproduce the symptom offline. The prompt quotes the real error text, or the observed misbehaviour, that the starter produces against the pin.
In full-system tasks the constraint the agent must honour lives in the repository (a comment in another module, a policy in docs/, a shared profile, a second host that must keep working), not in the prompt. The prompt points at the repository's rules without restating them.
Silent-failure tasks (3.2)
Every task added in 3.2 is a silent failure:
- **The starter evaluates cleanly.** Every host evaluates against the pin
with no error and no warning, or fails only on an unrelated, obvious error that is not the task.
- **The defect is invisible to evaluation.** Making evaluation errors go away
cannot find it. The agent finds it by reading the pinned tree (release notes under nixos/doc/manual/release-notes/, option descriptions, module source), by inspecting resolved configuration it had no reason to look at (a sibling host, a generated unit, a tmpfiles rule), or by reasoning about runtime.
- **The prompt quotes runtime output.** Since there is no evaluation error to
quote, the prompt pastes what the operator saw after deploying (a systemctl status, a journal excerpt, a failed login, the deploy pipeline's VM test log), the command that produced it, and the goal. It never names the option or the mechanism.
If the defect turns out to be visible as an evaluation error against the pin, the task does not qualify: couple it with a second silent constraint or drop it.
starter/
Starter files are copied into a temporary workdir. The agent edits only this copy.
Full-system starters are realistic repositories of roughly 10 to 40 files: an entry point (default.nix, often with a flake.nix beside it), hosts/, modules/ or profiles/, overlays or pkgs/, documentation, and stubs for secrets and hardware-configuration.nix.
tests/baseline/
Full-system evaluators need the unmodified starter to compute what must not change (a sibling host's resolved configuration, unrelated units, upstream sources the agent may not edit). They cannot read it from starter/: contract runs overlay each fixture's candidate files onto a copy of starter/ and validate that tree, so for a contract case starter/ and the candidate are the same files. tests/baseline/ is the evaluator's own copy of starter/, or of the starter files the evaluator reads (some tasks leave out documentation). Its files must stay identical to the starter's, so update both when editing the starter. An evaluator should exit 2 when the baseline is missing.
reference/
Reference files are overlaid onto the starter when running:
python3 bench.py validate --solution reference
The reference solution should pass the evaluator and act as a regression fixture for task authors.
tests/check.sh
The evaluator runs as:
/bin/sh tests/check.sh "$NIXBENCH_WORKDIR"
It exits 0 only when every required criterion passes, 1 for a candidate rejection, and 2 or greater for evaluator infrastructure errors. An ordinary candidate parse or evaluation error must still write a valid all-false or partial criterion payload before exiting 1.
{
"schema_version": 2,
"criteria": {
"package-source": true,
"build-contract": false
},
"notes": ["build contract failed"]
}
The runner gives the evaluator NIXBENCH_TASK_DIR, NIXBENCH_SCORE_FILE, NIXBENCH_EVALUATOR_EXIT (the exit-code helper), NIXBENCH_VENDOR_LIB (the vendored nixpkgs lib; see PIN.md), and NIXBENCH_NIXPKGS (the pinned full nixpkgs store path; see PIN.md). None of these reach the agent. The runner sets NIXBENCH_NIXPKGS only when the store path exists and its NAR hash matches the pin. A full-system evaluator exits 2 when the variable is unset or does not point at a nixpkgs tree.
Evaluate each criterion in its own process. Evaluators that source vendor/nixpkgs-lib/nixbench-evaluator.sh call nixbench_init <revision>, which exits 2 unless the vendored lib is at that revision, and then nixbench_score, which evaluates each criterion separately, writes the payload atomically, and exits through the helper. Other evaluators run a task-local probe per criterion and do the same. A criterion whose evaluation aborts is false; it does not erase the others.
Full-system evaluators follow the same rule with one Nix process per criterion, each with its stderr redirected to a scratch file. The release gate fails when a contract fixture's evaluator log contains error:, so expected candidate errors from rejecting fixtures must not reach it. Pass paths with --argstr:
nix-instantiate --eval --strict --json \
--option allow-import-from-derivation false \
--argstr nixpkgs "$NIXBENCH_NIXPKGS" \
--argstr workdir "$workdir" \
--argstr baseline "$NIXBENCH_TASK_DIR/tests/baseline" \
"$NIXBENCH_TASK_DIR/tests/$criterion.nix" \
>"$scratch/$criterion.out" 2>"$scratch/$criterion.err"
nix eval --file probe.nix --argstr ... does not call a top-level function on Nix 2.34; the arguments are applied only when an attribute path follows the file. Use nix-instantiate --eval --strict --json, or give nix eval an attribute (nix eval --json --file probe.nix <criterion> --argstr ...).
The payload contains exactly the declared criterion IDs with boolean values. The harness derives points and failure classes. A missing, malformed, or scalar payload makes a current task measurement invalid.
NIXBENCH_SCORE_FILE is only provided to the evaluator, not to the agent command. Bundled evaluators use the harness exit helper so exit 0 follows required criteria only. An optional criterion may fail without rejecting the candidate.
VM track
VM-track tasks (# track = "vm", new in 3.2) are graded by booting the candidate's hosts in a NixOS VM test instead of only evaluating them. The runner treats them like any other task; the differences are all inside the task directory. See tasks/state-directory-ownership-after-hardening/tests/ for a complete example.
- **Timeout.** Declare
timeout_seconds = 600. The runtime gate still
requires the evaluator to finish below 80 percent of it. Bound each nix build with timeout and give the test a globalTimeout so a hung guest cannot consume the whole budget.
- **
tests/vm.nix** is a function ofnixpkgsandrepo(both passed with
--argstr). It imports the pinned tree with system = "x86_64-linux"; config = { }; overlays = [ ]; and returns a testers.runNixOSTest whose nodes import the repository's host modules, plus evaluator-controlled variants (a relocated path, another device or port), and whose testScript is builtins.readFile ./vm-script.py.
- **At most three VM test runs per evaluation, counting the baseline.** Put
independent assertions for all nodes into one test script rather than one test per criterion. The shipped evaluators run two: baseline and candidate.
- **Baseline and candidate in parallel.**
check.shbuilds the test once
against tests/baseline and once against the candidate workdir, as two background nix build --no-link --print-out-paths -f tests/vm.nix jobs, with --option allow-import-from-derivation false, and waits for both.
- **
verdict.jsonmaps runtime results to criteria.** The test script runs
each criterion's checks inside its own try/except, records one boolean per criterion, and writes {"criteria": {...}, "notes": [...]} to $out/verdict.json. The script itself never raises, so the build succeeds and one failing criterion does not zero the others. check.sh copies the declared criteria from the candidate's verdict into the schema-2 payload; a candidate whose test does not evaluate or build gets every criterion false and exits 1.
- **Exit
2on infrastructure failure.** The evaluator exits2whennix
or python3 is missing, NIXBENCH_NIXPKGS is unset, tests/vm.nix, tests/vm-script.py, or tests/baseline/ is missing, or /dev/kvm does not exist. It also exits 2 when the baseline verdict differs from the starter's known verdict, which is hard-coded in check.sh: if the unmodified starter does not fail in exactly the expected way, the VM environment (no KVM, unreachable binary cache, an overloaded host) is broken, not the candidate.
- **Host requirements.** VM-track evaluators run on the host, with the Nix
daemon and binary-cache access; the first run fetches the VM closures. The agent has neither, so the prompt says it cannot boot VMs here and quotes the deploy pipeline's VM test output from the starter.
Authoring checklist
Before adding a task to the corpus:
- Run the reference solution and confirm it passes.
- Run the starter solution and confirm it fails.
- Keep evaluator assertions semantic: evaluate full systems and packages
against the pinned nixpkgs ($NIXBENCH_NIXPKGS), language tasks with the real lib, and execute generated shell instead of matching it.
- For full-system tasks, grade at least two hosts or profiles plus an
evaluator-controlled variation, compare against tests/baseline/ rather than the reference, and set timeout_seconds = 120 (eval track) or 600 (VM track, with a # track = "vm" comment).
- For new tasks, confirm the starter evaluates without errors or warnings and
that the prompt quotes runtime output, not an evaluation error (see Silent-failure tasks).
- Add a targeted rejecting fixture for every required criterion and at least
two passing alternatives that differ from the reference.
- Give every contract case a declared
criterion_id, and cover every required criterion. - Probe functions with more than one input and use opaque fake package values to catch hard-coded outputs.
- Add metamorphic probes where the contract implies them (reordered mirrors,
renamed directories, extra modules or overlays around the candidate).
- Avoid network access in the evaluator. The only exception is the VM track,
whose nix build may substitute VM closures from the binary cache; the test itself and the evaluator fetch nothing else.
- Avoid checking only for strings when Nix evaluation can inspect the value.
- Keep prompts symptom-driven and clear about required files and the interface.
- Record calibration evidence and lifecycle state before activation.
- Update the release note and
corpus/releases/<version>.json.
Metadata guidance
Use categories that describe the Nix skill being tested:
nix-languageflakespackagesmodulesoverlaysfetchersdevshellsdebuggingpurity
Use systems = ["any"] for pure evaluation tasks. Reserve real Nix systems for tasks that depend on platform-specific behavior.
Archived tasks
Deprecated tasks move out of tasks/ into a versioned archive with the same layout: archive/<version>/tasks/<task-id>/ and archive/<version>/contracts/<task-id>/. Corpus 3.0 archived 23 saturated tasks under archive/2.0/, corpus 3.1 archived 24 saturated 3.0 tasks under archive/3.0/, and corpus 3.2 archived 9 saturated 3.1 tasks under archive/3.1/. Each archived task has a record in corpus/task-deprecations.toml. The corpus digest covers only tasks/ and contracts/, so the archive does not affect corpus-id. Load it explicitly with python3 bench.py --tasks-dir archive/3.1/tasks list.