Skip to content

Documentation

Authoring Tasks

Good NixBench tasks should test Nix skill directly, not incidental knowledge of a particular package.

Maintained in docs/authoring.md

Documentation

Good NixBench tasks should test Nix skill directly, not incidental knowledge of a particular package.

Full-System Tasks (corpus 3.1)

New tasks since 3.1 grade against a real, pinned nixpkgs. The 3.0 calibration showed why: single-file idiom tasks graded with fake package sets saturated, and the only tasks frontier models missed were graded by executing real behaviour. The conventions:

  • **The oracle is the pinned nixpkgs.** The evaluator receives the pinned

store path as NIXBENCH_NIXPKGS (PIN.md) and exits 2 if it is unset or missing. Evaluate a complete NixOS configuration (forcing config.system.build.toplevel.drvPath runs every assertion and type check) or a real package expression called through import $NIXBENCH_NIXPKGS { system = "x86_64-linux"; config = { }; overlays = [ ]; }. Assert on resolved option values, derivation attributes, generated unit and configuration text, and assertions and warnings. Eval-track evaluators build nothing; only VM-track evaluators (below) build and boot. Never substitute a fake package set.

  • **Timeout.** Declare timeout_seconds = 120 for the eval track. Keep the

whole evaluator well under that; the release gate requires runtime below 80 percent of the timeout. Observed 3.1 runtimes are 4 to 29 seconds. VM-track tasks declare timeout_seconds = 600.

  • **A repository, not a snippet.** The starter is a realistic repository of

roughly 10 to 40 files with at least two hosts, profiles, or targets. The constraint the agent must honour lives in the repository (a comment in another module, a policy in docs/, a shared profile, a sibling host that must keep working), not in the prompt. Include a README or justfile that shows how the repository is evaluated, so the agent can reproduce the symptom with <nixpkgs> on NIX_PATH.

  • **Real symptom text.** Run the starter against the pin and paste the actual

error or warning into the prompt (for 3.2 silent-failure tasks, the runtime output instead; see below). Keep prompts under about 300 words, with no recipe and no helper or option names that give the fix away.

  • **Evaluator-owned baseline.** Copy the starter to tests/baseline/ and

compute "must not change" from it (a sibling host's resolved values, unrelated units, upstream sources). Do not read starter/ from the evaluator: contract runs overlay each fixture's candidate onto starter/, so it no longer holds the original. Do not snapshot the reference either.

  • **Metamorphic grading.** Evaluate the host in the prompt and at least one

sibling host or profile that must not regress, plus at least one variation the evaluator controls (another hostname, user set, storage layout, or feature toggle, layered as extra modules or through extendModules), so answers hard-coded to the starter fail.

  • **One process per criterion, stderr to scratch.** Run each criterion in its

own nix-instantiate --eval --strict --json process with --argstr for the pin, the workdir, and the baseline. On Nix 2.34, nix eval --file does not call a top-level function, so --argstr has no effect unless an attribute path follows the file; prefer nix-instantiate, or use nix eval --file probe.nix <criterion> --argstr .... Redirect each process's stderr to a scratch file: a contract fixture whose evaluator log contains error: fails the release gate. Pass --option allow-import-from-derivation false. Guard nested lookups and finish with the python3 "$NIXBENCH_EVALUATOR_EXIT" line.

  • **Two passing alternatives.** Keep at least two passing fixtures that solve

the task differently from the reference (raising priority against restructuring the module, an overlay against an override, splitting a module against inlining it), and a rejecting fixture for every criterion.

  • **Fresh material where it is discoverable.** Prefer nixpkgs changes whose

current shape differs from what a model trained before mid-2026 would write, but only when the agent can find the current shape in the pinned tree (option descriptions, rename modules, the 26.05 release notes). Verify the trap against the pin before writing the task; see fresh-nixpkgs-changes-2026-10.md.

  • **Other pinned trees need GC roots.** A task that needs a second source tree

(for example Home Manager in home-manager-nixos-module-boundary-fleet) pins it in the starter by store path and NAR hash and adds an indirect GC root under vendor/gcroots/; the registration commands are in PIN.md.

Set the pin for local testing with:

export NIXBENCH_NIXPKGS=$(python3 -c 'import json;print(json.load(open("vendor/nixpkgs/pin.json"))["store_path"])')

If a trap can be solved from memory in one obvious edit, make the repository larger, couple it to a second constraint, or drop the task.

Silent Failures And The VM Track (corpus 3.2)

The 3.1 calibration showed that frontier agents solve any task whose defect surfaces as an evaluation error or warning: they loop on nix-instantiate until it is clean. The one task every configuration failed had a requirement evaluation never reports. Every task added in 3.2 follows the silent-failure rule, and some are graded on a second track that boots the system.

The silent-failure rule

  • **The starter evaluates cleanly.** Every host evaluates against the pin

with no error and no warning (or fails only on an unrelated, obvious error that is not the task). Check this before writing anything else.

  • **The defect is invisible to evaluation.** The agent cannot find it by

making errors disappear. It has to read the pinned tree (release notes at $NIXBENCH_NIXPKGS/nixos/doc/manual/release-notes/, option descriptions, module source), inspect resolved configuration it did not think to look at (a sibling host, a generated unit, a tmpfiles rule), or reason about runtime. If the defect turns out to be visible as an evaluation error against the pin, the task is not a 3.2 task: couple it with a second silent constraint or drop it.

  • **The prompt quotes runtime output.** Describe what the operator observed

after deploying (a boot hang, a service that cannot write its state, a listener on the wrong address, data in the wrong directory, a refused login), the command the deploy pipeline runs, and the goal. Paste real output. Never name the option or the mechanism.

  • **Pair it with a regression trap.** Give each task a sibling host or

profile whose resolved configuration must not change, so the natural overbroad fix (a mkForce in a shared profile, a forced list) fails. The silent-regression traps T01 to T05 in silent-failure-changes-2026-10.md are verified examples.

Candidates and their pinned evidence are in silent-failure-changes-2026-10.md.

Two tracks

Each task declares its track as a comment line in metadata.toml, because the loader has no field for it:

# track = "vm"
timeout_seconds = 600

Eval-track tasks write # track = "eval" with timeout_seconds = 120. Tasks from before 3.2 carry no comment and are eval-track.

  • **Eval track.** As in 3.1: assert on resolved configuration, generated unit

text and configuration files, tmpfiles rules, kernel parameters, the initrd unit graph, PAM services, or secret manifests. The assertions must distinguish the old form from the new one, as the research probes do.

  • **VM track.** Use it when only runtime proves the requirement (ownership

inside a sandbox, packet forwarding, what survives a reboot). The conventions, with tasks/state-directory-ownership-after-hardening/tests/ as the template:

  • tests/vm.nix takes nixpkgs and repo and returns

(import nixpkgs { system = "x86_64-linux"; ... }).testers.runNixOSTest with the repository's host modules as nodes, plus evaluator-controlled variants, and testScript = builtins.readFile ./vm-script.py.

  • Keep at most three VM test runs per evaluation, counting the baseline.

Put every independent assertion into one test script.

  • tests/check.sh builds the test against tests/baseline and against the

candidate in parallel (two background nix build --no-link --print-out-paths -f tests/vm.nix --argstr nixpkgs ... --argstr repo ... jobs), each bounded by timeout.

  • The test script wraps each criterion's checks in try/except, never

raises, and writes {"criteria": {...}, "notes": [...]} to $out/verdict.json. check.sh maps that verdict onto the declared criteria, so one failing criterion does not zero the others.

  • Exit 2 on infrastructure failure: no /dev/kvm, no pinned tree, a

missing test file, or a baseline verdict that differs from the starter's known verdict (hard-coded in check.sh). A candidate whose test fails to evaluate or build gets all criteria false and exit 1.

  • The agent cannot boot VMs inside the sandbox. The prompt says so, says

that nix-instantiate and nix eval work offline, and quotes the deploy pipeline's VM test output from the starter.

  • Verify that the starter's test fails for the intended runtime reason and

the reference passes, and time a cold and a cached run: the evaluator must stay below 80 percent of the 600-second timeout.

The guidance below still applies to every task, including the language and module tasks that grade with the vendored lib.

Prefer

  • Self-contained evaluators.
  • The pinned full nixpkgs for anything that touches NixOS modules or real

packages (see above).

  • For pure language tasks, the real nixpkgs lib from $NIXBENCH_VENDOR_LIB,

with real fixed points. If the evaluator sources nixbench-evaluator.sh, call nixbench_init with the revision in PIN.md. Fake package sets (fakePackage in nixbench-eval.nix) remain only in the tasks carried over from 3.0; do not use them in new tasks.

  • Assertions on resolved values, not on how the candidate is written.

mkIf placement, mkMerge, mkDefault, list order, and helper choice must not matter.

  • Running candidate shell snippets (unit commands, wrappers, build and check

phases) against fake tools, in the sandbox helper (nixbench_sandbox.py) or a task-local probe, instead of matching them with regular expressions.

  • One Nix process per criterion (nixbench_score, or a task-local loop that

does the same), so an evaluation error in one criterion cannot zero the others.

  • Hidden inputs and variations that catch hardcoding.
  • Metamorphic probes that the task contract implies: reorder or rename what

the prompt says is irrelevant and expect the same result; add an unrelated module, overlay, or record and expect it to survive; change the one input that matters and expect the result to change. Do not assume a transformation is harmless unless the prompt says so.

  • Clear file targets in the prompt.
  • Symptom-driven prompts: state what the user observes (error text, failing

command, wrong behaviour) and the expected behaviour, plus the hard constraints that are part of the task. Do not name the helper, option path, or attribute shape that fixes it unless the evaluator genuinely cannot accept alternatives. When a detail is required, phrase it as a requirement, not a recipe. Arbitrary interface names (output attribute names, marker values) are fine to state.

  • Source links to the forum thread or issue that inspired a task belong in

research-derived-tasks.md, never in the prompt. Those threads usually contain the accepted fix.

  • Reference solutions that are boring and idiomatic.

Avoid

  • Network fetches in evaluators. VM-track evaluators may substitute VM

closures from the binary cache and nothing else.

  • Real builds in eval-track evaluators. Instantiate derivations and inspect

them instead; reserve building for the VM track, and never require the agent to build.

  • Prompts that leak the exact hidden assertions or enumerate the rubric, which

turns hidden evaluators into instruction-following checks.

  • Checks that only look for strings when Nix evaluation can inspect the value.
  • Grep bans on helpers that exist in nixpkgs, exact list equality where a

superset is valid, and checks that require one syntactic module shape.

Test The Evaluator, Not Only The Reference

A passing reference and failing starter are necessary, but they do not show that an evaluator draws the right boundary. For each task, also keep adversarial contract cases that cover:

  • A plausible-looking candidate that violates one important requirement and must fail.
  • At least two semantically valid alternatives that differ from the reference and must pass. Prefer a materially different idiom (for example mkMerge with per-attribute mkIf against one top-level mkIf) over cosmetic changes.
  • A second input when the candidate is a function, so hard-coded answers do not pass.
  • Opaque sentinel values for fake packages and builders, rather than strings that are identical to their attribute names.

Store evaluator cases under contracts/<task-id>/<case-id>/. Add a regression whenever a real benchmark run exposes a false positive or false negative.

Before activation, the corpus release check also requires deterministic repeated outcomes, no invalid measurements or known-issue skips, evaluator runtime below 80 percent of the task timeout, a release note, and a checked release manifest. Every task starts or returns to calibrating after a digest change. Generate draft evidence with calibration-report; do not hand-copy aggregate claims or edit lifecycle state from the command. Activation requires three materially different validated configurations with at least three valid observations per task in each, plus manual accepted-alternative and evaluator- dispute review and an explicit approval rationale. Tasks the report flags as saturated (pooled solve rate of at least 0.9) need an explicit reason to stay active. Further repetitions improve stability but do not replace those mandatory identities and reviews. Follow the full workflow in benchmark-governance.md.

Each case uses contract schema 3, names the rubric criterion it exercises, and declares the complete expected criterion vector:

schema_version = 3
task_id = "package-stdenv-cli"
outcome = "reject"
criterion_id = "install-contract"
description = "The install command is present only in a comment."

[expected_criteria]
package-source = true
build-contract = true
install-contract = false
package-metadata = true

Every required criterion needs its own targeted rejecting case; a criterion label on a passing fixture is not negative coverage. A passing case expects all criteria true. A rejecting case expects its named criterion false and should keep unrelated criteria true. If one plausible mutation necessarily causes additional failures, list every such criterion and a reason in [coupled_failures]. The loader rejects missing or unknown vector keys, undocumented coupling, and a rejecting case whose named criterion is expected true.

Candidate digests are part of the boundary inventory. Do not present the same rejecting candidate as independent evidence for multiple criteria unless the cases intentionally vary evaluator inputs and each manifest records a duplicate_candidate_reason. release-check requires at least one passing alternative whose materialized candidate differs from the reference. Corpus 3.0, 3.1, and 3.2 tasks keep at least two.

Evaluators initialize every outcome to false and write the schema-2 score file atomically. Make each criterion total: guard nested attribute access and keep candidate-dependent evaluation inside that criterion's boundary so one missing field does not erase unrelated credit. An ordinary whole-candidate syntax or import failure may exit 1 with an all-false valid payload, but an isolated missing field must not collapse the complete vector. Reserve exit codes 2 and greater for evaluator implementation or infrastructure failures.

Hidden cases may vary inputs and expose edge conditions. Hidden evaluators may not require an undocumented representation when common semantic alternatives exist.

Retiring Tasks

Do not delete a saturated or superseded task. Record it in corpus/task-deprecations.toml with deprecated_in, a reason such as saturated, and a note, then move tasks/<task-id>/ and contracts/<task-id>/ together into archive/<version>/. Keep its evaluator passing its contract fixtures there; the contract tests still load the archive. The move changes the corpus digest, so it ships with a release note and a new release manifest.

Task Ideas

  • Fix a NixOS module option type and conditional config.
  • Convert a single-system flake into a per-system flake.
  • Add overrideAttrs without dropping existing patches or metadata.
  • Replace impure host paths with package inputs.
  • Repair infinite recursion from self-referential attrsets.
  • Compose module paths passed through arguments without confusing path addition, string interpolation, and list syntax.
  • Package a Python, Rust, or Go app against the real package set, checking derivation attributes and inputs rather than building.