Skip to content

Documentation

Task selection for NixBench 3.0

Decision date: 2026-10-03. This selects the new tasks to author from

Maintained in docs/research/task-selection-3.0.md

Documentation

Decision date: 2026-10-03. This selects the new tasks to author from task-candidates-designed.md and task-candidates-community.md, following the recommendations in benchmark-design-research.md.

Selection rules

  • Take the highest-ranked candidates whose trap was verified locally or is

documented in a primary source.

  • Skip candidates whose trap overlaps an already selected one, so one bad

evaluator cannot sink several tasks at once.

  • Bring every category to at least five active tasks where candidates allow.
  • Prompts are symptom-driven: observed behaviour, failing command, the

interface the evaluator will call. No source links, no recipe.

  • Evaluators use the vendored nixpkgs lib ($NIXBENCH_VENDOR_LIB) and

real lib.evalModules or real fixed points wherever the trap involves the module system or overlays. Every criterion has a rejecting fixture and at least one passing alternative that differs from the reference.

Selected: designed candidates (16)

IdCategoryDifficulty
debug-freeform-config-cycledebugginghard
module-deferred-schema-defaultmoduleshard
overlay-finalattrs-reoverrideoverlayshard
scope-override-transitive-dependenciesoverlayshard
lazy-selected-report-validationnix-languagehard
debug-import-argument-cycledebugginghard
module-priority-submodule-applymoduleshard
flake-nested-follows-identityflakeshard
string-command-dependency-contextnix-languagehard
overlay-composed-final-prevoverlayshard
source-filter-traversal-stabilityflakeshard
module-migration-assertion-gatemoduleshard
lazy-recursive-update-closurenix-languagehard
fetcher-fixed-output-identityfetchershard
lazy-type-error-boundarynix-languagehard
purity-explicit-release-inputspuritymedium

Skipped: debug-contextful-attribute-name (overlaps the string-context task), string-shell-template-roundtrip and string-literal-replacement-order (reserve), flake-host-system-output-boundary (reserve), purity-sandbox-phase-confusion (reserve).

Selected: community candidates (8)

IdCategoryDifficulty
argument-default-forwardingnix-languageeasy
mkforce-across-option-reexportmodulesmedium
module-generated-jobs-fixed-pointmodulesmedium
python-build-backend-false-leaddebuggingmedium
follows-preserve-python-compatibilityflakesmedium
runtime-tools-outside-devshelldevshellsmedium
package-offline-test-selectionpackagesmedium
editor-project-state-isolationdevshellsmedium

Skipped because they overlap a designed task: context-preserving-trim, fetcher-stale-fixed-output-cache, source-filter-stable-store-name, module-import-source-before-pkgs. Skipped because the evaluator cannot be made deterministic offline within 60 seconds: nested-npm-dependency-rebuild, nuget-sdk-dependency-reconciliation, fhs-build-output-staging, gc-generations-shared-inodes, home-manager-fish-state-exclusion, samba-runtime-secret-provisioning. Remaining candidates stay in the pool for a later minor release.

Existing tasks

Calibration on the 2.0 digest showed 27 of 29 tasks solved by every configuration. Those tasks move to the deprecated lifecycle state with the reason saturated and stay in the repository for regression use. They do not enter 3.0 active scoring. rust-no-network-build (0% solve rate) and devshell-tooling-contract (mixed) stay active. Tasks whose evaluators were rewritten for semantic grading in 3.0 (module-service-options, string-escaping-systemd, module-system-boundaries, xdg-portal-merge) re-enter as calibrating because their digests changed and the old solve rate no longer describes them.

Resulting active corpus (target)

CategoryActive tasks
debugging3
devshells3
fetchers1
flakes3
modules8
nix-language6
overlays3
packages2
purity1

Thirty tasks. debugging, devshells, fetchers, flakes, overlays, packages, and purity remain below the five-task threshold and are reported descriptively only.