Skip to content

Documentation

Community candidates for NixBench 3.0

Research date: 2026-10-03. These are task proposals, not calibrated claims about model performance. The list contains 26 candidates, including 24 grounded in 2025 and 2026 discussions and two documentation-derived reserves. Entries appear in expected discrimination order. Build the first eight as pilots before expanding the corpus.

Maintained in docs/research/task-candidates-community.md

Documentation

Research date: 2026-10-03. These are task proposals, not calibrated claims about model performance. The list contains 26 candidates, including 24 grounded in 2025 and 2026 discussions and two documentation-derived reserves. Entries appear in expected discrimination order. Build the first eight as pilots before expanding the corpus.

I read README.md, docs/task-format.md, docs/research-derived-tasks.md, and every file in debug-network-false-lead and rust-no-network-build, including both starters, references, metadata, and evaluators. The proposed tasks preserve the small editable project format but replace answer-revealing checklists with symptoms and local evidence. A hidden constraint must be discoverable in the supplied project, logs, or public behavior contract. It must not be a secret requirement invented by the evaluator.

Evidence and ranking

Evidence labels distinguish what the sources establish:

  • **AI failure**: the reporter identifies AI advice or unsuccessful AI assistance and describes the failure.
  • **AI-assisted patch**: the PR discloses AI assistance and a reviewer corrects specific code. This does not prove the model authored that particular line.
  • **Community failure**: a concrete failure or misleading recommendation, without established AI involvement.
  • **Documentation**: an authoritative pitfall, without a recent incident or demonstrated model failure.

Discrimination rank is a research judgment, not a measured pass rate. Feasibility rank is an independent ordering, with 1 easiest to implement faithfully. Author difficulty estimates the work needed to create a sound fixture and evaluator, using easy, medium, and hard. Categories follow the vocabulary requested for this research; home-manager and darwin would need reconciliation with the current release vocabulary before task activation.

Discrimination rankWorking idCategoryAuthor difficultyFeasibility rankEvidence
1flake-input-graph-identityflakeshard12Community failure
2context-preserving-trimnix-languagemedium1Community failure
3nested-npm-dependency-rebuildpackageshard16Community failure
4override-nested-webapp-consumeroverlayshard14Community failure
5module-clone-definitions-not-valuesmoduleshard21Community failure
6scoped-qt-overlay-propagationoverlayshard18AI failure
7fhs-build-output-stagingpackageshard22Community failure
8nuget-sdk-dependency-reconciliationpackageshard17AI failure
9encoded-store-path-retentionpurityhard25Community failure
10module-generated-jobs-fixed-pointmodulesmedium5Community failure
11mkforce-across-option-reexportmodulesmedium4Community failure
12generated-palette-without-ifdpurityhard20Community failure
13darwin-import-selection-phasedarwinmedium8Community failure
14editor-project-state-isolationdevshellsmedium15AI-assisted investigation
15runtime-tools-outside-devshellpackagesmedium13AI-assisted patch
16home-manager-fish-state-exclusionhome-managerhard23Community failure
17follows-preserve-python-compatibilityflakesmedium10Community failure
18fetcher-stale-fixed-output-cachefetchersmedium7Community failure
19samba-runtime-secret-provisioningpurityhard24AI advice reported
20mkshell-cross-toolchain-hooksdevshellshard19AI failure
21gc-generations-shared-inodesdebugginghard26AI failure
22module-import-source-before-pkgsmodulesmedium6Community failure
23python-build-backend-false-leaddebuggingmedium9AI failure
24package-offline-test-selectionpackagesmedium11AI-assisted patch
25source-filter-stable-store-namepuritymedium3Documentation
26argument-default-forwardingnix-languageeasy2Documentation

Evaluator rules shared by all candidates

Use nix eval --offline --json on explicit local files, with an empty NIX_PATH, substitutes disabled, and IFD disabled except where an intentional negative probe checks that it fails. Do not resolve remote flakes, import installed nixpkgs, build packages, boot VMs, or start services. For file-based evaluation, offline operation and pure evaluation are different properties: do not blindly turn on pure-eval for an arbitrary temporary absolute --file path. Use a local flake or explicitly admitted source root when testing pure evaluation itself.

Vendor a pinned real nixpkgs lib closure for module priorities, evalModules, makeScope, and fixed-point behavior. Minimal service and Home Manager schemas must declare their approximation openly. An identity mkIf or attrset-merging mkMerge cannot grade these tasks. Fake builders must implement the relevant override layer faithfully, including finalAttrs when supported by the chosen fixture. They must not reject a valid solution just because it uses another ordinary library helper.

Extract script text with evaluation and execute it only against temporary fixture directories and fake tools. Supply relevant real shell helpers, or compatible implementations with contract tests, instead of grading source substrings. Fail on attempted network use. Keep all generated paths, timestamps, environment variables, and test input orders controlled. Paths can vary between two fixed fixture layouts without randomness.

Target less than 10 seconds per evaluator and a hard cap comfortably below 60 seconds. These are design budgets, not measured timings. Run alternate probes in one evaluation where possible, and deep-force the graded projection rather than an entire package set or recursive module configuration. Each task still needs starter-fails, reference-passes, important valid-alternative contracts, and invalid mutations for every required criterion before activation.

1. flake-input-graph-identity

Category: flakes. Author difficulty: hard. Expected discrimination: high. Feasibility rank: 12.

Sources: offline rebuild dependency crawler, posts 1 to 2, 2026-09-24/26, same source path after follows, 2025-07-29.

Real failure: A dependency crawler used genericClosure with each input's outPath as its visited key. A respondent demonstrated that two occurrences of identical flake source can have different transitive inputs after follows overrides, causing the crawler to omit a needed source.

Trap: Deduplicating source paths before expanding graph nodes conflates source identity with resolved input-graph identity. Changing the key to the input's short name also fails when different parents use the same name; unrestricted recursion instead loops on aliases or cycles.

Prompt sketch: "The deployment bundle rebuilds offline on one host, but another tries to fetch a private input that should already be included. Repair collect-inputs.nix; the two hosts' local resolved lock graphs and the failing bundle inventory are included."

Evaluator: Supply a finite resolved lock graph with explicit node IDs, outPath, and edges, including follows arrays resolved relative to the lock root. Evaluate the collector and assert the exact set of reachable source paths, excluding unrelated nodes. One graph has distinct nodes with equal source paths but different children, a diamond, a cycle, and a non-flake leaf. Rename every node and reorder edges in an alternate graph; then change only one equal-source node's child and require the newly reachable path. Accept node-ID closure followed by source deduplication, a visited-node traversal, or an equivalent lock-graph resolver. Do not require the source thread's proposed crawler, which its author explicitly described as not fully verified.

Why models may fail: outPath looks like the obvious stable identity, and the initial published solution used exactly that shortcut. The finite lock-graph adaptation tests this distinction without claiming to reproduce every flake override rule.

2. context-preserving-trim

Category: nix-language. Author difficulty: medium. Expected discrimination: high. Feasibility rank: 1.

Sources: trim loses context, 2026-05-28/29, Nix string context reference.

Real failure: A user found that lib.strings.trim returned correct-looking text while dropping the context of interpolated store paths. The discussion traced this to regex captures and supplied lib.addContextFrom s (lib.trim s) as a workaround.

Trap: Comparing JSON or string equality misses the lost dependency. Replacing trim with a different regex repeats the failure; returning the untrimmed string preserves context but breaks the consumer's format.

Prompt sketch: "The generated config has the right resource filename, but deployment to a clean machine omits that resource. It started after normalize-config.nix began removing surrounding whitespace; repair that helper without changing the config format."

Evaluator: Use builtins.toFile to create two real, tiny local store inputs and interpolate them into strings. Assert both the trimmed text and equality of builtins.getContext with the original input, including whitespace-only and context-free strings. Alternate probes change both store inputs, include two dependencies in one string, and preserve internal spaces and newlines. Accept appendContext, addContextFrom, a context-preserving substring algorithm, or any equivalent transformation. No derivation build is needed. A local offline probe on Nix 2.34.8 confirmed that builtins.match loses this context and appendContext restores it; pin the evaluator version or retain a reproducing helper if upstream behavior changes.

Why models may fail: The rendered output and ordinary unit tests pass. The source establishes this invisible failure directly, though no particular frontier model was tested.

3. nested-npm-dependency-rebuild

Category: packages. Author difficulty: hard. Expected discrimination: high. Feasibility rank: 16.

Source: Sunshine nested UI override, especially posts 7 to 11, 2025-09-25/26.

Real failure: Updating a nested UI's src, postPatch, and npmDepsHash still reused the old npm dependency output. The accepted repair regenerated npmDeps with fetchNpmDeps, and follow-up advice exposed the separate distinction between a fetcher's hash argument and its derivation's outputHash.

Trap: The outer package's finalAttrs does not make every nested legacy builder override-aware. Setting hash through a plain derivation's overrideAttrs can leave outputHash unchanged; assuming .override exists also failed in the thread.

Prompt sketch: "The server uses the forked source, but its web UI fails with a package-lock mismatch against the old npm cache. Repair overlay.nix using the supplied offline release records and builder definitions; the server and UI must still ship together."

Evaluator: Supply an outer final-attributes builder and a deliberately legacy inner npm builder, documented in starter code. The fake fetcher records its effective source, patch-produced lockfile identity, and outputHash; a tiny shell fixture executes lockfile preparation. Assert that the final consumer uses UI assets built from the new source and that the effective npm cache matches those assets. Alternate releases change source, lockfile, and hash independently, with an existing UI patch and unrelated server metadata to preserve. Accept a fresh fetchNpmDeps, a correct outputHash override where exposed, or reconstructing the nested package through its supported builder interface. Do not assume modern and historical buildNpmPackage versions behave identically.

Why models may fail: Several plausible fixes, including maintainer suggestions, failed before the nested derivation boundary was identified. This is substantially deeper than checking a top-level cargoHash or npmDepsHash attribute.

4. override-nested-webapp-consumer

Category: overlays. Author difficulty: hard. Expected discrimination: high. Feasibility rank: 14.

Sources: Mattermost nested attribute, 2025-01-10/13, upstream repair PR #373085, nixpkgs override semantics.

Real failure: Overriding Mattermost's exposed webapp attribute did not change the web files installed by the full application. The install hook had already interpolated the old value from a recursive attrset, and the maintainer changed the packaging so the override reached the actual consumer.

Trap: Patching passthru or an attrset does not retroactively rewrite a previously evaluated Nix string. Returning the webapp alone makes its tests pass while silently dropping the server.

Prompt sketch: "Our patched web UI builds, and the package inspector shows it under webapp, but the installed server still serves the upstream page. Fix the small package so downstream webapp overrides affect what is installed."

Evaluator: Provide local old/new webapp directories and a fake override-capable mkDerivation. Evaluate the package, apply a downstream webapp override, execute the resulting install hook, and inspect both the installed UI marker and server executable. Alternate probes replace the webapp twice, add an unrelated install action, and override the server version separately. Accept shell-time $webapp consumption, a finalAttrs package definition, or a correct consumer-hook override. Preserve unrelated hook behavior and existing patches. Do not require eliminating the old dependency's context unless that is an explicit public requirement.

Why models may fail: The wrong answer looks idiomatic, and replies in the thread repeated it. Executing the consumer distinguishes it from the existing corpus's shallow overlay checks.

5. module-clone-definitions-not-values

Category: modules. Author difficulty: hard. Expected discrimination: high. Feasibility rank: 21.

Source: replicating a systemd service, posts 6 to 14, 2025-11-12.

Real failure: Assigning config.systemd.services.a = config.systemd.services.b produced infinite recursion. Respondents also identified copied generated names, conflicting environment defaults, and the mistake of merging already-merged option values into another service.

Trap: Treating an evaluated service as reusable plain data copies computed fields and defaults. Adding mkForce around the whole copy can suppress one conflict while retaining the wrong unit identity or duplicating list defaults.

Prompt sketch: "Adding a second worker based on the first makes evaluation recurse; an earlier workaround gave both workers the same unit name. Repair workers.nix so both instances keep the shared behavior and their own state directories."

Evaluator: Use real evalModules with a small service submodule whose generated name, environment, list defaults, and a read-only field model the documented problem. Evaluate two workers with different names, ports, and directories; require shared dependencies exactly once and distinct computed identities. Alternate probes add an operator-level list extension and rename both instances. Accept a shared pre-merge module/function, explicit reusable definitions, or a careful projection of permitted evaluated fields that passes all probes. Do not demand the exact allowlist shown in the thread. This tests the reduced module contract, not full systemd behavior.

Why models may fail: Attrset copying is an attractive repair, and error suppression can conceal the second-order failure. The thread documents both recursion and incorrect semantics after copying.

6. scoped-qt-overlay-propagation

Category: overlays. Author difficulty: hard. Expected discrimination: high. Feasibility rank: 18.

Source: fcitx5 package-set override, 2025-10-26 and 2026-02-01.

Real failure: The reporter tried nested assignment, //, and recursive updates while an LLM failed to supply a useful fix. A response showed an override of the actual qt6Packages scope, while the discussion explained that package search does not expose every package-set relationship.

Trap: Replacing one exported attr can leave scope-internal consumers closed over the old dependency; replacing the entire nested set also removes siblings. The thread demonstrates the failed approaches and the scope-based proposal, but does not include a final confirmation of the reporter's result.

Prompt sketch: "The patched input-method plugin appears when inspected directly, but the desktop bundle still pulls the broken one; another attempted fix loses breeze-icons. Repair the overlay against the supplied package-set definitions."

Evaluator: Build a real lib.makeScope fixture in which desktopBundle depends on scope.fcitx5-qt, and export a KDE-facing view of that scope. Check the transitive dependency marker, retained siblings, and existing patch order. An alternate scope changes the dependency marker and has a second consumer; apply a later overlay to test composition. Accept overrideScope, an equivalent recomputation of the scope, or explicit correct overrides of all consumers exposed by the task contract. Require behavior, not the token overrideScope.

Why models may fail: The direct package attr looks fixed, while the closed-over consumer remains broken. The reported LLM failure supports the problem choice, but the precise fixture is an authored reduction and needs calibration.

7. fhs-build-output-staging

Category: packages. Author difficulty: hard. Expected discrimination: high. Feasibility rank: 22.

Source: buildFHSEnv cannot write under the build output, 2025-03-12/16.

Real failure: A Node script could write outside a derivation but failed with EROFS when an FHS-wrapped build set HOME=$out. The accepted explanation pointed to the wrapper's read-only /nix bind, and running in the build directory fixed the problem.

Trap: $out is writable to the outer builder but becomes read-only inside this additional mount namespace. Generic chmod, sudo, or a global writable store bind addresses the wrong boundary.

Prompt sketch: "The asset generator works in the development shell, but packaging it fails creating $HOME/cache even though the install phase can write $out. Repair package.nix; the vendor tool still needs the provided FHS environment."

Evaluator: Evaluate phases with a fake FHS launcher that rejects writable working directories or HOME locations beneath the supplied output prefix and writes known assets elsewhere. Execute the phases in a temporary build directory, then inspect the installed assets. Alternate probes use an output path containing spaces, a different asset name, and a pre-existing build directory. Accept staging in the build directory or another explicit temporary directory followed by installation outside the wrapper. If supporting a narrowly scoped writable bind, implement its actual ordering and path semantics in the launcher and accept it; otherwise state the provided wrapper's fixed mount policy in starter documentation. No real bubblewrap or privileged mounts are necessary.

Why models may fail: The error invites ordinary permission fixes, but the same path has different permissions across the wrapper boundary. This adds an actual write-contract check to the existing FHS task shape.

8. nuget-sdk-dependency-reconciliation

Category: packages. Author difficulty: hard. Expected discrimination: high. Feasibility rank: 17.

Source: OpenRA .NET update and Copilot attempts, 2025-11-02.

Real failure: An SDK update left duplicate SDK-provided packages in the NuGet fallback cache, producing a symlink collision. Copilot changes got past that error but led to sandbox download failures; regenerating dependencies with the package's own fetch-deps workflow removed duplicates and restored missing dependencies.

Trap: Deleting the colliding path, forcing ln, or filtering all Microsoft.* packages hides the collision while losing required restore inputs. SDK-supplied identity includes normalized package ID, version, and platform applicability, not just a name prefix.

Prompt sketch: "After the SDK update, the offline restore first fails with ‘File exists'; the attempted cleanup instead gives NU1101 errors. Repair the dependency preparation helper using the included restore manifest and SDK inventory."

Evaluator: Supply small JSON restore graphs and SDK package inventories. Evaluate the Nix helper's generated dependency records and execute a fake fallback-cache builder that rejects duplicate normalized identities and missing required packages. Alternate inputs change SDK version and runtime identifier, use case variants of one NuGet ID, and include a required non-SDK Microsoft.* package. Accept inventory-aware filtering plus complete pinned records, or generation from the supplied canonical restore manifest. Provide the local manifest-generation interface so the task does not require a real .NET restore or knowledge absent from the project.

Why models may fail: The source contains an explicit failed Copilot attempt and a locally tempting workaround that changed the symptom. The challenge is preserving the complete dependency contract while removing duplicates.

9. encoded-store-path-retention

Category: purity. Author difficulty: hard. Expected discrimination: high, with evaluator caveat. Feasibility rank: 25.

Source: string context versus output references, 2025-06-12.

Real failure: An unattended installer used a percent-encoded flake store path and retained its Nix string context, yet the flake was absent from the installed image. The reply explained that output references depend on literal store-path bytes, so build-time context alone does not keep an encoded runtime reference alive.

Trap: appendContext repairs input dependency tracking but does not make the raw path appear in the produced file. The superficially advanced answer can therefore pass a getContext test and still fail deployment.

Prompt sketch: "The install script works on the build machine and passes its dependency check, but on the copied image disko-install cannot find the encoded flake path. Repair installer.nix while preserving URL handling for unusual path characters."

Evaluator: Evaluate script text and declared retention artifacts, execute the script against a recording fake installer, and inspect the produced files or symlink targets for a literal reference to the fixture source. Separately require context for any embedded source-derived content. Alternate probes change the source path and exercise URI-escaping characters in the supplied path suffix. Accept an unencoded safe path: reference with correct quoting, a retained literal path plus encoded URL, or an explicit dependency file/symlink included in the output contract. The reference scan is a deliberately reduced model, not proof of a real Nix runtime closure; author a contract test for each supported retention mechanism.

Why models may fail: It requires distinguishing evaluation context, derivation inputs, and output references. Rank it below the simpler context task for implementation priority because a fake closure scanner can easily overclaim correctness.

10. module-generated-jobs-fixed-point

Category: modules. Author difficulty: medium. Expected discrimination: medium-high. Feasibility rank: 5.

Sources: mapped Borg jobs recurse, 2026-04-11/12, list-shape variant, 2025-05-05.

Real failure: Mapping a configuration-provided job list into a root-level mkMerge caused recursion while the module system determined its configuration structure. Moving the generated definitions below a statically known option branch resolved the reported failure.

Trap: Wrapping the root merge in mkIf does not necessarily defer evaluation of the list's structure. A singleton literal list can work because its length is known without forcing its element, which makes a simplified reproduction misleading.

Prompt sketch: "One inline backup job works, but moving jobs into the host's options makes evaluation recurse. Fix backup-jobs.nix so any number of enabled jobs produces the expected sockets and timers."

Evaluator: Use real evalModules with typed job submodules and socket/timer output options. Deep-force only names and selected fields. Probe zero, one, and three jobs, disabled jobs, and an unrelated service contributed by another module. The metamorphic probe renames jobs and adds one job, requiring exactly one additional socket/timer pair. Accept merges under static branches, mapAttrs/listToAttrs construction there, or a submodule design with equivalent behavior. Preserve enable semantics and other modules' definitions.

Why models may fail: The diagnostic often points at imports even when the cause is root structure discovery. Familiarity with the standard mkIf advice can lead to a repair that remains recursive.

11. mkforce-across-option-reexport

Category: modules. Author difficulty: medium. Expected discrimination: medium-high. Feasibility rank: 4.

Source: GTK_IM_MODULE conflict despite mkForce, 2025-07-08/09.

Real failure: Forcing environment.sessionVariables.GTK_IM_MODULE did not fix a conflicting environment.variables.GTK_IM_MODULE definition. The successful repair applied priority at the destination option identified in the error.

Trap: A module reads the merged value of one option and defines another option from it; that value no longer carries the original definition's priority wrapper. Applying a stronger priority at the source repeats the mistake, while forcing the entire destination attrset can delete unrelated variables.

Prompt sketch: "The input-method override still conflicts after adding mkForce, and another workaround makes unrelated login variables disappear. Repair input-method.nix; the error and the module that forwards session variables are included."

Evaluator: Real evalModules evaluates a source option, a forwarding module, and a conflicting destination definition. Assert the intended destination value and preservation of unrelated environment variables and input-method settings. Alternate probes change the variable name and desired value through a public option and add another forwarded variable. Accept priority at the destination leaf, removing the conflicting source through a supported option, or changing a task-owned forwarding module to preserve the intended override contract. Avoid asserting a particular priority integer.

Why models may fail: "Use mkForce" is already present in the failure and is insufficient without following the option flow. The source explicitly verifies the failed and successful placements.

12. generated-palette-without-ifd

Category: purity. Author difficulty: hard. Expected discrimination: medium-high. Feasibility rank: 20.

Sources: Home Manager palette generation, 2025-01-26, readFile and IFD distinction, 2025-02-25.

Real failure: A Home Manager module defined a palette-producing runCommand but never consumed the option, so nothing ran. The discussion then exposed that its output was a directory rather than JSON text and recommended moving generation/merging out of evaluation or committing generated data.

Trap: Making the lazy value live by calling readFile on the derivation output introduces IFD. Reading a checked-in JSON file is fine; it is not true that every read from the store is IFD, despite an overbroad explanation in the thread.

Prompt sketch: "Theme generation appears to do nothing; forcing the value makes our evaluation-only CI fail because it tries to realize a builder. Repair theme.nix so the final application config combines the supplied generated palette with user overrides."

Evaluator: Supply a tiny local palette generator and a JSON merge tool through fake package paths. Evaluate with IFD disabled, capture the generated build/runtime script, execute it on two palette fixtures, and verify merged JSON with user values taking precedence while unrelated defaults survive. An alternate palette changes keys and includes nested objects; the module-disabled case must not force a deliberately throwing generator. Accept build-time generation connected as a file source, runtime generation through a wrapper, or a checked-in palette only when the task's supplied inputs make that palette valid. Do not accept a fixed palette for an input-dependent generator.

Why models may fail: Fixing laziness exposes a second phase-boundary error. This requires deciding what must be known during evaluation rather than merely suppressing IFD checks.

13. darwin-import-selection-phase

Category: darwin. Author difficulty: medium. Expected discrimination: medium-high. Feasibility rank: 8.

Source: shared Stylix module recurses, 2025-05-24.

Real failure: A shared module chose its NixOS or Darwin import using pkgs.stdenv.isDarwin, causing recursion. Advice to import both and enable one did not address the reporter's incompatible module families; the reporter moved the import selection to system construction.

Trap: pkgs can depend on module evaluation, so using it to determine the import graph closes a cycle. Importing both platform modules can also declare incompatible or unknown options even if their feature is disabled.

Prompt sketch: "The shared appearance module evaluates on its own but recurses when used by the Linux and Mac host constructors. Fix the composition so each host keeps its appearance settings and can still disable the feature."

Evaluator: Use two small platform module families with disjoint declarations and real evalModules; make pkgs a module argument resolved only after imports. Evaluate both host constructors and a disabled-feature case. Alternate probes rename hosts and swap an explicit platform parameter independently of the machine running the evaluator. Accept constructor-level imports, an early specialArgs discriminator independent of config, or separate platform adapters sharing data. Importing both is acceptable only if the candidate makes both schemas genuinely compatible, rather than hiding errors with disabled checks.

Why models may fail: Conditional imports and conditional definitions are easily confused. The source records an initially plausible recommendation that did not fit the actual platform boundary.

14. editor-project-state-isolation

Category: devshells. Author difficulty: medium. Expected discrimination: medium. Feasibility rank: 15.

Source: VSCodium project shells share the first instance, 2026-07-08.

Real failure: Opening a second project's wrapped VSCodium reused the first project's extension set. The AI-assisted investigation reported that distinct user-data directories fixed the observed behavior while sharing only selected human-authored settings; its IPC explanation was explicitly marked unverified.

Trap: Changing only --extensions-dir does not establish separate application state. Sharing all of User also shares mutable databases and workspace trust, and keying state by directory basename collides for different repositories with the same name.

Prompt sketch: "The editor launched from project B still has project A's extensions whenever A is open. Repair editor-wrapper.nix; user settings should remain shared, and reopening a project should retain that project's trust decision."

Evaluator: Capture the produced wrapper, execute it under temporary HOME/XDG directories, and use a fake editor that models instance identity by user-data directory and records argv. Launch two different absolute project paths with equal basenames, then reopen one; assert distinct persistent state, correct extension paths, shared settings/keybindings, and unshared globalStorage. Alternate probes set a custom XDG directory, omit optional settings, and pass arguments containing spaces. Accept per-project directories or a stable full-path-derived key, direct symlinks or a documented refreshed copy of the selected settings. Do not claim this fake proves real Electron IPC behavior.

Why models may fail: An extension-path-only fix looks directly related to the symptom but does not solve the observed reuse. Discrimination depends on the state-sharing probes; simple quoting alone would be weak.

15. runtime-tools-outside-devshell

Category: packages. Author difficulty: medium. Expected discrimination: medium. Feasibility rank: 13.

Sources: AI disclosure in claude-squad PR #424155, 2025-07-10, runtime wrapper correction, 2025-07-21.

Real failure: The Claude Code-generated PR put runtime requirements tmux and gh in nativeBuildInputs. A reviewer instructed the author to wrap the installed program's PATH, because availability during packaging does not provide runtime lookup.

Trap: Moving the tools to buildInputs still does not teach the program where to find them at runtime. Testing inside the dev shell masks the omission; wrapping the pre-rename binary can leave the exported command unwrapped.

Prompt sketch: "The packaged CLI passes its smoke test in nix develop, but the installed command cannot find its helper tools on a clean account. Repair package.nix, including the existing executable rename."

Evaluator: Use a fake main executable that invokes two helper names, extract and run install/fixup phases with wrapper support, and run the installed command with a minimal PATH containing neither helper. Verify each fixture helper was actually called and original argv/exit status were preserved. Alternate probes change both helper locations, include an unrelated executable to preserve, and put a wrong helper earlier in the caller's PATH. Accept a correct wrapper, patched absolute tool references, or another self-contained launcher. Do not require nativeBuildInputs to be empty or enforce a single wrapper hook.

Why models may fail: The source is explicitly AI-attributed, but this is common Nix knowledge. Expect only moderate discrimination unless the clean-runtime and rename probes matter.

16. home-manager-fish-state-exclusion

Category: home-manager. Author difficulty: hard. Expected discrimination: medium. Feasibility rank: 23.

Source: Fish universal variables remain read-only after recursive linking, 2025-05-20/21.

Real failure: Linking the whole Fish config directory made its universal-variable state unwritable. Adding recursive = true still failed because fish_variables remained among the individually managed files; the replies clarified that the state file itself must be excluded.

Trap: A writable directory does not make a linked store file writable. Copying an initial state file on every activation overwrites subsequent user changes, while excluding the whole directory loses declarative functions and config.

Prompt sketch: "Fish reports EACCES writing universal variables even after we made the Home Manager directory recursive. Repair fish-files.nix so declarative config updates still work and existing interactive state survives repeated activation."

Evaluator: Use a small Home Manager file-plan schema with real module evaluation, then materialize the evaluated plan in a temporary home using a fixture linker. Inspect that static files resolve to the expected source and fish_variables is an ordinary writable file or absent for Fish to create. Run two activations around a simulated state update and require that update to survive. Alternate input adds another static function, starts with no state, and changes the source state snapshot. Accept source filtering plus recursive linking, per-file declarations, or one-time initialization with explicit non-overwrite behavior. Do not require an activation hook if simply leaving state unmanaged satisfies the prompt.

Why models may fail: The thread directly demonstrates that the standard recursive-link fix is incomplete. This overlaps mutable-config-home-manager; include it only if the two-activation and recursive-source probes improve discrimination.

17. follows-preserve-python-compatibility

Category: flakes. Author difficulty: medium. Expected discrimination: medium. Feasibility rank: 10.

Source: follows changes Python dependency compatibility, 2025-10-16.

Real failure: Changing a package's nixpkgs follow target from a 24.11 input to a 25.05 input broke its Python fontconfig dependency. The reply offered three valid directions: keep the package's own pin, select an appropriate Python version, or port the package to the newer API.

Trap: Deduplicating all nixpkgs inputs is not semantics-preserving. Fixing only the Python executable can leave dependencies from an incompatible interpreter package set.

Prompt sketch: "A lockfile cleanup reduced duplicate inputs, but the font tool now fails while the rest of the system still works. Repair the flake/package composition using the included API compatibility fixtures."

Evaluator: Import local flake.nix as data and call outputs with fake old/new package sets, with a tiny offline resolver honoring its declared follows edges. Each interpreter and dependency carries an opaque ABI marker and a recorded API capability. Assert a compatible application closure while unrelated host tools use the chosen host input. Alternate probes change the host default Python and provide the compatible interpreter under a different supplied release record. Accept retaining an independent pin, choosing a coherent compatible interpreter set, or porting the local adapter so fixture calls succeed on the new API. Do not grade a specific count of nixpkgs nodes.

Why models may fail: "Make everything follow nixpkgs" is common generic advice. The thread confirms breakage, but a frontier agent may solve the reduced case easily once it inspects compatibility metadata.

18. fetcher-stale-fixed-output-cache

Category: fetchers. Author difficulty: medium. Expected discrimination: medium. Feasibility rank: 7.

Source: jai-jail version bump correction, 2026-06-05.

Real failure: A patch bumped the package version without updating the source hash. Reviewer andersk warned that Nix could continue using old source matching the unchanged hash rather than fetching the newly named release.

Trap: Successful evaluation and an updated URL do not prove updated contents. With the same fixed-output identity and cached output, changing the URL alone can preserve old data; on a cache miss the same mismatch instead fails.

Prompt sketch: "The package says version 0.3 and has the new release URL, but a warm-cache build still contains version 0.2. Repair package.nix; the repository includes verified release hashes and source markers."

Evaluator: Model the effective fetch result by fixed-output name, hash mode, algorithm, and hash, with a local release-content table. Check that declared version, requested revision, effective source marker, and the authoritative digest agree in both warm-cache and cold-cache fixtures. Alternate probes vary the release and use equal source names across releases; another changes only URL while preserving the hash and must still expose stale content. Accept hash or compatible sha256, a different fixed-output fetcher with the same verified contents, and reconstruction through an overridable source argument. Do not pretend the fake fetcher computes a network hash or accept arbitrary SRI-looking strings.

Why models may fail: The source documents a real review correction, but not AI authorship of this patch. It improves on fetcher-source-pin by checking content identity rather than field shape; expected discrimination remains moderate.

19. samba-runtime-secret-provisioning

Category: purity. Author difficulty: hard. Expected discrimination: medium. Feasibility rank: 24.

Sources: Reddit Samba setup discussion, 2026-08-23/24, Samba configuration reference.

Real failure: A user reported that Gemini supplied instructions to include a Samba password in configuration.nix after Claude and ChatGPT had suggested manual password setup. Replies recommended runtime secret management and explicitly noted that configuration-embedded passwords can end up in the Nix store; the actual generated recipe was not posted, so its precise defect is unverified.

Trap: builtins.readFile on a secret followed by writeText or string interpolation embeds its bytes in store-bound artifacts. A mode-0600 installed file does not undo that exposure; a runtime path must remain a path until the provisioner runs.

Prompt sketch: "The Samba account is not recreated on a fresh machine, and our attempted automation exposes the password in the generated script. Repair samba-account.nix to use the supplied runtime credential file and preserve an unattended setup."

Evaluator: Evaluate while the runtime credential file does not exist, inspect all emitted scripts/configuration, and assert that none contains the fixture password. Then create the file only for execution and run the emitted provisioner against fake smbpasswd/account-query tools that record stdin and argv. Alternate probes change username, secret path, and password content with percent signs, backslashes, and shell metacharacters, then run twice with an existing account. Assert no secret in argv/logs, correct input delivery, and the task's explicit existing-account policy. Accept systemd credentials, a sops/agenix-provided runtime path, or equivalent deferred file reads. Do not require real decryption, Samba, privilege changes, or a live account database. Do not copy the wiki activation example blindly: it puts the password into the printf format string and shows one file used for both a Unix hashed-password option and Samba password input. The fixture must distinguish those credential formats and require literal byte handling.

Why models may fail: It spans evaluation and runtime, but the security rule is familiar. The source supports the concern and reported advice, not a verified exploitable Gemini implementation; rank accordingly.

20. mkshell-cross-toolchain-hooks

Category: devshells. Author difficulty: hard. Expected discrimination: medium. Feasibility rank: 19.

Sources: Claude's packages/buildInputs explanation corrected, 2025-02-20, nixpkgs dependency roles.

Real failure: Claude 3.5 Sonnet claimed that packages only adds binaries to PATH while buildInputs supplies deeper library and manual integration. Replies explained that packages feeds nativeBuildInputs and noted that choosing a compiler requires the appropriate stdenv, rather than merely adding GCC to the shell.

Trap: The native case masks build/host confusion, and putting every dependency in packages is also wrong for an explicitly cross-compiling shell. A library's existence does not imply automatic LD_LIBRARY_PATH configuration.

Prompt sketch: "The shell works for native builds, but the cross shell tries to run the target's code generator and still invokes the old compiler. Repair shell.nix using the supplied toolchain and library roles."

Evaluator: Use a pinned mkShell reduction with real package-to-native-input translation and stdenv override behavior. Fake tools carry build/host markers and setup hooks; run the generated shell setup and a fake compile invocation. Check that the generator executes on the build platform, linked libraries match the host platform, and the selected compiler belongs to the requested stdenv. Alternate probes swap build/host platforms and include a package needed in both roles. Accept packages or nativeBuildInputs for executable tools and equivalent explicit setup preserving hook semantics. Do not require a particular list location when two routes normalize identically.

Why models may fail: The source verifies a confident false explanation by a named model. Current frontier models may know the distinction, so the cross and compiler-selection constraints are authored extensions that need calibration.

21. gc-generations-shared-inodes

Category: debugging. Author difficulty: hard. Expected discrimination: medium, with scope caveat. Feasibility rank: 26.

Source: garbage-collection size estimates and Claude advice, 2026-06-26/28.

Real failure: Suggested --print-dead | du | awk recipes, including one attributed to Claude, did not match the space freed by nix-collect-garbage -d. Replies identified both generation-root removal and store hard-link sharing as reasons simple path-size sums fail.

Trap: Currently dead paths are not the same as paths becoming dead after old generation roots are removed. Summing all removed paths double-counts shared inodes and counts bytes that remain linked from retained paths.

Prompt sketch: "Our cleanup preview disagrees with the supplied before/after store inventory and sometimes counts a retained shared file as freed. Fix estimate-gc.nix; it must remain a read-only preview."

Evaluator: Give a bounded JSON/Nix snapshot with store-reference edges, current/old generation roots, external roots, and each file's inode ID and allocated-byte count. Define an intentionally limited GC model with no implicit derivation retention and no hidden links. Assert removed-path sets and bytes whose final link is removed. Alternate probes share one inode across live/dead paths, keep an old generation via an external root, add a dependency cycle, and permute paths. Accept any correct reachability and inode-accounting algorithm; the result must be labeled as a snapshot-model estimate, not a guarantee of real disk bytes. No actual garbage collection or system inspection occurs.

Why models may fail: The source verifies misleading recipes, but a good coding model may solve the explicit graph problem readily. Real Nix GC fidelity is out of scope, and this risks measuring general algorithms more than Nix; keep it as a reserve.

22. module-import-source-before-pkgs

Category: modules. Author difficulty: medium. Expected discrimination: medium. Feasibility rank: 6.

Source: fetchFromGitHub in Home Manager imports recurses, 2025-09-29.

Real failure: Replacing an evaluation-time source fetch with pkgs.fetchFromGitHub caused recursion while importing Home Manager modules. The reply explained that pkgs is itself supplied through the module system and cannot be used to determine that same module import graph.

Trap: Fixed hashes make a fetch reproducible but do not move it to an earlier evaluation phase. Passing the source through _module.args can preserve the same import-time cycle, and realizing the fetched derivation introduces an additional IFD dependency.

Prompt sketch: "The test host began recursing after we standardized all source fetching on fetchFromGitHub. Repair the module composition using the already-vendored Home Manager fixture; evaluation must not realize a builder."

Evaluator: Real evalModules supplies pkgs via a module and imports a local dependency module from the editable entrypoint. Evaluate with IFD disabled and a fake package fetcher that throws if forced. Alternate probes relocate the supplied dependency source and add another ordinary pkgs use in the module's config, which must still work. Accept a constructor-supplied source through specialArgs, a lexical parameter, a relative vendored import, or another early local input. Do not demand actual builtins.fetchTarball, which would violate offline evaluation.

Why models may fail: Generic "use a pinned fetcher" advice ignores when imports are resolved. This is more specific than the existing extraSpecialArgs task, though the recursion pattern is familiar.

23. python-build-backend-false-lead

Category: debugging. Author difficulty: medium. Expected discrimination: medium-low. Feasibility rank: 9.

Source: whisper-overlay and ChatGPT's Python-version diagnosis, 2025-09-30.

Real failure: ChatGPT led the reporter to attribute a package failure to Python 3.13 versus 3.12. A respondent followed the actual error to a missing Python build format/backend declaration and identified the dependency derivation that needed repair.

Trap: Downgrading the interpreter, changing a service module, or patching only the top-level application leaves the failing dependency's build contract unchanged. The stack includes several module and flake-parts frames that are not the cause.

Prompt sketch: "Enabling the speech service fails during evaluation; the previous attempt switched Python versions but the same error remains. Repair the local package definitions so the service evaluates with the selected interpreter."

Evaluator: Provide a small nested application/dependency graph and a Python builder that checks the backend contract selected by local project metadata. Evaluate the final service package with two interpreter markers and assert coherent dependencies and the correct build backend for the failing leaf. Alternate probes rename that leaf and change the selected interpreter without changing its build metadata; unrelated leaf metadata must survive. Accept pyproject plus the required build-system dependency or another backend mode explicitly supported by the pinned builder. A compatible older pin is valid only if permitted by the project contract; do not silently require the latest interpreter.

Why models may fail: This is an explicit ChatGPT false lead, but the error itself supplies much of the answer. It is a lower-priority calibration candidate and overlaps the existing Python packaging task.

24. package-offline-test-selection

Category: packages. Author difficulty: medium. Expected discrimination: medium-low. Feasibility rank: 11.

Sources: QLever PR #512128 AI disclosure, 2026-04-21, requests-sse tests require internet, 2026-04-22.

Real failure: In an explicitly LLM-assisted packaging PR, tests for the included requests-sse dependency eventually ran but all failed because they needed internet access. The author stated that those tests would be explicitly disabled.

Trap: Enabling upstream tests indiscriminately violates the build's network boundary; disabling every check loses useful offline tests when the package has a mixed suite. The mixed-suite requirement is an authored extension, since the cited comment says all tests in that dependency needed internet.

Prompt sketch: "The package works, but its check phase stalls in the sandbox; disabling the phase also hides a failing local parser regression. Repair test selection in package.nix using the included test inventory."

Evaluator: Evaluate the check configuration and execute its command against a small fake runner with offline and network-tagged tests. Require all offline tests to execute, all network-tagged tests to be excluded, and an injected offline failure to propagate. Alternate probes add another network-tagged test under a different name and a new offline test, preventing a fixed name list from satisfying an inventory-based contract. Accept runner markers, explicit dynamically derived selections, a supported exclusion mechanism, or local replacements that exercise the same assertions without networking. If all tests genuinely need network, disabling that entire suite must remain a valid alternative.

Why models may fail: The patch provides AI-assisted evidence, but selective test execution is familiar. Avoid promoting this merely because the word "network" appears; its value is the failure-propagation and newly added test probes.

25. source-filter-stable-store-name

Category: purity. Author difficulty: medium. Expected discrimination: medium-low. Feasibility rank: 3.

Sources: nix.dev reproducible source paths, Nix filterSource warning.

Real failure pattern: nix.dev documents that an unnamed local source path incorporates the checkout directory's name into its store identity. The Nix manual separately warns that filtering an already-store-backed source can preserve an input-derived name, so excluded-file changes can still trigger rebuilds.

Trap: Correct filtered file contents do not imply a stable source path. Replacing ./. with filterSource can retain the very name dependency that the task needs to eliminate.

Prompt sketch: "The source contents are unchanged, yet cloning the project into another directory or editing an excluded note changes the package input path. Repair source.nix without dropping build-relevant files."

Evaluator: Create equivalent temporary trees under two different basenames and evaluate the candidate's source constructor. Compare the actual resulting store paths and inventories, then change an excluded file and require identity stability. An alternate probe changes a required source file and requires a changed identity, defeating constant-path answers. Include empty directories and nested excluded files according to a public filter policy. Accept named builtins.path, cleanSourceWith with a stable name, a fileset conversion with equivalent semantics, or another local source constructor. Vendor any needed library; no build or fetch is involved.

Why models may fail: A source-content-only test misses the bug. This is documentation-derived, not a recent AI failure, and should remain a reserve until pilot results show discrimination.

26. argument-default-forwarding

Category: nix-language. Author difficulty: easy. Expected discrimination: medium-low. Feasibility rank: 2.

Sources: official wiki, default values are not bound in @ syntax, revision dated 2026-05-20, nix.dev scope pitfalls.

Real failure pattern: The wiki shows that args@{ a ? "a" }: args called with {} returns {}, rather than an attrset containing the default. A wrapper forwarding args therefore loses defaults that are available as local parameter bindings.

Trap: Adding defaults to the destructuring signature does not normalize the forwarded attrset. Merging defaults in the wrong order overrides caller values, while using with can select an outer lexical binding instead of a desired package attr.

Prompt sketch: "The wrapper works when every option is specified, but its advertised defaults disappear in the downstream builder. Repair wrapper.nix so omitted options and explicit overrides behave consistently."

Evaluator: Use a recording downstream function and evaluate omitted, partial, and fully explicit arguments. Assert default values, preservation of permitted extra arguments, and caller precedence. Alternate probes vary a dependent default such as a filename derived from name, use explicit false, and distinguish a supplied null from an omitted field where the public type allows it. Accept explicit forwarding of normalized bindings, defaults merged with caller args followed by dependent-default computation, or another equivalent wrapper. Do not require a specific syntax or inject an unrelated scoping puzzle.

Why models may fail: The syntax invites an incorrect mental model, but this is a small and widely documented language quirk. It is an inexpensive reserve, not a strong claim about frontier-model failure.

Search coverage and exclusions

The web search used Discourse's public search JSON endpoint with a 2025-01-01 lower date bound for ChatGPT, Claude, LLM, Copilot, AI, infinite recursion, overlay, overrideAttrs, follows, specialArgs, mkForce, mkMerge-related threads, string context, IFD, impure, fetchFromGitHub hash, buildFHSEnv, hardening, nix-darwin, and home-manager. I read the selected threads beyond search snippets. Google returned a CAPTCHA; DuckDuckGo and Reddit's own search were usable. Reddit dates were read from post timestamp metadata rather than inferred from relative labels.

GitHub research located and read maintainer review comments, then the selected wrapper, hash, and test-failure claims were checked against the GitHub API. The proposed Yashiki tag task was dropped after inspecting the diff: the initial rev already had the correct yashiki-v prefix, so describing that review as a fabricated release-tag failure would overstate the evidence. Redundant Darwin entries in lib.platforms.unix ++ lib.platforms.darwin were also too weak a semantic failure to justify a task.

The requested old Nix_Pitfalls wiki URL was inaccessible, and that title returned 404 on the official wiki. I used the official Nix_Language_Quirks page and nix.dev's best-practices page, alongside the nixpkgs manual's override and dependency warnings. These documentation sources have no demonstrated model failure and are labeled accordingly.

Excluded or deferred:

  • Technitium DynamicUser permissions has a 2026 workaround but no isolated verified cause in the thread. An eval-only check of hardening flags would be too weak to certify a permissions fix.
  • Python dependency overlay for searxng had no resolution in the retrieved thread. Prefer the scoped Qt candidate with a concrete response and an explicit local fixed-point model.
  • Container mountPoint impurity shows a quoted string in the reported configuration and does not establish a path-literal fix. Do not turn it into a misleading "just add quotes" task.
  • Input-remapper compilation detour directly documents ChatGPT failure and the missing service enablement, but package lookup and service-option tasks already cover most of it.
  • Hardware freezes, driver regressions, live portal discovery, and real systemd namespace behavior need execution environments beyond the stated evaluator budget. A matching-looking attrset is not evidence that those runtime problems are fixed.

The strongest next step is calibration of candidates 1 to 8 with real alternate solutions and plausible wrong patches. An unsolved community question or an older model's mistake is evidence for investigation, not evidence that today's frontier agents will fail.