Research date: 2026-10-03. Sources are benchmark papers, official repositories, maintainer publications, and statistical documentation. Numeric recommendations below are proposed NixBench policy, not universal thresholds endorsed by these sources.
NixBench should preserve its reproducibility and governance machinery while replacing most ranking tasks with empirically calibrated repair problems graded through real Nix behavior. More repetitions alone cannot make universally solved tasks discriminate.
This report starts from README.md, benchmark design, governance, scoring, and calibration records. The checked 2.0.0 calibration has three configurations with one observation each; every discrimination estimate is unavailable. Its approvals explicitly support activation, not stable ranking. Current documentation already specifies repeated studies, Wilson intervals, leave-one-task-out discrimination, independent contract fixtures, lifecycle transitions, and immutable digests. Strengthen these mechanisms rather than commission replacements.
One distinction matters for the redesign: required criteria determine binary success, but do not mathematically prevent partial credit. The calibration records already contain 75/100 failures. The opportunity is to make criterion scores informative and behaviorally independent, while reporting them separately from complete success.
1. Task selection, saturation, and discrimination
Mature benchmarks use several complementary methods. Expert review establishes that a problem is legitimate; empirical screening establishes whether it distinguishes the systems being compared. Neither substitutes for the other.
| Benchmark | Selection and calibration method | Implication for NixBench |
|---|---|---|
| SWE-bench Verified | Human reviewers selected 500 instances with clear descriptions, correct tests, and enough information to solve them. Verified is a validity intervention, not principally a harder replacement. | Audit solvability and acceptance boundaries before interpreting failures as difficulty. |
| SWE-bench Pro | The original benchmark spans 1,865 tasks in 41 repositories, with public, held-out, and commercial partitions. Tasks involve realistic, often multi-file work requiring hours or days for professionals. The V2 repository documents 642 validated public tasks, down from 731, and a HARD-51 subset selected using failures across multiple model families after removing ambiguity. | Increase dependency reasoning and repository navigation; remove invalid difficulty before selecting hard tasks. Preserve selection provenance. |
| LiveCodeBench | Continuously collects dated contest problems from LeetCode, AtCoder, and Codeforces, with versioned releases and selectable evaluation windows. Its documented generation protocol uses ten samples and reports pass@1 and pass@5. | Maintain a task pipeline and dated cohorts, rather than repeatedly scoring one static release. Contest ratings are an external difficulty cue, not proof of agent difficulty. |
| Terminal-Bench 2.0 | Selected 89 of 229 contributed tasks using author difficulty estimates and three human reviewers. Authors investigated model failures for task defects. Empirical difficulty uses solve-rate bands below 33.3%, 33.3–66.7%, and at least 66.7%; these remain separate from human difficulty. | Combine independent review with pilot runs. Genuine environmental complexity is preferable to missing APIs or brittle graders. |
| Aider polyglot | Replaced a saturating Python benchmark by testing 697 exercises in six languages with seven models, retaining 225 solved by at most three models. | Use an explicit selection panel spanning abilities. Freeze selection before evaluating the next model to limit selection bias. |
| HumanEval+ / EvalPlus | Expanded HumanEval tests approximately 80-fold using model-generated and mutation-generated inputs. Stronger tests exposed previously accepted incorrect programs and changed rankings. | Some apparent saturation is evaluator weakness. Strengthen behavioral coverage before inventing harder prompts. |
| BigCodeBench | Curated 1,140 tasks using 139 libraries, with expert review and behavioral tests averaging 99% branch coverage. Its Instruct variant removes nonessential docstring detail. Independent humans solved 32 of 33 sampled tasks. | Test composition across Nix facilities and validate that concise requests remain solvable. Coverage alone does not validate the specification. |
| τ-bench | Constructs tasks from domain policies, APIs, and database states. Iterative trajectory inspection removes ambiguity; its original retail construction used more than 40 pilot trials per task and inspected low-success cases. | Low success triggers evaluator and prompt review, rather than automatic promotion to “hard.” |
| MLE-bench | Screens Kaggle competitions for relevance, reproducible grading, and viable train/test splits. Its 75 tasks have human-effort complexity bands below two hours, two to ten hours, and above ten hours; seven additional competitions form a development split. Medal thresholds provide external performance anchors. | Record expert repair time and independent human success, alongside empirical agent rates. Keep development tasks separate. |
| Cybench | Selects 40 professional CTF tasks across categories and uses human first-solve times as difficulty evidence. Seventeen tasks also have intermediate subtasks; guided and unguided outcomes are separate. | Measure useful intermediate outcomes without turning the primary repair prompt into a solution checklist. |
What NixBench should do
Build an initial pool of 60 candidate tasks, aiming for a 40-task discriminating release and expansion toward 60 independent active tasks. Keep the existing corpus as a development/regression suite. Favor multi-file repairs involving module precedence, overlay fixed points, pinned dependency migrations, generated configuration, and offline package execution.
Pilot every candidate with six materially different configurations across at least three model families, five trials each: 30 observations per task, or 1,800 attempts for the pool. Require an independent human solution and two reviewers for prompt/test agreement. Target at least 70% of active tasks with panel-average solve rates between 20% and 80%; reserve roughly 15% each for easier anchors and hard extensions. These are portfolio targets, not automatic rejection rules. Prioritize underrepresented categories already named in governance; seek five independent tasks per reported category.
2. Statistical design and retirement rules
For task-specific success probability p, pass@k is the chance of at least one success in k independent attempts, 1 − (1 − p)^k. Pass^k is the chance all k attempts succeed, p^k. τ-bench distinguishes capability under repeated attempts from reliability. Average these quantities across tasks; applying either formula to the corpus-average success rate gives a different result.
With c successes among n independent trials, use the estimators 1 − C(n−c,k)/C(n,k) and C(c,k)/C(n,k), respectively, only when n ≥ k. Aider's repair after test feedback is a sequential interaction, not independent pass@2. NixBench's agent can already retry inside one run; count that complete run as one attempt. Pass@k also assumes an oracle can identify successful candidates, so it is not the deployment success rate of an unvalidated candidate selector.
Adding Error Bars to Evals, by Anthropic's Evan Miller, recommends reporting standard errors and sample counts, accounting for related-question clusters, comparing question-level paired differences, and planning sample sizes through power analysis. Resampling answers reduces within-question noise but cannot remove between-question variation. Its example finds diminishing returns from additional samples; it does not prescribe five trials universally. Lowering temperature solely to suppress noise changes the estimand. Next-token probability scoring, another recommendation, is unsuitable for multi-step coding agents.
NixBench's existing Student-t interval concerns repeated runs of a fixed corpus. Keep that interpretation. Task resampling remains a composition-sensitivity analysis, not a confidence statement about all Nix work. Compare configurations on the same tasks and use paired task differences for composition sensitivity. For fixed-corpus run uncertainty, use genuinely matched run blocks or independent-run variance estimates; arbitrary pairing of trial numbers does not create an experimental pairing.
Use Wilson intervals for each configuration/task pass rate. Five successes in five trials yield a two-sided 95% interval of approximately 57–100%; ten successes yield 72–100%. An observed 100% is weak evidence of universal reliability. At 29 tasks, one task changes the single-run rate by 3.45 percentage points. Displaying extra decimals cannot supply missing evidence.
The existing corrected item-total statistic is appropriate for screening. Point-biserial correlation relates binary task success to a continuous score; excluding the task from the total prevents mechanical self-correlation. Repeated runs and related model families still create dependence, and a task with zero outcome variance has undefined discrimination, not a measured coefficient of zero.
tinyBenchmarks demonstrates item response theory for modeling item difficulty and discrimination using historical model responses. A two-parameter logistic model can distinguish the ability level where an item becomes solvable from how sharply success changes with ability. NixBench's three configurations cannot support credible item-level slope estimates. Repetitions improve success estimates but do not create new independent ability levels. Also, specialization across Nix domains can violate a one-dimensional ability model.
What NixBench should do
Require five complete valid trials per published configuration, replacing the current budget exception for ranking claims. Use ten for preregistered close comparisons and for pass@5/pass^5 reporting. Keep pass@1 primary. Randomize task order, reset workspaces, record model/scaffold/effort identities, and predeclare budgets. Retain invalid attempts separately; valid agent timeouts remain failures.
Use corrected discrimination r ≥ 0.20 as a provisional retention target. Review r < 0.10 and every negative value for ambiguity, specialization, or grader defects. Start this review at 30 observations across six configurations, with uncertainty and family composition visible; do not turn a noisy correlation into an automatic exclusion rule. Defer IRT until dozens of diverse configurations exist and validate predictions on held-out model families.
Flag saturation when every reference configuration scores at least 95% over ten or more trials. Review again next release, inspect interval bounds and accepted alternatives, then retire from ranking while retaining regression coverage. Below 5% pooled success, require another independent human solve and an evaluator audit. Freeze retirement decisions before new comparison results. Plan close comparisons for 80% power at a five-percentage-point difference using pilot variance; if the budget is inadequate, report the detectable difference rather than a definitive rank.
3. Contamination, freshness, and prompt hygiene
LiveCodeBench records problem publication dates and supports post-cutoff evaluation windows. LiveBench explicitly designs for monthly question refreshes, recent source material, objective answers, and release-specific evaluation. These reduce exposure opportunities; neither freshness nor secrecy proves absence from training or retrieval.
SWE-bench Pro separates repositories across public, held-out, and commercial sets. MLE-bench supplies a development split, constructs new data splits, and investigates both solution similarity and memorization. Its paper distinguishes copying a public solution from learning a general strategy. NixBench should make the same distinction. Removing a URL does not remove a recognizable issue title, exact error passage, or duplicated repair pattern.
Prompt hygiene must preserve fair disclosure. A symptom-driven request should state the observed failure, intended user behavior, constraints, and available evidence. The agent should discover the implementation. BigCodeBench-Instruct is a useful precedent for shorter intent-oriented prompts; Terminal-Bench's verification process also requires tested behavior to be adequately specified. Concealing essential requirements simply manufactures failure.
What NixBench should do
For the next major corpus, remove all 18 diagnosed source links from agent-visible prompts and starters. Keep attribution and provenance in evaluator-side records. Audit issue numbers, distinctive quotations, comments, filenames, and generated task metadata for answer retrieval shortcuts.
Split by underlying issue, repository, and repair family before creating variants. Keep all variants of a source issue in one split. Start with 40 held-out ranking tasks plus at least 20 distinct public development tasks. Preserve the existing private storage and trusted-launcher requirements; do not infer that hiding evaluator paths is isolation.
Review freshness quarterly and target replacement of 20–25% of the held-out corpus per cycle, subject to calibration quality. Publish immutable cohort identities and retain old results separately. Describe public results as development performance. Retire and disclose private tasks only through existing authorization rules.
Rewrite checklist prompts as requests such as “the service fails after enabling this module; restore startup and preserve the existing firewall behavior.” Supply reproducible logs and pinned dependencies, but omit the prescribed option names or patch sequence. Freeze network policy: either supply offline documentation, or allow a declared documentation environment with retrieval logged. Never mix these protocols in one ranking.
4. Evaluator design
Semantic grading asks whether the resulting system behaves correctly. Structural grading asks whether it resembles a chosen implementation. Nix attrset assertions can be semantic when they inspect evaluated observable configuration, but fake module libraries can bypass merge priorities, type checks, laziness, and fixed-point behavior. Passing such tests need not imply a usable NixOS configuration.
EvalPlus shows that additional adversarial inputs expose false positives. Its repository history also records oracle and contract fixes, demonstrating that more tests can preserve mistaken assumptions. Reference equality is only valid where outputs are uniquely specified. Different derivation structures, dependency ordering, or helper choices can produce equally correct behavior.
BigCodeBench deliberately checks behavior while allowing different function choices. τ-bench checks final database state rather than requiring a reference trajectory, but acknowledges that correct final state can miss policy violations. NixBench similarly needs explicit impurity or offline-execution checks where final files alone cannot establish compliance.
Property-based testing exercises generated inputs and edge cases. Metamorphic testing checks relationships across transformed inputs without requiring one complete reference output. For Nix, adding an unrelated package should preserve existing selected packages; renaming a module instance should rename only corresponding generated resources; disabling a service should remove its own activation effects. Such properties must follow the task contract. Reordering modules, for example, is not universally semantics-preserving.
LLM judges are useful for subjective dimensions or failure analysis, but MT-Bench's judge study documents position, verbosity, and self-enhancement biases. Terminal-Bench uses calibrated LLM judgments for error analysis while retaining executable task verification. That is a useful separation for NixBench.
What NixBench should do
Create three explicit evaluation profiles: fast pure Nix evaluation, pinned real module evaluation, and sandboxed package build/runtime verification. Use actual pinned nixpkgs/Home Manager module implementations where their semantics matter. Preload dependencies and measure warm/cold setup separately. Real builds need a supported evaluator environment, not access granted to the agent through the forbidden host daemon socket.
For devshell-tooling-contract, execute the emitted hook in a clean shell and inspect resulting values. For string-escaping-systemd, execute the generated script against controlled arguments and environment. For overlay-override-package, evaluate against at least three package variants and check preserved metadata plus observable behavior. For packaging, run the built executable and test outputs and offline operation; regex presence of a build input is insufficient.
Require at least two independently authored valid alternatives per task, including a materially different Nix idiom. Retain targeted negative fixtures for every required criterion. Add boundary cases and at least two justified metamorphic properties where applicable. During authoring, generate roughly 100 small inputs, then freeze a reviewed, seed-versioned test set for deterministic release grading. Test that deliberate mutants fail and valid alternatives survive.
Keep binary complete success primary and macro normalized criterion score secondary. Use three to six independently evaluated behavioral criteria; do not award points for merely attempting a named technique. Required functional criteria can still contribute partial points on failed tasks. Keep optional criteria within the current objective maintainability/formatting policy unless governance explicitly changes. Human-reviewed or rubric-judge assessments of issue-report usefulness should remain a separate diagnostic, calibrated against at least 50 diverse human-rated examples with disagreements retained.
5. Reporting and presentation
The strongest leaderboard designs expose evidence behind totals:
| Source | Useful presentation practice |
|---|---|
| Terminal-Bench and its 2.0 paper | The current site labels 95% confidence whiskers and displays agent/model identity. The paper provides per-task model grids and runtime/token distributions. Live site versions advance, so cite the exact corpus. |
| SWE-bench | Separate benchmark subsets, resolved-instance matrices, repository breakdowns, and resolution-versus-cost views. Its default Verified comparison also identifies the common agent environment. |
| Aider | Percent correct alongside cost, edit format, and drill-down fields for first/second attempts and seconds per case. These identify operational tradeoffs hidden by correctness alone. |
| LiveBench | Release selectors and category/task breakdowns keep results interpretable across changing test sets. |
| Epoch AI | Separates internally administered results from external reports, cites provenance, provides downloadable data, and documents repeated runs. Its described main-plot bars are plus/minus one standard error, not 95% intervals. |
| Artificial Analysis | Presents intelligence against cost and time per task, with token accounting and a versioned methodology. Aggregate index precision should not be transferred to individual benchmark scores. |
What NixBench should do
Make the public report a task-by-configuration heatmap. Each cell should show successes/trials, a Wilson interval, timeout count, and links to criterion failures and traces. Distinguish zero success from missing or invalid measurement. Provide sortable failure sets: always failed, intermittent, solved only by A, solved only by B, and shared failures. Show overlap counts or Jaccard similarity alongside pairwise score differences; identical totals can hide entirely different weaknesses.
Above the matrix, display pass@1, its explicitly named uncertainty method, task/trial counts, invalid-attempt rate, and timeout rate. Offer pass@k and pass^k selectors only when sampling supports them. Keep rubric score separately labeled. For active private tasks, this matrix remains operator-only: existing governance forbids publishing even task hashes or per-task outcomes. Public held-out reports retain the aggregate allowlist and minimum-stratum rules.
Add accuracy-versus-total-cost and accuracy-versus-wall-time plots with Pareto frontiers. Include failed attempts in cost accounting, distinguish agent time from evaluator/setup time, and stratify timing by environment identity. Record actual token charges, cache accounting, pricing date, and whether monetary cost is estimated.
Every view and export should carry corpus digest, protocol version, scaffold/model/effort identity, run dates, and network policy. Never merge changed acceptance boundaries into one trend line. Export public task-level CSV/JSON and pairwise difference intervals; describe unresolved comparisons as inconclusive, not equivalent. A rank number without these details is too strong a claim for this corpus.
Prioritized top ten
- Replace ranking tasks with a reviewed candidate pool targeting 20–80% solve rates; retain saturated tasks for regression use.
- Grade real Nix semantics and candidate execution, prioritizing modules, shell hooks, overlays, and package runtime behavior.
- Require five valid trials per published configuration; use ten for close comparisons and reliability metrics.
- Remove source-answer links and implementation checklists while preserving explicit behavioral requirements.
- Calibrate with six configurations across at least three model families; freeze selection before subsequent comparisons.
- Require two independent valid alternatives, targeted negative fixtures, and justified property checks per task.
- Report pass@1, criterion scores, paired differences, and correctly labeled intervals; defer IRT until the response matrix is adequate.
- Publish public failure matrices and disagreement sets, with private equivalents restricted to authorized operators.
- Apply reviewed saturation/discrimination thresholds and quarterly held-out rotation through existing lifecycle governance.
- Add versioned cost/time comparisons and downloadable evidence without pooling corpus identities or rewriting historical results.