Scoring policy (corpus 3.0)
**The headline metric is the task pass rate.** A task observation passes when every *required* criterion passes and the agent did not time out. For each configuration, reports lead with the share of valid task-trial observations that pass, shown with a 95% Wilson interval. Repeated-trial reliability is reported as pass@1 (expected single-trial solve rate, averaged over tasks), pass@k, and pass^k (the chance that all k sampled trials of a task pass). See study-matrix below.
Criteria remain diagnostics. Criterion booleans, points, normalized scores, and failure classes explain *why* a task failed: which criteria failed and in which failure class. They do not define the headline number. Through corpus 2.0 every criterion was required, so the score was always pass × 100 and point-based partial credit carried no information beyond the pass bit. That remains true for most 3.0 tasks, and the policy no longer pretends otherwise.
Optional criteria (required = false) exist for checks that are objective but secondary to the task's stated goal, such as package metadata or a downstream-facing annotation. Rules:
- An optional criterion must still be objectively checkable. Subjective style
judgments are out of scope.
- Its
failure_classmust bemaintainabilityorformatting. The loader
rejects an optional criterion in any other class, and every task keeps at least one required functional criterion.
- A task may declare **at most one** optional criterion. The loader enforces
this.
- A failing optional criterion lowers the diagnostic score but never fails the
task, so it never affects the headline pass rate.
- Do not invent criteria to have something optional. Mark an existing
criterion optional only when failing it leaves the task's stated goal met.
No active 3.0 task declares an optional criterion. The policy was first applied to three 2.0 tasks that are now archived under archive/2.0; their rows stay here as worked examples:
| Task | Optional criterion | Why it is secondary |
|---|---|---|
package-stdenv-cli | package-metadata | meta fields (description, homepage, license, platforms, mainProgram) do not affect fetching, building, testing, or installing the CLI, which the other three criteria check. |
package-python-application | package-metadata | Same reasoning: identity, inputs, and the test contract define the package, and meta is packaging hygiene. |
purity-wrapper-derivation | purity-marker | passthru.pure is an annotation for downstream tooling. Purity itself is checked by pure-install-phase and wrapper-inputs. |
All other criteria stay required, because each one is part of what its task measures.
Evaluator contract fixtures follow the same split. A rejecting fixture whose expected-false criteria are all optional expects the named criterion to be false while the task still passes validly. Targeted rejecting coverage is required only for required criteria.
Criteria and score payloads
Current NixBench tasks declare objective binary criteria in metadata.toml. Each criterion has a stable ID, point value, required flag, and failure class. The points must sum exactly to the task's max_score.
[[criteria]]
id = "preserves-runtime-inputs"
points = 25
required = true
failure_class = "wrong-value"
Evaluators write a schema-2 payload to the evaluator-only $NIXBENCH_SCORE_FILE path:
{
"schema_version": 2,
"criteria": {
"evaluates": true,
"preserves-runtime-inputs": false
},
"notes": ["runtime input preservation check failed"]
}
The criterion keys must exactly match task metadata, and each value must be a JSON boolean. Evaluators do not submit totals or failure classes. The harness computes both from metadata. Notes are bounded diagnostics and cannot change a score.
An evaluator exits 0 only when every required criterion passes. Exit 1 rejects the candidate and may retain partial credit. A disagreement between the exit code and required criteria makes the measurement invalid.
Criterion booleans are independent measurement outputs, not a post-hoc split of one binary result. Evaluators guard nested lookups and localize candidate evaluation so an absent field in one dimension does not zero unrelated dimensions. Only a whole-candidate syntax or import failure may legitimately use an all-false rejection fallback.
Evaluator contract schema 3 records the exact expected criterion vector for every fixture. Release health runs each fixture repeatedly, compares the whole actual vector, verifies the named rejecting criterion is false, requires a targeted negative for every required criterion, checks candidate-digest independence, and rejects evaluator error logs or nondeterministic vectors. Historical trials remain bound to their original corpus digest and are never rescored with the hardened boundary.
Measurement validity and task outcome
measurement_status records whether the run measured the candidate:
validmeans the evaluator, score payload, agent process, and any required
completion attestation were usable.
invalidmeans an evaluator, process, transport, score, or attestation
failure prevented measurement.
incompletemeans a study attempt stopped before it ran the selected task
set.
For valid measurements, task_outcome is pass, fail, or agent-timeout. The harness records its own timeout event, so an attested agent timeout is a valid outcome. Evaluator timeouts, evaluator exit codes 2 or greater, malformed score payloads, non-timeout agent process errors, and missing or failed required attestations are invalid measurements. Invalid and incomplete attempts stay in the study attempts ledger but never enter trials, score denominators, or estimates.
Failure classes
Task metadata may use these controlled classes:
syntaxevaluationmissing-attrwrong-valueunavailable-helperimpurityoverfitmaintainabilityformatting
Optional criteria are limited to objective maintainability or formatting checks, at most one per task (see the scoring policy above). Infrastructure events do not use these classes.
Compatibility
The harness can read historical scalar score files for legacy tasks and marks them legacy-binary. Current publishable corpus releases require criteria-v2. export-site accepts legacy scoring only with the explicit --allow-legacy-protocol compatibility flag. Historical binary scores are not converted and must not be pooled with rubric-scored trials.
Study observations
Study schema version 3 retains one normalized observation for every task in every valid trial. The primitive evidence is the task identity and digest, controlled category and difficulty, criterion booleans plus their rubric points, required flags and failure classes, timeout state, measurement status, task outcome, infrastructure events, and agent and evaluator durations. score and max_score are checked against that rubric evidence.
normalized_score is redundant: the shared current-study validator recomputes it as score / max_score (with a documented 1e-12 floating-point tolerance) and requires the stored value to agree and remain in [0, 1]. The same validator derives pass state from required criteria and timeout state, derives passed/failed criterion lists and failure classes, and derives every trial's score, maximum score, score rate, pass/fail counts, task count, timeout count, and timing totals. Current reporting and publication reject an altered normalized_score, impossible score, duplicate task cell, criterion/pass disagreement, or redundant trial total rather than using it.
Historical study summaries are hydrated from their referenced run summaries when those files remain available. A historical summary without those files is marked aggregate_only = true. This is an explicit legacy path: historical aggregate-only data is never upgraded into a current schema-3 publication by filling defaults or inventing task observations. Reports do not infer task observations from a corpus total.
Reported estimands
nixbench.reporting is the canonical implementation for new statistical reports. Reports define three score summaries:
- Macro task score averages each task's normalized score over valid trials,
then averages those task means without task weights.
- Macro pass rate applies the same calculation to each task's binary pass
outcome.
- Point-weighted score divides all earned points by all available points in
the included task-trial cells.
Reports never pool task-trial cells first when calculating a macro result. They publish the task count, valid observation count, trial count, raw points, range, invalid-attempt count, timeout rate where applicable, and exclusion reasons beside the estimates. Category and difficulty groups contain no editorial weights. Groups with fewer than five tasks have descriptive_only = true and no task-resampling interval.
Cross-configuration study matrix
study-matrix summarizes several validated schema-3 studies of one corpus digest. Its JSON output is meant for site/src/data and the documentation:
python3 bench.py study-matrix --studies-dir results/studies --json > matrix.json
python3 bench.py study-matrix --studies-dir results/studies \
--output site/src/data/study-matrix.json
Studies are grouped by configuration_id, and several study files for one configuration are pooled. A trial that appears twice in one configuration is rejected. Studies that span corpus digests are rejected unless --corpus-digest selects one. Pre-schema-3 studies are rejected rather than hydrated.
Per configuration, the output contains:
pass_rate: passing over valid task-trial observations, with a 95% Wilson
interval. Observations of the same task are not independent, so treat the interval as descriptive. Fewer than three trials adds a warning.
pass_at_1,pass_at_k, andpass_hat_kfor every k up to the trial count.
pass@k uses the unbiased estimator 1 - C(n-c,k)/C(n,k), and pass^k uses C(c,k)/C(n,k). Both are averaged over tasks.
- per-task
solve_rate,timeouts,mean_agent_seconds, and the task's
failure set: counts of failed_criteria and failure_classes. Optional criteria appear here even when the task passed.
failed_task_ids(not solved in every trial) andunsolved_task_ids(never
solved).
mean_agent_seconds_per_task,mean_agent_seconds_per_trial,
timeout_count, timeout_rate, and invalid and incomplete attempt counts.
Per task, across configurations, the output contains the pooled solve rate, the solve rate per configuration, the solve-rate band, the empirical difficulty (saturated, easy, mixed, or hard; see benchmark governance), and the point-biserial discrimination. Discrimination comes from the same implementation as corpus health, with the same thresholds by default (20 valid observations across 2 configurations). The --discrimination-min-* flags only lower them for exploratory reports.
Uncertainty methods
The student-t-fixed-corpus-run-variation method estimates the mean statistic from hypothetical independent repetitions of the same fixed corpus under the same correctness configuration. It assumes the repeated-trial statistic is approximately normal. The report retains raw interval bounds. It may also provide bounds clipped to the display range as separate values. One trial has no interval, and fewer than five trials produces a warning.
This interval does not estimate a single future-run prediction interval, evaluator correctness, representativeness of all Nix work, model identity certainty, or a direct significance test between configurations.
The wilson-task-pass-stability method reports per-task pass stability over valid repeated observations within one correctness configuration. Corpus health reports keep these intervals separated by configuration. Their pooled empirical pass rate is descriptive only and does not receive a Wilson interval. Criterion and failure-class reports use raw numerators and denominators rather than intervals for small samples.
Task discrimination uses the point-biserial correlation between the binary task outcome and the sum of the other valid normalized task scores in the same trial. Sample-size and configuration-diversity thresholds are reapplied after rows without a leave-one-task-out score are excluded.
The trial-task-resampling-sensitivity method measures sensitivity to the observed trial and fixed-corpus task composition. It requires a complete rectangular matrix for one corpus and configuration and never imputes missing cells. Each of 10,000 replicates samples trial indices with replacement, then samples task IDs with replacement and computes the macro normalized score. The seed is the SHA-256 digest of the corpus ID, configuration ID, stratum ID, and method version. Bounds use linear interpolation equivalent to Hyndman-Fan type
- This is a sensitivity interval, not evidence of generalization to all Nix
work.
Timing summaries are grouped by timing_environment_id. A correctness configuration with more than one timing environment receives separate timing results and no combined timing interval.