Skip to content

Results history

Earlier NixBench results

Each corpus generation is scored against its own digest. Its results stay here for reference and are never pooled with, or ranked against, another generation or the current results.

Archived corpus generation

NixBench 3.1.0

Corpus nixbench-public 3.1.0, digest 7add4fb2607fb150…. These results are bound to that digest and are never pooled with another generation.

Study matrix (JSON)Release notes

Tasks
21
Configurations
5
Trials each
3
Run dates
2026-10-04 to 2026-10-04

9 of 21 tasks saturated; see the calibration report.

3.1.0 leaderboard

Configurations ordered by task pass rate. Bars are 95% Wilson intervals; n is shown for every estimate.

Metric
  • Claude Code
  • Codex CLI

Headline: share of valid task-trial observations that pass, with a 95% Wilson interval. Observations of the same task are not independent, so the interval is descriptive.

  1. 1–4

    Claude Sonnet 5.5 via Claude Code

    claude-sonnet-5-5 · default effort · Claude Code

    Task pass rate
    95.2%
    86.9–98.4%
    Passed
    60/63
    task-trials
    Trials
    3
    21 tasks
    Timeouts
    0
    0.0%
    Agent / task
    35 s
    limit 300s
  2. 1–4

    GPT-6 Astra via Codex CLI

    gpt-6-astra · medium effort · Codex CLI

    Task pass rate
    95.2%
    86.9–98.4%
    Passed
    60/63
    task-trials
    Trials
    3
    21 tasks
    Timeouts
    0
    0.0%
    Agent / task
    1m 35s
    limit 300s
  3. 1–4

    Claude Opus 5.5 via Claude Code

    claude-opus-5-5 · default effort · Claude Code

    Task pass rate
    93.7%
    84.8–97.5%
    Passed
    59/63
    task-trials
    Trials
    3
    21 tasks
    Timeouts
    0
    0.0%
    Agent / task
    58 s
    limit 300s
  4. 1–4

    GPT-6.1 Sol via Codex CLI

    gpt-6.1-sol · medium effort · Codex CLI

    Task pass rate
    92.1%
    82.7–96.6%
    Passed
    58/63
    task-trials
    Trials
    3
    21 tasks
    Timeouts
    0
    0.0%
    Agent / task
    1m 51s
    limit 300s
  5. 5

    GPT-6 Luna via Codex CLI

    gpt-6-luna · medium effort · Codex CLI

    Task pass rate
    47.6%
    35.8–59.7%
    Passed
    30/63
    task-trials
    Trials
    3
    21 tasks
    Timeouts
    0
    0.0%
    Agent / task
    73 s
    limit 300s

Rank is a plausible range: a configuration is placed behind another only when their 95% intervals do not overlap, so “1–3” means the data cannot order those configurations. Configurations with fewer than 3 trials are shown but never ranked. Agent time is the mean wall-clock agent seconds per task observation.

Pass rate vs. agent time

Up is more tasks passed; left is less agent time per task. Vertical whiskers are 95% Wilson intervals. Every value is also in the leaderboard above.

  • Claude Code
  • Codex CLI

3.1.0 per-task results

Every task against every configuration. Cells count passing trials; select a task for discrimination, difficulty, criterion failures, and each trial's outcome.

Solve rate across trials
  1. 0
  2. ≤ 1/3
  3. ≤ 2/3
  4. < 1
  5. 1
no valid observation timeout in at least one trial
Solve rate per task and configuration. Each cell shows passing trials over valid trials. Select a task for details.
TaskClaude Sonnet 5.5 via Claude CodeGPT-6 Astra via Codex CLIClaude Opus 5.5 via Claude CodeGPT-6.1 Sol via Codex CLIGPT-6 Luna via Codex CLIPooled
debugging · 1 task
100%15/15
devshells · 2 tasks
80%12/15
100%15/15
modules · 12 tasks
0%0/15
67%10/15
80%12/15
80%12/15
80%12/15
80%12/15
93%14/15
93%14/15
100%15/15
100%15/15
100%15/15
100%15/15
nix-language · 1 task
80%12/15
overlays · 2 tasks
87%13/15
87%13/15
packages · 3 tasks
87%13/15
87%13/15
100%15/15

Columns follow the pass-rate order above. Rows are grouped by category, hardest first. A clock marks a cell with at least one agent timeout; a dashed cell has no valid observation, which is different from zero passes. Hover or focus a cell for trial outcomes; select a task for its full breakdown.

debugging

Find The Recursion Behind A Fleet-Wide Rebuild Failure

debug-infinite-recursion-fleet

This repository holds the NixOS configurations for our three-machine fleet. Since the ops/site-moves merge (see CHANGELOG.md), no host evaluates.

Pooled solve rate
100.0% 15/15
descriptive; no interval
Empirical difficulty
saturated
band: high
Author difficulty
hard
Discrimination
unavailable
insufficient sample size: n=15, 5 configs (needs 20 across 2)

Criterion failures

  • hosts-evaluateno failures
  • site-from-host-configno failures
  • deploy-site-overridesno failures
  • roles-stay-scopedno failures
  • fleet-services-intactno failures

Trials by configuration

Claude Sonnet 5.5 via Claude Code3/3 · 63 s avg

  • 20261004T091313Z-0115ff7f pass62 s
  • 20261004T092822Z-e606619f pass50 s
  • 20261004T094254Z-c4b4110e pass78 s

GPT-6 Astra via Codex CLI3/3 · 2m 05s avg

  • 20261004T091313Z-abbbb990 pass1m 54s
  • 20261004T095010Z-83efac71 pass2m 09s
  • 20261004T102516Z-b8cad2e1 pass2m 11s

Claude Opus 5.5 via Claude Code3/3 · 86 s avg

  • 20261004T091313Z-0f370c7f pass85 s
  • 20261004T093724Z-3a3e7929 pass1m 38s
  • 20261004T100053Z-c7a17c27 pass75 s

GPT-6.1 Sol via Codex CLI3/3 · 2m 55s avg

  • 20261004T091313Z-940d4c58 pass1m 60s
  • 20261004T095218Z-0462b91f pass2m 53s
  • 20261004T103343Z-3ba03744 pass3m 52s

GPT-6 Luna via Codex CLI3/3 · 1m 53s avg

  • 20261004T091313Z-6984ff98 pass1m 38s
  • 20261004T094332Z-e82b44ff pass2m 32s
  • 20261004T101241Z-f6bed25d pass90 s

3.1.0 failure sets

Where configurations agree and disagree. Shared failures point at hard tasks or evaluator issues; solo solves show distinct strengths.

Never solved by any configuration

Solved in every trial by every configuration

Solved by only one configuration

None.

Intermittent: passed in some trials, failed in others

Failure overlapTasks each pair failed in at least one trial: shared / union, with Jaccard similarity. Similar totals can hide different weaknesses.
ConfigurationGPT-6 Luna via Codex CLIClaude Opus 5.5 via Claude CodeGPT-6 Astra via Codex CLIGPT-6.1 Sol via Codex CLIClaude Sonnet 5.5 via Claude Code
GPT-6 Luna via Codex CLI14 failed tasks—2/140.141/140.072/140.141/140.07
Claude Opus 5.5 via Claude Code2 failed tasks2/140.14—1/20.501/30.331/20.50
GPT-6 Astra via Codex CLI1 failed tasks1/140.071/20.50—1/20.501/11.00
GPT-6.1 Sol via Codex CLI2 failed tasks2/140.141/30.331/20.50—1/20.50
Claude Sonnet 5.5 via Claude Code1 failed tasks1/140.071/20.501/11.001/20.50—

Archived corpus generation

NixBench 3.0.0

Corpus nixbench-public 3.0.0, digest 0d4e05063591d17d…. These results are bound to that digest and are never pooled with another generation.

Study matrix (JSON)Release notes

Tasks
30
Configurations
5
Trials each
3
Run dates
2026-10-03 to 2026-10-04

24 of 30 tasks saturated; see the calibration report.

3.0.0 leaderboard

Configurations ordered by task pass rate. Bars are 95% Wilson intervals; n is shown for every estimate.

Metric
  • Claude Code
  • Codex CLI

Headline: share of valid task-trial observations that pass, with a 95% Wilson interval. Observations of the same task are not independent, so the interval is descriptive.

  1. 1–4

    Claude Opus 5.5 via Claude Code

    claude-opus-5-5 · default effort · Claude Code

    Task pass rate
    98.9%
    94.0–99.8%
    Passed
    89/90
    task-trials
    Trials
    3
    30 tasks
    Timeouts
    0
    0.0%
    Agent / task
    25 s
    limit 300s
  2. 1–4

    Claude Sonnet 5.5 via Claude Code

    claude-sonnet-5-5 · default effort · Claude Code

    Task pass rate
    97.8%
    92.3–99.4%
    Passed
    88/90
    task-trials
    Trials
    3
    30 tasks
    Timeouts
    0
    0.0%
    Agent / task
    16 s
    limit 300s
  3. 1–4

    GPT-6 Astra via Codex CLI

    gpt-6-astra · medium effort · Codex CLI

    Task pass rate
    97.8%
    92.3–99.4%
    Passed
    88/90
    task-trials
    Trials
    3
    30 tasks
    Timeouts
    0
    0.0%
    Agent / task
    81 s
    limit 300s
  4. 1–4

    GPT-6.1 Sol via Codex CLI

    gpt-6.1-sol · medium effort · Codex CLI

    Task pass rate
    96.7%
    90.7–98.9%
    Passed
    87/90
    task-trials
    Trials
    3
    30 tasks
    Timeouts
    0
    0.0%
    Agent / task
    77 s
    limit 300s
  5. 5

    GPT-6 Luna via Codex CLI

    gpt-6-luna · medium effort · Codex CLI

    Task pass rate
    82.2%
    73.1–88.8%
    Passed
    74/90
    task-trials
    Trials
    3
    30 tasks
    Timeouts
    0
    0.0%
    Agent / task
    47 s
    limit 300s

Rank is a plausible range: a configuration is placed behind another only when their 95% intervals do not overlap, so “1–3” means the data cannot order those configurations. Configurations with fewer than 3 trials are shown but never ranked. Agent time is the mean wall-clock agent seconds per task observation.

Pass rate vs. agent time

Up is more tasks passed; left is less agent time per task. Vertical whiskers are 95% Wilson intervals. Every value is also in the leaderboard above.

  • Claude Code
  • Codex CLI

3.0.0 per-task results

Every task against every configuration. Cells count passing trials; select a task for discrimination, difficulty, criterion failures, and each trial's outcome.

Solve rate across trials
  1. 0
  2. ≤ 1/3
  3. ≤ 2/3
  4. < 1
  5. 1
no valid observation timeout in at least one trial
Solve rate per task and configuration. Each cell shows passing trials over valid trials. Select a task for details.
TaskClaude Opus 5.5 via Claude CodeClaude Sonnet 5.5 via Claude CodeGPT-6 Astra via Codex CLIGPT-6.1 Sol via Codex CLIGPT-6 Luna via Codex CLIPooled
debugging · 3 tasks
100%15/15
100%15/15
100%15/15
devshells · 3 tasks
60%9/15
87%13/15
93%14/15
fetchers · 1 task
93%14/15
flakes · 3 tasks
100%15/15
100%15/15
100%15/15
modules · 8 tasks
80%12/15
93%14/15
100%15/15
100%15/15
100%15/15
100%15/15
100%15/15
100%15/15
nix-language · 6 tasks
73%11/15
100%15/15
100%15/15
100%15/15
100%15/15
100%15/15
overlays · 3 tasks
87%13/15
100%15/15
100%15/15
packages · 2 tasks
80%12/15
93%14/15
purity · 1 task
100%15/15

Columns follow the pass-rate order above. Rows are grouped by category, hardest first. A clock marks a cell with at least one agent timeout; a dashed cell has no valid observation, which is different from zero passes. Hover or focus a cell for trial outcomes; select a task for its full breakdown.

debugging

Debug A Freeform Module Recursion

debug-freeform-config-cycle

nix eval --json --file demo.nix fails.

Pooled solve rate
100.0% 15/15
descriptive; no interval
Empirical difficulty
saturated
band: high
Author difficulty
hard
Discrimination
unavailable
insufficient sample size: n=15, 5 configs (needs 20 across 2)

Criterion failures

  • status-messageno failures
  • message-overrideno failures
  • freeform-settingsno failures
  • freeform-type-enforcedno failures
  • legacy-optionsno failures

Trials by configuration

Claude Opus 5.5 via Claude Code3/3 · 24 s avg

  • 20261003T224805Z-15f08116 pass28 s
  • 20261003T230035Z-075d7c8a pass24 s
  • 20261003T231337Z-1d65a276 pass21 s

Claude Sonnet 5.5 via Claude Code3/3 · 13 s avg

  • 20261003T224805Z-a04908d5 pass15 s
  • 20261003T225644Z-e20bde37 pass14 s
  • 20261003T230501Z-53dac9d6 pass11 s

GPT-6 Astra via Codex CLI3/3 · 73 s avg

  • 20261003T224805Z-935844c5 pass62 s
  • 20261003T232922Z-0dfaaa8c pass75 s
  • 20261004T001146Z-d8b06e7c pass82 s

GPT-6.1 Sol via Codex CLI3/3 · 60 s avg

  • 20261003T224805Z-27e547ae pass50 s
  • 20261003T232521Z-ac4b1f33 pass65 s
  • 20261004T000530Z-44c4ad23 pass63 s

GPT-6 Luna via Codex CLI3/3 · 47 s avg

  • 20261003T224805Z-e2817d96 pass33 s
  • 20261003T231201Z-7789019b pass48 s
  • 20261003T233632Z-2e8ff65d pass60 s

3.0.0 failure sets

Where configurations agree and disagree. Shared failures point at hard tasks or evaluator issues; solo solves show distinct strengths.

Failure overlapTasks each pair failed in at least one trial: shared / union, with Jaccard similarity. Similar totals can hide different weaknesses.
ConfigurationGPT-6 Astra via Codex CLIClaude Sonnet 5.5 via Claude CodeClaude Opus 5.5 via Claude CodeGPT-6 Luna via Codex CLIGPT-6.1 Sol via Codex CLI
GPT-6 Astra via Codex CLI1 failed tasks—0/30.001/11.000/90.001/11.00
Claude Sonnet 5.5 via Claude Code2 failed tasks0/30.00—0/30.001/90.110/30.00
Claude Opus 5.5 via Claude Code1 failed tasks1/11.000/30.00—0/90.001/11.00
GPT-6 Luna via Codex CLI8 failed tasks0/90.001/90.110/90.00—0/90.00
GPT-6.1 Sol via Codex CLI1 failed tasks1/11.000/30.001/11.000/90.00—

Legacy results · corpus 2.0 and earlier

NixBench 2.0 results

These runs used the 29-task 2.0 corpus and an earlier 26-task corpus. They predate the study-matrix format, so they keep their own charts. They are never pooled with, or ranked against, 3.x results.

2.0 corpus · 20 configurations · 76 runs
all models
8
all configurations
35
all recorded runs
91
task rows
29

Why this corpus was retired: it saturated. In the 2.0 calibration, 27 of 29 tasks were solved by every configuration. A task everyone solves cannot separate agents, and more repeated runs cannot make it discriminate, so 2.0 rankings rested on a handful of tasks.

NixBench 3.0 replaced those tasks with about 30 harder, symptom-driven repairs graded through real Nix behaviour, ran every configuration several times in a clean sandbox, and reported pass rates with intervals. The saturated 2.0 tasks remain in the repository under archive/2.0/ for regression use only.

Benchmark design research 3.0 task selection

2.0 leaderboard

Historical configurations on the retired corpora. Filter by model or evidence strength, then compare tasks solved with normalized time per task.

Corpus
Runs
Task axis
3 models · 14 configurations · 70 runs14 averages shown14/14 repeated

Tasks solved vs. time

Higher is better; farther left is faster. Select a model to inspect effort levels and uncertainty.

Zoomed: 18–29 tasksLinear time↑ more tasks← less time

Best observed average for each visible model

  1. GPT-5.6 Luna23.2/2930.5s/task · medium
  2. GPT-5.6 Sol24.0/2938.5s/task · medium
  3. GPT-5.6 Terra22.8/2938.4s/task · high
Ordered effort path + meanSingle observation or legacy composite; no CISelect a model below to reveal effort labels and uncertainty.
  1. GPT-5.6 Lunavia Codex CLI
    • low23.0/29tasks26.6sper task5 runs95% CI 22.1–23.9 · 1 timeouts
    • medium23.2/29tasks30.5sper task5 runs95% CI 22.2–24.2 · 0 timeouts
    • high22.6/29tasks44.4sper task5 runs95% CI 21.2–24.0 · 0 timeouts
    • xhigh22.4/29tasks53.9sper task5 runs95% CI 21.7–23.1 · 0 timeouts
    • max22.4/29tasks75.0sper task5 runs95% CI 20.7–24.1 · 4 timeouts
  2. GPT-5.6 Solvia Codex CLI
    • low23.6/29tasks27.4sper task5 runs95% CI 22.2–25.0 · 0 timeouts
    • medium24.0/29tasks38.5sper task5 runs95% CI 22.5–25.5 · 0 timeouts
    • high23.6/29tasks47.4sper task5 runs95% CI 22.9–24.3 · 0 timeouts
    • xhigh24.0/29tasks60.7sper task5 runs95% CI 23.1–24.9 · 0 timeouts
    • max23.2/29tasks81.4sper task5 runs95% CI 20.1–26.3 · 3 timeouts
  3. GPT-5.6 Terravia Codex CLI
    • low22.0/29tasks25.4sper task5 runs95% CI 21.1–22.9 · 0 timeouts
    • medium21.8/29tasks28.4sper task5 runs95% CI 20.2–23.4 · 0 timeouts
    • high22.8/29tasks38.4sper task5 runs95% CI 20.8–24.8 · 0 timeouts
    • xhigh22.8/29tasks48.8sper task5 runs95% CI 21.8–23.8 · 0 timeouts
About this data

Corpora are kept separate and runtime is normalized per task. Lines connect configurations from lower to higher effort; they do not imply continuous or monotonic scaling. See the reproducibility method and Astra/Pi run provenance. Raw run IDs appear in individual-run tooltips.

Environments: 1 agent version(s) · 1 host(s) · 1 repository revision(s) · 1 network state(s) · timeout budgets 240s.

Coverage by model family, without picking a lucky winner.

Ranges summarize configuration means. Trial and replication counts show how much evidence sits behind each family.

Methodology
ModelCorpusMean task rangeSeconds / task rangeEffort settingsEvidence
codexGPT-5.526-task corpus19.0–22.0 tasks44.8–94.7slow, medium, high, xhigh4 trials · 4 configurations0/4 replicated · 0 timeouts
codexGPT-5.426-task corpus20.0–22.0 tasks41.0–92.4slow, medium, high, xhigh4 trials · 4 configurations0/4 replicated · 0 timeouts
codexGPT-5.4 mini26-task corpus19.0–21.0 tasks32.0–91.8slow, medium, high, xhigh4 trials · 4 configurations0/4 replicated · 4 timeouts
claudeClaude Opus 4.826-task corpus19.0–21.0 tasks20.8–50.9slow, high, xhigh3 trials · 3 configurations0/3 replicated · 0 timeouts
codexGPT-5.6 Sol29-task corpus23.2–24.0 tasks27.4–81.4shigh, low, max, medium, xhigh25 trials · 5 configurations5/5 replicated · 3 timeouts
codexGPT-5.6 Terra29-task corpus21.8–22.8 tasks25.4–48.8shigh, low, medium, xhigh20 trials · 4 configurations4/4 replicated · 0 timeouts
codexGPT-5.6 Luna29-task corpus22.4–23.2 tasks26.6–75.0shigh, low, max, medium, xhigh25 trials · 5 configurations5/5 replicated · 5 timeouts
piGPT-6 Astra via Pi, no skills29-task corpus22.0–23.0 tasks29.7–36.8shigh, low, medium3 trials · 3 configurations0/3 replicated · 0 timeouts

Every task, shown against fixed baseline runs.

These are named comparison runs—not each model’s best row. Codex columns use xhigh effort; the historical Claude column preserves the original default composite.

Showing 7 of 7 model columns

Comparison contextHistorical 26-task and 2.0 29-task corpora · 240-second per-task timeoutRows marked (+2) combine a 24-task run with two supplemental task artifacts.Inspect run provenance

5.5

GPT-5.5 · 26-task corpus · xhigh

22/264 failed

Average task time: 1m 35s

5.4

GPT-5.4 · 26-task corpus · xhigh

21/265 failed

Average task time: 1m 32s

mini

GPT-5.4 mini · 26-task corpus · xhigh

19/267 failed

Average task time: 1m 32s

opus

Claude Opus 4.8 · 26-task corpus · default

21/265 failed

Average task time: 59s

sol

GPT-5.6 Sol · 29-task corpus · xhigh

21/298 failed

Average task time: 55s

terra

GPT-5.6 Terra · 29-task corpus · xhigh

19/2910 failed

Average task time: 42s

luna

GPT-5.6 Luna · 29-task corpus · xhigh

19/2910 failed

Average task time: 49s

Pass/fail status and elapsed task seconds for the selected model columns.
TaskAreaGPT-5.526-task corpus · xhigh20260624T182835Z-4ad8b555 (+2)GPT-5.426-task corpus · xhigh20260624T190640Z-fa04a19c (+2)GPT-5.4 mini26-task corpus · xhigh20260624T194359Z-268b0abe (+2)Claude Opus 4.826-task corpus · default20260624T202141Z-881ef1e9 (+2)GPT-5.6 Sol29-task corpus · xhigh20260709T175817Z-cb86c575GPT-5.6 Terra29-task corpus · xhigh20260709T175820Z-36a5f2f2GPT-5.6 Luna29-task corpus · xhigh20260709T175826Z-b8e5f041
container-native-vs-ociModulesFail53.0sFail48.3sFail75.9sFail27.8sFail41.5sFail32.8sFail26.5s
debug-infinite-recursionDebuggingPass62.8sPass42.7sPass37.9sPass54.0sPass40.9sPass43.5sPass47.4s
debug-network-false-leadDebuggingFail212.9sFail193.1sFail240.0sFail182.5sFail105.0sFail82.8sFail130.7s
devshell-tooling-contractDev shellsPass46.9sPass131.1sPass53.2sPass39.0sPass60.5sPass56.8sPass51.0s
fetcher-source-pinFetchersPass96.8sPass138.4sPass107.0sPass36.9sPass53.8sPass26.1sPass23.0s
fhs-binary-wrapperPackagingPass98.0sPass68.5sPass142.7sPass47.1sPass51.1sPass47.6sPass43.7s
flake-input-package-selectionFlakesPass29.3sPass63.9sPass24.4sPass21.7sPass31.3sPass36.8sPass32.1s
flake-per-system-outputsFlakesPass212.4sPass184.7sFail240.0sPass74.9sPass74.1sPass53.9sPass87.6s
home-manager-extra-special-argsFlakesPass135.5sPass138.4sPass80.5sPass131.9sPass43.2sPass25.8sPass21.9s
home-manager-wsl-module-importModulesPass39.3sPass78.0sPass44.3sPass36.5sPass39.7sPass36.6sPass31.6s
home-manager-xdg-filesModulesPass45.7sPass49.1sPass56.9sPass33.1sPass34.3sPass25.9sPass28.1s
issue-report-qualityDebuggingFail38.8sFail50.2sFail39.1sFail89.3sFail49.1sFail36.8sFail28.7s
lang-attrsets-normalizeNix languagePass87.3sPass80.0sPass74.6sPass67.7sPass80.5sPass56.1sPass68.4s
module-path-compositionNix languagePass98.1sPass51.7sPass49.6sPass38.2sPass47.1sPass36.1sPass43.8s
module-service-optionsModulesFail124.8sPass106.6sPass84.6sPass81.2sFail64.1sFail41.0sFail67.8s
module-stale-option-migrationModulesPass41.5sPass37.7sPass63.1sPass29.9sPass38.3sPass37.3sPass19.7s
module-system-boundariesModulesPass55.6sPass81.8sPass40.8sPass45.3sPass41.1sPass40.6sPass29.8s
mutable-config-home-managerModulesPass82.8sPass42.8sPass63.3sPass38.2sPass44.2sPass41.2sFail46.7s
nushell-command-not-foundModulesNo data--No data--No data--No data--Fail50.4sFail60.0sFail44.1s
overlay-module-boundaryOverlaysPass69.1sPass58.4sPass95.8sPass34.2sPass58.4sPass29.0sPass34.2s
overlay-override-packageOverlaysPass44.0sPass70.9sPass50.5sPass48.6sPass53.9sFail32.8sPass61.2s
package-name-lookup-contractPackagingPass51.2sPass54.0sPass26.6sPass36.4sPass39.0sPass41.3sPass53.9s
package-python-applicationPackagingPass133.5sFail185.6sFail240.0sPass51.5sPass69.3sFail46.0sFail42.1s
package-stdenv-cliPackagingPass182.2sFail203.1sFail184.5sFail30.3sFail72.6sFail63.3sPass67.4s
purity-wrapper-derivationPurityPass145.0sPass107.3sPass79.1sPass100.4sPass39.0sPass38.0sFail52.5s
python-cuda-uv2nix-patchPackagingPass44.4sPass52.8sPass68.9sPass35.4sPass51.8sPass34.8sPass32.1s
rust-no-network-buildPackagingNo data--No data--No data--No data--Fail56.5sFail37.7sFail64.1s
string-escaping-systemdNix languagePass231.9sPass83.1sFail122.8sFail111.6sPass70.3sPass43.2sPass111.0s
xdg-portal-mergeModulesNo data--No data--No data--No data--Fail90.8sFail38.3sFail36.1s

Elapsed task time for the same fixed baseline runs.

Timing follows the exact columns above: xhigh Codex runs and the historical Claude default composite, each with a 240-second per-task timeout.

Charting 7 of 7 model columns

Run provenance (7)
  • GPT-5.526-task corpus · xhigh20260624T182835Z-4ad8b555 (+2)
  • GPT-5.426-task corpus · xhigh20260624T190640Z-fa04a19c (+2)
  • GPT-5.4 mini26-task corpus · xhigh20260624T194359Z-268b0abe (+2)
  • Claude Opus 4.826-task corpus · default20260624T202141Z-881ef1e9 (+2)
  • GPT-5.6 Sol29-task corpus · xhigh20260709T175817Z-cb86c575
  • GPT-5.6 Terra29-task corpus · xhigh20260709T175820Z-36a5f2f2
  • GPT-5.6 Luna29-task corpus · xhigh20260709T175826Z-b8e5f041

Several outcome patterns repeat across runs.

Native-container outcome

The xhigh/default baseline rows did not pass the native NixOS container task; GPT-5.5 low passed it in the effort sweep.

Debugging and reports

The false-lead debugging task and issue-report task were not passed by any row in this set.

Packaging variation

The Python application and stdenv CLI tasks show different pass patterns across models and effort levels.

String escaping

GPT-5.5 and GPT-5.4 passed the baseline string-escaping task; GPT-5.4 mini and Claude Opus 4.8 did not.