Coding-agent evidence

Model benchmarks, separated by task and methodology.

Compare patches, objective correctness, manual review, solution time, and known provider cost without treating availability or infrastructure failures as code quality.

Current · isolated

Retry-policy patch

A focused JavaScript maintenance task executed through the clean-room OpenCode and Codex pipeline, with public tests, a private objective evaluator, immutable patches, time, and provider-reported cost.

  • 42 runtime routes
  • 4 review criteria
  • public + objective checks
View current results
Patch benchmark

Ledger reconciliation

The first broad patch comparison: 53 coding-agent candidates solving the same ledger task, with availability kept separate from quality scoring.

  • 53 candidates
  • 31 scored patches
  • 5 quality dimensions
View ledger results
Prepared · results pending

Bug fixing

Diagnose and repair duplicate IDs, exact money parsing, sales, refunds, invalid rows, and daily ledger totals.

  • public + objective checks
  • 4 review criteria
  • clean-room execution
View test definition
Prepared · results pending

Feature implementation

Implement feature overrides and deterministic percentage rollout from a detailed behavioral contract.

  • public + objective checks
  • 4 review criteria
  • clean-room execution
View test definition
Prepared · results pending

Refactoring

Extract one shared event-filtering helper while preserving the public API and all existing behavior.

  • behavior + structure
  • 4 review criteria
  • clean-room execution
View test definition
Prepared · results pending

Repository navigation

Explore an unfamiliar module layout, locate the correct formatter, and repair account status labels.

  • public + objective checks
  • 4 review criteria
  • correct-layer change
View test definition
Historical · provisional

Five-task Phase 2

Older multi-task evidence covering bug fixing, feature work, refactoring, repository navigation, and edge-case tests. Audit found shared-history contamination, so this is not presented as a blind benchmark.

  • 31 candidates
  • 5 tasks
  • 155 records
View historical results
Methodology

Runner and source evidence

The public benchmark runner documents clean-room isolation, separate candidate testing and judging, objective evaluation, timing, cost policy, and immutable result artifacts.

  • OpenCode + Codex
  • MIT licensed
  • public artifacts
Open benchmark repository ↗