Ledger reconciliation patch

Patch test: repair transaction accounting

Fifty-nine coding-agent candidates received the same ledger-reconciliation task. Quality scores review visible code and are constrained by recorded public and hidden correctness; availability outcomes remain separate from patch quality.

You are viewing the first patch benchmark. Continue to the current isolated retry-policy patch results, or browse the benchmark index.

Prices verified: August 16, 2026
72provider routes
52scored patches
20N/A availability outcomes
5quality dimensions

Test assignment

Ledger reconciliation: repair transaction accounting

The model received a small, deliberately faulty settlement module. Its job was to diagnose the defects and make the smallest correct patch, without changing the tests, package metadata, benchmark files, or public API.

What the repaired code had to do

  • Read sales and refunds and calculate net totals for each account, currency, and UTC calendar day.
  • Parse money exactly: accept normal dollar values with optional $ and comma separators, but reject negatives, malformed values, and values with more than two decimal places.
  • Treat transaction IDs as idempotent: process the first occurrence and count every later occurrence as a duplicate.
  • Apply a refund to the original sale's account and currency on the refund date; ignore and count refunds whose original sale cannot be found.
  • Count as processed only rows that actually change a total, while preserving the existing exported functions.
Allowed file: fixtures/phase1-ledger/src/reconcile.js

How each reviewed patch is scored

Each quality column receives a score from 0 to 10. Higher scores indicate stronger performance; N/A means that no valid patch was available for review.

Review policy. GPT-5.6 Sol was the single canonical expert reviewer to keep benchmark costs controlled. Independent Gemini cross-checks broadly agreed with its scoring. Claude Opus 4.8 was removed from future judging after repeatedly missing seeded hidden defects while evaluating candidate patches.

Independent verification. The test results, submitted patches, evaluation artifacts, and testing pipeline are publicly available. You can inspect the evidence, reproduce the process, and make your own independent assessment.

Simplicity

How direct and economical the solution is. Unnecessary branches, abstractions, or complicated control flow lower the score.

Readability

How clearly the code communicates its intent through structure, naming, and idiomatic JavaScript.

No extra code

Whether the patch stays focused on the required repair without unrelated helpers, rewrites, or changes outside the task.

Reliability

How consistently the implementation satisfies the documented ledger contract and avoids brittle behavior.

Edge cases

How correctly the patch handles malformed money, duplicates, orphan refunds, UTC dates, and other boundary inputs.

Overall is the arithmetic mean of the five quality scores. Rank follows Overall, while unavailable outcomes remain separate from patch quality.

Complete candidate results

Activate any column heading to sort. Activate it again to reverse direction; N/A scores always remain last.

72 rows

All 72 phase-1 provider routes with quality, speed, access, and price information. This includes 59 original routes and 13 later OpenRouter reruns of previously unavailable models.

Pricing notes and official sources

OpenCode Free labels describe the tested endpoints and are not a permanent availability or pricing guarantee. DeepSeek V4 Pro and Flash use separate peak/off-peak rates; peak-hour windows and context-length thresholds can change which rate applies, so consult the current rate card before estimating a run.

Cache-write is shown where supplied. A dash means no separate verified cache-write value was provided for this comparison.

Candidate patches, run records, deterministic evaluation summaries, and expert reviews for the latest completion wave are public in the models-test repository.