Simplicity
How direct and economical the solution is. Unnecessary branches, abstractions, or complicated control flow lower the score.
Fifty-nine coding-agent candidates received the same ledger-reconciliation task. Quality scores review visible code and are constrained by recorded public and hidden correctness; availability outcomes remain separate from patch quality.
You are viewing the first patch benchmark. Continue to the current isolated retry-policy patch results, or browse the benchmark index.
Test assignment
The model received a small, deliberately faulty settlement module. Its job was to diagnose the defects and make the smallest correct patch, without changing the tests, package metadata, benchmark files, or public API.
$ and comma separators, but reject negatives,
malformed values, and values with more than two decimal places.
Each quality column receives a score from 0 to 10. Higher scores indicate stronger performance; N/A means that no valid patch was available for review.
Review policy. GPT-5.6 Sol was the single canonical expert reviewer to keep benchmark costs controlled. Independent Gemini cross-checks broadly agreed with its scoring. Claude Opus 4.8 was removed from future judging after repeatedly missing seeded hidden defects while evaluating candidate patches.
Independent verification. The test results, submitted patches, evaluation artifacts, and testing pipeline are publicly available. You can inspect the evidence, reproduce the process, and make your own independent assessment.
How direct and economical the solution is. Unnecessary branches, abstractions, or complicated control flow lower the score.
How clearly the code communicates its intent through structure, naming, and idiomatic JavaScript.
Whether the patch stays focused on the required repair without unrelated helpers, rewrites, or changes outside the task.
How consistently the implementation satisfies the documented ledger contract and avoids brittle behavior.
How correctly the patch handles malformed money, duplicates, orphan refunds, UTC dates, and other boundary inputs.
Overall is the arithmetic mean of the five quality scores. Rank follows Overall, while unavailable outcomes remain separate from patch quality.
Activate any column heading to sort. Activate it again to reverse direction; N/A scores always remain last.
72 rows
OpenCode Free labels describe the tested endpoints and are not a permanent availability or pricing guarantee. DeepSeek V4 Pro and Flash use separate peak/off-peak rates; peak-hour windows and context-length thresholds can change which rate applies, so consult the current rate card before estimating a run.
Cache-write is shown where supplied. A dash means no separate verified cache-write value was provided for this comparison.
Candidate patches, run records, deterministic evaluation summaries, and expert reviews for the latest completion wave are public in the models-test repository.