Model comparison grounded in the patches.
Fifty-three coding-agent candidates received the same ledger-reconciliation task. Quality scores review visible code and are constrained by recorded public and hidden correctness; availability outcomes remain separate from patch quality.
You are viewing Test V1 / Phase 1. Continue to the Test V1 / Phase 2 aggregate. Test V2 begins with Phase 3, which is planned; no Phase 3 results exist yet.
Complete candidate results
Activate any column heading to sort. Activate it again to reverse direction; N/A scores always remain last.
53 rows
Pricing notes and official sources
OpenCode Free labels describe the tested endpoints and are not a permanent availability or pricing guarantee. DeepSeek V4 Pro and Flash use separate peak/off-peak rates; peak-hour windows and context-length thresholds can change which rate applies, so consult the current rate card before estimating a run.
Cache-write is shown where supplied. A dash means no separate verified cache-write value was provided for this comparison.