Test V1 / Phase 1

Model comparison grounded in the patches.

Fifty-three coding-agent candidates received the same ledger-reconciliation task. Quality scores review visible code and are constrained by recorded public and hidden correctness; availability outcomes remain separate from patch quality.

You are viewing Test V1 / Phase 1. Continue to the Test V1 / Phase 2 aggregate. Test V2 begins with Phase 3, which is planned; no Phase 3 results exist yet.

Prices verified: August 16, 2026
53tested candidates
31scored patches
22N/A availability outcomes
5quality dimensions

Complete candidate results

Activate any column heading to sort. Activate it again to reverse direction; N/A scores always remain last.

53 rows

All 53 phase-1 benchmark candidates with quality, speed, access, and price information.

Pricing notes and official sources

OpenCode Free labels describe the tested endpoints and are not a permanent availability or pricing guarantee. DeepSeek V4 Pro and Flash use separate peak/off-peak rates; peak-hour windows and context-length thresholds can change which rate applies, so consult the current rate card before estimating a run.

Cache-write is shown where supplied. A dash means no separate verified cache-write value was provided for this comparison.