Historical multi-task benchmark · Provisional

Historical five-task benchmark: aggregate results

An aggregate of five different coding tasks: repairing ledger logic, implementing feature flags, refactoring event processing, navigating a repository to fix status labels, and strengthening range parsing and edge-case tests. Each model has one record per task; the aggregate combines objective results and four-dimensional expert reviews across all five.

This is provisional historical evidence, not part of the current isolated ranking. See the retry-policy patch results or the benchmark index.

Historical / Provisional — not a blind benchmark
31candidates
155task records
146 / 134public / hidden passes
0forbidden changes

Complete candidate results

Activate any column heading to sort. Activate it again to reverse direction.

31 rows

All 31 Phase 2 candidates, with aggregate arithmetic means and status summaries across five records.
1 opencode-go/deepseek-v4-pro04-deepseek-v4-pro 5/5 5/5 5 5 10.0 10.0 9.0 10.0 9.75 65.574 5/5 completed 5/5 5/5 records ↗
2 opencode-go/glm-5.205-glm-5-2 5/5 5/5 5 5 10.0 10.0 9.0 10.0 9.75 119.742 5/5 completed 5/5 5/5 records ↗
3 opencode-go/gpt-5.6-luna01-gpt-5-6-luna 5/5 5/5 5 5 10.0 10.0 8.8 10.0 9.70 38.673 5/5 completed 5/5 5/5 records ↗
4 opencode/deepseek-v4-flash-free22-deepseek-v4-flash-free 5/5 5/5 5 5 10.0 10.0 8.8 10.0 9.70 51.981 5/5 completed 5/5 5/5 records ↗
5 opencode/claude-sonnet-552-claude-sonnet-5 5/5 5/5 5 5 10.0 10.0 8.8 10.0 9.70 68.213 5/5 completed 5/5 5/5 records ↗
6 opencode-go/deepseek-v4-flash11-deepseek-v4-flash 5/5 5/5 5 5 10.0 10.0 8.8 10.0 9.70 74.143 5/5 completed 5/5 5/5 records ↗
7 gpt-5.6-terra48-gpt-5-6-terra 5/5 5/5 5 5 10.0 10.0 8.8 10.0 9.70 87.055 5/5 completed 5/5 5/5 records ↗
8 opencode-go/kimi-k2.616-kimi-k2-6 5/5 5/5 5 5 10.0 10.0 8.8 10.0 9.70 91.592 5/5 completed 5/5 5/5 records ↗
9 opencode-go/kimi-k2.7-code03-kimi-k2-7-code 5/5 5/5 5 5 10.0 10.0 8.8 10.0 9.70 99.379 5/5 completed 5/5 5/5 records ↗
10 opencode-go/minimax-m306-minimax-m3 5/5 5/5 5 5 10.0 10.0 8.8 10.0 9.70 101.459 5/5 completed 5/5 5/5 records ↗
11 gpt-5.449-gpt-5-4 5/5 5/5 5 5 10.0 10.0 8.8 10.0 9.70 119.703 5/5 completed 5/5 5/5 records ↗
12 opencode-go/glm-5.115-glm-5-1 5/5 5/5 5 5 10.0 10.0 8.8 10.0 9.70 135.099 5/5 completed 5/5 5/5 records ↗
13 opencode-go/grok-4.509-grok-4-5 5/5 5/5 5 5 10.0 10.0 8.6 10.0 9.65 29.899 5/5 completed 5/5 5/5 records ↗
14 opencode-go/mimo-v2.5-pro07-mimo-v2-5-pro 5/5 5/5 5 5 10.0 10.0 8.6 10.0 9.65 105.134 5/5 completed 5/5 5/5 records ↗
15 gpt-5.6-sol47-gpt-5-6-sol 5/5 5/5 5 5 10.0 10.0 8.6 10.0 9.65 121.107 5/5 completed 5/5 5/5 records ↗
16 gpt-5.4-mini50-gpt-5-4-mini 5/5 5/5 5 5 10.0 10.0 8.6 10.0 9.65 124.081 5/5 completed 5/5 5/5 records ↗
17 gpt-5.6-luna46-gpt-5-6-luna 5/5 5/5 5 5 10.0 10.0 8.6 10.0 9.65 195.330 5/5 completed 5/5 5/5 records ↗
18 opencode-go/kimi-k310-kimi-k3 5/5 5/5 5 5 10.0 10.0 8.4 10.0 9.60 189.396 5/5 completed 5/5 5/5 records ↗
19 gpt-5.3-codex-spark18-gpt-5-3-codex-spark 5/5 5/5 5 5 10.0 10.0 8.2 10.0 9.55 60.725 5/5 completed 5/5 5/5 records ↗
20 opencode/claude-opus-553-claude-opus-5 5/5 5/5 5 5 10.0 10.0 8.2 10.0 9.55 143.420 5/5 completed 5/5 5/5 records ↗
21 opencode-go/mimo-v2.513-mimo-v2-5 5/5 4/5 5 4 9.4 9.0 8.8 10.0 9.30 72.584 5/5 completed 5/5 4/5 records ↗
22 opencode/hy3-free27-hy3-free 5/5 4/5 5 4 9.4 9.0 8.6 10.0 9.25 59.818 5/5 completed 5/5 4/5 records ↗
23 opencode/big-pickle19-big-pickle 5/5 4/5 5 4 9.0 8.6 9.0 10.0 9.15 68.279 5/5 completed 5/5 4/5 records ↗
24 opencode-go/hy314-hy3 5/5 4/5 5 4 9.0 8.6 8.8 10.0 9.10 42.575 5/5 completed 5/5 4/5 records ↗
25 opencode/claude-haiku-4-551-claude-haiku-4-5 5/5 4/5 5 4 9.0 8.6 8.2 10.0 8.95 55.642 5/5 completed 5/5 4/5 records ↗
26 opencode-go/minimax-m2.717-minimax-m2-7 5/5 3/5 5 3 8.6 7.8 8.8 10.0 8.80 88.774 5/5 completed 5/5 3/5 records ↗
27 opencode/nemotron-3.5-lightning-free33-nemotron-3-5-lightning-free 5/5 3/5 5 3 8.4 7.6 9.0 10.0 8.75 55.641 5/5 completed 5/5 3/5 records ↗
N/A opencode/nemotron-3-ultra-free35-nemotron-3-ultra-free 4/5 3/5 4 3 8.8 8.3 9.0 10.0 9.00 81.606 5/5 completed 4/5 3/5 records ↗
N/A opencode/mimo-v2.5-free28-mimo-v2-5-free 4/5 2/5 4 2 8.0 7.0 8.8 10.0 8.44 168.850 5/5 completed 5/5 2/5 records ↗
N/A opencode/laguna-s-2.1-free20-laguna-s-2-1-free 3/5 3/5 3 3 10.0 10.0 8.7 10.0 9.67 499.955 5/5 completed 3/5 3/5 records ↗
N/A opencode-go/qwen3.7-plus02-qwen3-7-plus 0/5 0/5 0 0 N/A N/A N/A N/A N/A 900.024 5/5 completed 0/5 0/5 records ↗

Source data and evidence

Generated from the Phase 2 records and expert reviews at commit f6e0d4ed4df2690f66b66d656012238d9cfe5add. Times are agent wall-clock solution durations. Public and hidden results remain separate objective evidence; expert scores use four 0–10 criteria and Overall is their arithmetic mean. No-patch outcomes remain N/A and do not count as code-quality failures.