Model battery: every shipped model, per role, as permutations
On this page
Store: 144 permutation(s), 10459 record(s), measured 2026-09-10 – 2026-09-13 · cost $154.37 · summary generated 2026-09-13
What this page is
Every model @gadhs/pi ships, measured by the battery one permutation
at a time — a task (what is asked and how it is graded) × a subject (a
model and its knobs: thinking level, thinking budget, output cap,
temperature, sampling) × a stack (which extensions the lane runs under:
bare pi, or the shipped @gadhs/pi distribution) — and scored
mechanically. The numbers below are the battery’s store
(tools/model-battery/store/), one record per call, accumulated across
runs rather than produced whole; the page shows the permutations the
manifest (tools/model-battery/battery.json) declares, and a declared
permutation with no records reads not measured. The recommendation
matrix at the end is derived from the production cells — the bare
model ref under the bare stack, the settings production runs — by the
rules in the same manifest, and every cell prints the numbers that decided
it. Nothing on this page is editorial: change a threshold, re-render, and
the matrix follows.
The shipped defaults are starred (★). They are argued from the same table, and changing one is its own issue, citing this page — never a battery run.
How to read it
-
Prompts are production’s. The reviewer and planning tasks send the shipped review prompts (imported from pi-workflow) with production’s output budget and no reasoning level — which is what the pre-commit review sends. The judge task runs the shipped
runJudgewith the shipped mode policy at its effort. The explore task runs the packagedExplore.mdverbatim at its frontmatter level. The coder task runs pi’s own default system prompt atmedium, what a developer session runs at. A subject’s knobs override; the record carries both what was asked and the level the request was built with (a Gemini asked foroffis sentlow, the floor pi-vertex applies; a Claude whose catalogue marksoffunsupported is sent no thinking field and thinks at its own discretion — the Measured column’s cell says so where it applies), and each record’srequest-knobs.jsonevidence holds the request fields the provider actually built. -
Scores. Reviewer/planning: pooled recall over planted defects (the scorer was calibrated against hand scores;
cases/reviewer/CALIBRATION.md), false flags on the control case, unscored extras. Judge: exact verdicts, UNSAFE (a critical case allowed against its label — the security failure), friction (an acceptedallowanswered otherwise). Explore: expected paths present in what the parent receives (the yield, else the final prose), per corpus — this repository (a small TypeScript monorepo) andruff0.16.6 (a widely used Rust workspace, ~800k lines across 52 crates) — with the matrix reading the worst corpus; yield rate beside it, tool-call failures withyieldcounted apart (a rejected yield is the contract’s retry ladder, reported as retries per yield). The twelve questions per corpus were written blind by a fresh-context model that read the repository and was told only the shape — what a new developer asks, naming no identifier from the repository’s code and stating no mechanism — and never saw a key; the keys (the files a correct answer cites, with alternatives where two are legitimate) were derived from the tree afterwards, no function names are scored, and a test forbids any key identifier in a question’s text. The author’s own difficulty label (grep / multi-hop / design) is reported per tier. Coder: hidden tests passed over total, tasks solved, tool-call failures (bash exits counted apart), whether the run ended by itself. Verify: the packagedVerify.mdasked to run a check in a small project with a known outcome — a passing suite, a failing test, two failing tests, a type error, a syntax error, a command that does not exist — and scored on the object it yields against the case’s key: faithful when the verdict, whether it ran, the quoted evidence and the named failures all match; FALSE PASS counts apassed: trueon a check that failed or never ran, the one failure that makes a verifier worse than none, and the matrix’sfalsePassMax: 0admits none (#91). Integrity: whole answers over completed answers on twelve tiny prompts — a truncated answer is a truncated tool call. -
Spread. Every case runs N times (the manifest’s
runsfor the permutation); a mean is shown with its min-max where the task has one. The production subject adds no temperature: variance is measured, not hidden. A permutation with a temperature knob is an experiment and renders as one. -
Capped is the number of calls that ended on the task’s output budget (
length) rather than on the model’s own stop: production’s judge sends 300 tokens, the review 8k or 16k, and a reasoning model can spend those thinking before it answers. A low score with a high capped count is the budget’s verdict, not the model’s — the matrix marks such cells budget-starved (#81 is the per-model budgets issue). -
Cost is what the wire reported at the catalog’s rates. Wall is the median wall time per call, including tool use for the agent-loop tasks.
-
Measured is the day the cell’s newest record was made, and it is all the page says about a cell’s age. Every record carries the fingerprint of what it was measured against - prompt, case, corpus pin, stack, the catalog entry, pi’s versions - as provenance, and nothing turns that into a to-do: no cell is flagged for its age, and no verb re-measures because the tree moved (#130). A re-measure is a person’s decision - after a prompt change meant to move a cell, after a model changes upstream (not detectable, pi-ai reports no served-model version), when a default is re-argued - made with
battery fill --refresh <model>, abattery run --forceon the slice, or abattery run --store DIRfor a confirmation that leaves the store alone. A prompt that is being revised - a judge rubric, a helper’s agent file - is iterated on its cheap gate first (task evalfor the judge, abattery runon one model for the rest) and the store re-measured once, when the wording has settled: the judge slice is 8,000 calls (#120’s second wording cost a second re-measure). Cost incomplete marks a lane whose spend the battery could not fully see — a helper run it could not price, or a mode-judge call without a price (pi-modes before 0.17 emitted none, #110; from 0.17 the judge’s calls are read priced from the lane’s review log and folded in). -
Experiments sit beside the role they vary: the same task under a knob (a thinking budget, an output cap) or under the shipped stack, and the task’s experiment arms (
explore-ripwire: explore plus a read-onlyripwirecall-graph tool and one paragraph on when to reach for it, #85;explore-ripwire-firstinverts the priority, #88). The matrix ignores them. Under the shipped stack an agent-loop lane is a top-level session, not a subagent: pi-modes' ownyieldtells a parentless session to answer normally, so yielded is structurally false there and the answer is scored as prose — the stack axis measures the bolt-ons' effect on the work, not the subagent contract.
To measure: node tools/task.mjs battery fill fills the manifest’s gaps
(--retry-errors remakes the rate-limited and timed-out records, at
--jobs 1);
battery run --task T --model M [knobs] measures one permutation, listed
or not; then battery render. Every verb that calls a model prints a cost
projection and refuses a run over --budget (default US$300) before any
call, and a knob the model cannot honour is refused before any call
(battery knobs <model>). Details: Local Development.
reviewer
| Model | Recall | Caught / planted | Per-case spread | False flags (control) | Extras | Capped | Errors | Wall | Cost / case | Measured |
|---|---|---|---|---|---|---|---|---|---|---|
|
5 % |
3 / 58 |
0.05 (0.00–0.75) |
0 |
2 |
0 |
0 |
2.4 s |
$0.0006 |
2026-09-12 |
|
38 % |
22 / 58 |
0.38 (0.25–0.75) |
0 |
0 |
0 |
0 |
4.2 s |
$0.0009 |
2026-09-12 |
|
66 % |
38 / 58 |
0.67 (0.50–1.00) |
0 |
6 |
0 |
0 |
29.0 s |
$0.0009 |
2026-09-12 |
|
26 % |
15 / 58 |
0.27 (0.00–0.67) |
0 |
0 |
0 |
0 |
5.3 s |
$0.0010 |
2026-09-12 |
|
50 % |
29 / 58 |
0.51 (0.25–1.00) |
0 |
10 |
0 |
0 |
1.4 s |
$0.0042 |
2026-09-12 |
|
45 % |
26 / 58 |
0.41 (0.00–1.00) |
0 |
10 |
0 |
0 |
23.1 s |
$0.0054 |
2026-09-12 |
|
74 % |
43 / 58 |
0.78 (0.33–1.00) |
0 |
1 |
0 |
0 |
19.1 s |
$0.0062 † |
2026-09-12 (off requested; pi-vertex floors a Gemini request to low (#78)) |
|
74 % |
43 / 58 |
0.78 (0.33–1.00) |
0 |
1 |
0 |
0 |
15.0 s |
$0.0066 † |
2026-09-12 (off requested; pi-vertex floors a Gemini request to low (#78)) |
|
78 % |
45 / 58 |
0.80 (0.50–1.00) |
0 |
6 |
0 |
0 |
26.6 s |
$0.0074 |
2026-09-12 |
|
34 % |
20 / 58 |
0.33 (0.00–1.00) |
0 |
13 |
2 |
1 |
30.6 s |
$0.0088 |
2026-09-12 |
|
81 % |
47 / 58 |
0.82 (0.50–1.00) |
0 |
9 |
0 |
0 |
14.6 s |
$0.01 |
2026-09-12 |
|
43 % |
25 / 58 |
0.39 (0.00–1.00) |
0 |
3 |
0 |
0 |
15.3 s |
$0.02 |
2026-09-12 |
|
76 % |
44 / 58 |
0.77 (0.50–1.00) |
0 |
10 |
0 |
0 |
12.5 s |
$0.02 |
2026-09-12 (off requested; pi-vertex floors a Gemini request to low (#78)) |
|
62 % |
36 / 58 |
0.65 (0.17–1.00) |
0 |
17 |
0 |
0 |
19.4 s |
$0.02 |
2026-09-12 (off requested; pi-vertex floors a Gemini request to low (#78)) |
|
57 % |
33 / 58 |
0.60 (0.25–1.00) |
0 |
2 |
0 |
0 |
16.9 s |
$0.02 |
2026-09-12 (off requested; pi-vertex floors a Gemini request to low (#78)) |
|
88 % |
51 / 58 |
0.89 (0.50–1.00) |
0 |
4 |
0 |
0 |
24.6 s |
$0.03 |
2026-09-12 |
|
88 % |
51 / 58 |
0.89 (0.67–1.00) |
0 |
6 |
0 |
0 |
18.6 s |
$0.06 |
2026-09-12 |
|
97 % |
56 / 58 |
0.96 (0.75–1.00) |
0 |
1 |
0 |
0 |
22.0 s |
$0.06 |
2026-09-12 |
|
100 % |
58 / 58 |
1.00 |
0 |
2 |
0 |
0 |
24.2 s |
$0.12 |
2026-09-12 (off requested; claude-fable-5-1 does not offer it and thinks at its own discretion) |
† introductory rate, through 2026-12-31: vertex-gemini/gemini-3.8-flash, vertex-gemini/gemini-3.7-flash. A cost argument from this table does not survive that date; the catalog’s standard rate is in its //cost-google note.
planning
| Model | Recall | Caught / planted | Per-case spread | False flags (control) | Extras | Capped | Errors | Wall | Cost / case | Measured |
|---|---|---|---|---|---|---|---|---|---|---|
|
57 % |
8 / 14 |
0.57 |
– |
0 |
0 |
0 |
5.3 s |
$0.0007 |
2026-09-12 |
|
71 % |
10 / 14 |
0.71 |
– |
0 |
0 |
0 |
24.1 s |
$0.0008 |
2026-09-12 |
|
64 % |
9 / 14 |
0.64 (0.29–1.00) |
– |
0 |
0 |
0 |
10.8 s |
$0.0014 |
2026-09-12 |
|
36 % |
5 / 14 |
0.36 (0.29–0.43) |
– |
0 |
0 |
0 |
10.6 s |
$0.0015 |
2026-09-12 |
|
29 % |
4 / 14 |
0.29 (0.00–0.57) |
– |
0 |
0 |
0 |
11.3 s |
$0.0017 |
2026-09-12 |
|
43 % |
6 / 14 |
0.43 |
– |
0 |
0 |
0 |
1.0 s |
$0.0040 |
2026-09-12 |
|
71 % |
10 / 14 |
0.71 |
– |
0 |
0 |
0 |
10.9 s |
$0.0045 † |
2026-09-12 (off requested; pi-vertex floors a Gemini request to low (#78)) |
|
79 % |
11 / 14 |
0.79 (0.71–0.86) |
– |
0 |
0 |
0 |
11.1 s |
$0.0045 † |
2026-09-12 (off requested; pi-vertex floors a Gemini request to low (#78)) |
|
64 % |
9 / 14 |
0.64 (0.57–0.71) |
– |
0 |
0 |
0 |
11.2 s |
$0.0053 |
2026-09-12 |
|
71 % |
10 / 14 |
0.71 |
– |
0 |
0 |
0 |
29.6 s |
$0.0057 |
2026-09-12 |
|
57 % |
8 / 14 |
0.57 (0.43–0.71) |
– |
0 |
0 |
0 |
11.4 s |
$0.0099 |
2026-09-12 |
|
79 % |
11 / 14 |
0.79 (0.71–0.86) |
– |
0 |
0 |
0 |
9.6 s |
$0.01 |
2026-09-12 (off requested; pi-vertex floors a Gemini request to low (#78)) |
|
50 % |
7 / 14 |
0.50 (0.43–0.57) |
– |
0 |
0 |
0 |
14.1 s |
$0.02 |
2026-09-12 (off requested; pi-vertex floors a Gemini request to low (#78)) |
|
57 % |
8 / 14 |
0.57 |
– |
0 |
0 |
0 |
20.6 s |
$0.02 |
2026-09-12 (off requested; pi-vertex floors a Gemini request to low (#78)) |
|
79 % |
11 / 14 |
0.79 (0.71–0.86) |
– |
0 |
0 |
0 |
17.4 s |
$0.02 |
2026-09-12 |
|
71 % |
10 / 14 |
0.71 |
– |
0 |
0 |
0 |
24.4 s |
$0.04 |
2026-09-12 |
|
86 % |
12 / 14 |
0.86 |
– |
0 |
0 |
0 |
25.4 s |
$0.06 |
2026-09-12 |
|
79 % |
11 / 14 |
0.79 (0.71–0.86) |
– |
0 |
0 |
0 |
20.5 s |
$0.07 |
2026-09-12 |
|
93 % |
13 / 14 |
0.93 (0.86–1.00) |
– |
0 |
0 |
0 |
24.0 s |
$0.13 |
2026-09-12 (off requested; claude-fable-5-1 does not offer it and thinks at its own discretion) |
† introductory rate, through 2026-12-31: vertex-gemini/gemini-3.7-flash, vertex-gemini/gemini-3.8-flash. A cost argument from this table does not survive that date; the catalog’s standard rate is in its //cost-google note.
judge
| Model | Exact | UNSAFE | Friction | Deferred | Capped | Errors | Wall | Cost / case | Measured |
|---|---|---|---|---|---|---|---|---|---|
|
85 % |
3 |
43 |
41 |
0 |
0 |
0.5 s |
$0.0001 |
2026-09-12 |
|
90 % |
14 |
20 |
35 |
0 |
0 |
0.5 s |
$0.0003 |
2026-09-12 (low requested; grok-4.20-non-reasoning does not reason and is sent no level) |
|
97 % |
0 |
20 |
18 |
0 |
0 |
0.7 s |
$0.0000 |
2026-09-12 (low requested; qwen3-235b does not reason and is sent no level) |
|
96 % |
0 |
26 |
30 |
0 |
0 |
1.3 s |
$0.0029 |
2026-09-12 |
|
97 % |
0 |
21 |
23 |
0 |
0 |
1.5 s |
$0.0001 |
2026-09-12 (low requested; grok-4.1-fast-reasoning does not reason and is sent no level) |
|
97 % |
0 |
24 |
27 |
0 |
0 |
1.6 s |
$0.0025 |
2026-09-12 |
|
87 % |
0 |
48 |
32 |
0 |
0 |
1.6 s |
$0.0001 |
2026-09-12 (low requested; qwen3-coder-480b does not reason and is sent no level) |
|
93 % |
0 |
32 |
32 |
0 |
1 |
1.8 s |
$0.0024 |
2026-09-12 |
|
64 % |
0 |
41 |
139 |
0 |
104 |
1.8 s |
$0.0027 |
2026-09-12 |
|
95 % |
0 |
22 |
24 |
0 |
1 |
2.1 s |
$0.0016 |
2026-09-12 |
|
90 % |
16 |
23 |
31 |
2 |
0 |
2.8 s |
$0.0006 |
2026-09-13 |
|
92 % |
0 |
35 |
27 |
0 |
0 |
4.1 s |
$0.0027 |
2026-09-12 |
|
97 % |
0 |
23 |
29 |
0 |
0 |
4.3 s |
$0.0032 |
2026-09-12 |
|
98 % |
0 |
22 |
30 |
0 |
0 |
4.6 s |
$0.0011 † |
2026-09-12 |
|
97 % |
0 |
24 |
24 |
0 |
0 |
4.7 s |
$0.0003 |
2026-09-12 (low requested; grok-4.20-reasoning does not reason and is sent no level) |
|
96 % |
0 |
25 |
23 |
0 |
0 |
4.9 s |
$0.0032 |
2026-09-12 |
|
90 % |
2 |
39 |
49 |
0 |
0 |
5.1 s |
$0.0013 |
2026-09-12 |
|
73 % |
0 |
94 |
92 |
0 |
0 |
8.1 s |
$0.0065 |
2026-09-12 |
|
94 % |
0 |
32 |
30 |
0 |
0 |
10.1 s |
$0.0013 † |
2026-09-12 |
† introductory rate, through 2026-12-31: vertex-gemini/gemini-3.7-flash, vertex-gemini/gemini-3.8-flash. A cost argument from this table does not survive that date; the catalog’s standard rate is in its //cost-google note.
judge: experiments
Not read by the matrix: the same task under other knobs or another stack, and the task’s experiment arms.
| Permutation | Exact | UNSAFE | Friction | Deferred | Capped | Errors | Wall | Cost / case | Measured |
|---|---|---|---|---|---|---|---|---|---|
|
92 % |
0 |
54 |
51 |
0 |
0 |
1.1 s |
$0.0013 |
2026-09-12 |
|
97 % |
0 |
36 |
45 |
0 |
0 |
3.8 s |
$0.0029 |
2026-09-12 |
|
97 % |
0 |
39 |
48 |
0 |
0 |
4.2 s |
$0.0032 |
2026-09-12 |
|
96 % |
0 |
37 |
48 |
0 |
0 |
4.0 s |
$0.0029 |
2026-09-12 |
|
97 % |
0 |
32 |
43 |
0 |
0 |
3.9 s |
$0.0030 |
2026-09-12 |
|
95 % |
0 |
44 |
41 |
0 |
0 |
1.9 s |
$0.0023 |
2026-09-12 |
|
98 % |
0 |
33 |
40 |
0 |
0 |
1.5 s |
$0.0021 |
2026-09-12 |
explore
| Model | Path recall (worst corpus) | By corpus | By tier | Yield rate | Yield retries / yield | Capped | Tool failures (non-yield) | Ended by itself | Wall | Cost / case | Measured |
|---|---|---|---|---|---|---|---|---|---|---|---|
|
0.00 |
pi 0.00 · ruff 0.00 |
grep 0.00 · multi-hop 0.00 · design 0.00 |
0 % |
– |
0 |
0 % (0/26) |
100 % |
4.2 s |
$0.0005 |
2026-09-11 |
|
0.85 |
pi 1.00 · ruff 0.85 (0.00–1.00) |
grep 1.00 · multi-hop 0.94 · design 0.84 |
52 % |
0.00 |
0 |
3 % (26/777) |
100 % |
25.3 s |
$0.01 |
2026-09-11 (low requested; grok-4.1-fast-reasoning does not reason and is sent no level) |
|
0.81 |
pi 0.82 (0.00–1.00) · ruff 0.81 (0.00–1.00) |
grep 0.84 · multi-hop 0.79 · design 0.81 |
98 % |
0.30 |
0 |
3 % (15/523) |
98 % |
23.2 s |
$0.01 |
2026-09-11 (low requested; qwen3-coder-480b does not reason and is sent no level) |
|
0.00 |
pi 0.00 · ruff 0.06 (0.00–1.00) |
grep 0.06 · multi-hop 0.03 · design 0.00 |
0 % |
– |
17 |
58 % (7/12) |
65 % |
54.9 s |
$0.02 |
2026-09-11 |
|
0.66 |
pi 0.96 (0.00–1.00) · ruff 0.66 (0.00–1.00) |
grep 0.81 · multi-hop 0.84 · design 0.77 |
98 % |
0.17 |
0 |
1 % (10/782) |
100 % |
28.0 s |
$0.02 |
2026-09-13 |
|
0.46 |
pi 0.90 (0.00–1.00) · ruff 0.46 (0.00–1.00) |
grep 0.94 · multi-hop 0.42 · design 0.69 |
38 % |
0.00 |
1 |
2 % (32/1328) |
73 % |
28.7 s |
$0.03 |
2026-09-11 (low requested; qwen3-235b does not reason and is sent no level) |
|
1.00 |
pi 1.00 · ruff 1.00 |
grep 1.00 · multi-hop 1.00 · design 1.00 |
100 % |
0.00 |
0 |
1 % (3/515) |
100 % |
50.5 s |
$0.05 † |
2026-09-11 |
|
1.00 |
pi 1.00 · ruff 1.00 |
grep 1.00 · multi-hop 1.00 · design 1.00 |
100 % |
0.00 |
0 |
0 % (2/458) |
100 % |
54.8 s |
$0.05 † |
2026-09-11 |
|
0.85 |
pi 0.85 (0.00–1.00) · ruff 0.91 (0.00–1.00) |
grep 0.94 · multi-hop 0.81 · design 0.90 |
96 % |
0.00 |
0 |
5 % (14/307) |
100 % |
48.7 s |
$0.07 |
2026-09-11 |
|
0.69 |
pi 0.83 (0.00–1.00) · ruff 0.69 (0.00–1.00) |
grep 0.94 · multi-hop 0.78 · design 0.57 |
54 % |
0.08 |
0 |
2 % (13/696) |
100 % |
27.3 s |
$0.08 |
2026-09-11 |
|
0.90 |
pi 0.90 (0.00–1.00) · ruff 0.99 (0.67–1.00) |
grep 1.00 · multi-hop 0.92 · design 0.92 |
96 % |
0.00 |
0 |
0 % (0/482) |
98 % |
39.5 s |
$0.10 |
2026-09-11 |
|
0.86 |
pi 0.94 (0.33–1.00) · ruff 0.86 (0.33–1.00) |
grep 1.00 · multi-hop 0.82 · design 0.89 |
100 % |
0.00 |
0 |
0 % (1/346) |
100 % |
40.9 s |
$0.10 |
2026-09-11 |
|
0.89 |
pi 1.00 · ruff 0.89 (0.00–1.00) |
grep 0.94 · multi-hop 1.00 · design 0.90 |
92 % |
0.41 |
0 |
3 % (33/981) |
94 % |
52.0 s |
$0.12 |
2026-09-11 |
|
0.73 |
pi 0.73 (0.00–1.00) · ruff 0.84 (0.00–1.00) |
grep 0.97 · multi-hop 0.72 · design 0.67 |
29 % |
0.00 |
0 |
11 % (71/662) |
100 % |
21.8 s |
$0.12 |
2026-09-11 (low requested; grok-4.20-reasoning does not reason and is sent no level) |
|
0.98 |
pi 0.98 (0.50–1.00) · ruff 0.99 (0.67–1.00) |
grep 1.00 · multi-hop 0.97 · design 0.98 |
100 % |
0.00 |
0 |
1 % (4/576) |
100 % |
44.7 s |
$0.13 |
2026-09-11 |
|
0.71 |
pi 0.71 (0.00–1.00) · ruff 0.74 (0.00–1.00) |
grep 0.72 · multi-hop 0.81 · design 0.65 |
2 % |
0.00 |
0 |
6 % (42/699) |
100 % |
17.6 s |
$0.15 |
2026-09-11 (low requested; grok-4.20-non-reasoning does not reason and is sent no level) |
|
1.00 |
pi 1.00 · ruff 1.00 |
grep 1.00 · multi-hop 1.00 · design 1.00 |
96 % |
0.43 |
0 |
1 % (4/371) |
100 % |
37.6 s |
$0.21 |
2026-09-11 |
|
1.00 |
pi 1.00 · ruff 1.00 |
grep 1.00 · multi-hop 1.00 · design 1.00 |
100 % |
0.15 |
0 |
1 % (3/383) |
100 % |
32.2 s |
$0.22 |
2026-09-11 |
|
0.98 |
pi 1.00 · ruff 0.98 (0.50–1.00) |
grep 1.00 · multi-hop 1.00 · design 0.97 |
100 % |
0.04 |
0 |
0 % (2/417) |
100 % |
39.3 s |
$0.32 |
2026-09-11 |
† introductory rate, through 2026-12-31: vertex-gemini/gemini-3.8-flash, vertex-gemini/gemini-3.7-flash. A cost argument from this table does not survive that date; the catalog’s standard rate is in its //cost-google note.
explore: experiments
Not read by the matrix: the same task under other knobs or another stack, and the task’s experiment arms.
| Permutation | Path recall (worst corpus) | By corpus | By tier | Yield rate | Yield retries / yield | Capped | Tool failures (non-yield) | Ended by itself | Wall | Cost / case | Measured |
|---|---|---|---|---|---|---|---|---|---|---|---|
|
1.00 |
pi 1.00 · ruff 1.00 |
grep 1.00 · multi-hop 1.00 · design 1.00 |
0 % |
– |
0 |
0 % (1/592) |
100 % |
66.6 s |
$0.07 † |
2026-09-13 (judge ×4 $0.02) |
† introductory rate, through 2026-12-31: vertex-gemini/gemini-3.8-flash. A cost argument from this table does not survive that date; the catalog’s standard rate is in its //cost-google note.
verify
| Model | Faithful | FALSE PASS | Yield rate | Bash calls | Tool failures | Ended by itself | Errors | Wall | Cost / case | Measured |
|---|---|---|---|---|---|---|---|---|---|---|
|
0 % |
0 |
0 % |
13 |
11 |
100 % |
0 |
2.7 s |
$0.0003 |
2026-09-11 |
|
58 % |
0 |
75 % |
12 |
10 |
100 % |
0 |
5.5 s |
$0.0006 |
2026-09-11 (low requested; grok-4.1-fast-reasoning does not reason and is sent no level) |
|
100 % |
0 |
100 % |
12 |
10 |
100 % |
0 |
3.0 s |
$0.0008 |
2026-09-11 (low requested; qwen3-235b does not reason and is sent no level) |
|
67 % |
0 |
100 % |
12 |
10 |
100 % |
0 |
3.6 s |
$0.0008 |
2026-09-11 (low requested; qwen3-coder-480b does not reason and is sent no level) |
|
92 % |
0 |
100 % |
12 |
11 |
100 % |
0 |
9.2 s |
$0.0018 |
2026-09-13 |
|
92 % |
0 |
100 % |
16 |
12 |
100 % |
0 |
4.9 s |
$0.0020 |
2026-09-11 |
|
100 % |
0 |
100 % |
12 |
10 |
100 % |
0 |
5.7 s |
$0.0037 † |
2026-09-11 |
|
100 % |
0 |
100 % |
12 |
10 |
100 % |
0 |
6.0 s |
$0.0038 † |
2026-09-11 |
|
83 % |
0 |
100 % |
15 |
12 |
100 % |
0 |
2.4 s |
$0.0041 |
2026-09-11 (low requested; grok-4.20-non-reasoning does not reason and is sent no level) |
|
83 % |
0 |
100 % |
14 |
10 |
100 % |
0 |
3.9 s |
$0.0066 |
2026-09-11 (low requested; grok-4.20-reasoning does not reason and is sent no level) |
|
92 % |
0 |
100 % |
12 |
10 |
100 % |
0 |
9.1 s |
$0.0093 |
2026-09-11 |
|
75 % |
0 |
100 % |
14 |
11 |
100 % |
0 |
6.8 s |
$0.0097 |
2026-09-11 |
|
100 % |
0 |
100 % |
14 |
17 |
100 % |
0 |
7.9 s |
$0.01 |
2026-09-11 |
|
75 % |
0 |
100 % |
15 |
13 |
100 % |
0 |
8.5 s |
$0.01 |
2026-09-11 |
|
0 % |
0 |
0 % |
0 |
0 |
75 % |
0 |
28.9 s |
$0.01 |
2026-09-11 |
|
100 % |
0 |
100 % |
12 |
6 |
100 % |
0 |
7.5 s |
$0.02 |
2026-09-11 |
|
100 % |
0 |
100 % |
13 |
2 |
100 % |
0 |
6.6 s |
$0.03 |
2026-09-11 |
|
100 % |
0 |
100 % |
14 |
11 |
100 % |
0 |
11.4 s |
$0.04 |
2026-09-11 |
|
100 % |
0 |
100 % |
15 |
0 |
100 % |
0 |
7.2 s |
$0.05 |
2026-09-11 |
† introductory rate, through 2026-12-31: vertex-gemini/gemini-3.8-flash, vertex-gemini/gemini-3.7-flash. A cost argument from this table does not survive that date; the catalog’s standard rate is in its //cost-google note.
coder
| Model | Hidden tests | Solved | Load failures | Capped | Tool failures (non-bash) | Bash exits ≠ 0 | Ended by itself | Wall | Cost / case | Measured |
|---|---|---|---|---|---|---|---|---|---|---|
|
0.03 (0.00–0.20) |
0 / 6 |
0 |
0 |
0 % (0/2) |
0 / 0 |
100 % |
1.6 s |
$0.0002 |
2026-09-11 |
|
0.52 (0.00–1.00) |
3 / 6 |
4 |
0 |
3 % (2/73) |
13 / 19 |
100 % |
66.4 s |
$0.0045 |
2026-09-11 (medium requested; grok-4.1-fast-reasoning does not reason and is sent no level) |
|
0.83 (0.00–1.00) |
6 / 6 |
1 |
0 |
14 % (17/120) |
15 / 27 |
100 % |
34.7 s |
$0.0098 |
2026-09-12 (medium requested; qwen3-235b does not reason and is sent no level) |
|
0.92 (0.80–1.00) |
7 / 6 |
0 |
0 |
3 % (4/117) |
8 / 45 |
100 % |
36.4 s |
$0.02 |
2026-09-11 (medium requested; qwen3-coder-480b does not reason and is sent no level) |
|
0.75 (0.00–1.00) |
6 / 6 |
0 |
0 |
4 % (4/113) |
10 / 47 |
100 % |
67.1 s |
$0.02 |
2026-09-13 |
|
0.17 (0.00–0.80) |
0 / 6 |
0 |
6 |
33 % (1/3) |
0 / 0 |
50 % |
183.1 s |
$0.03 |
2026-09-12 |
|
0.97 (0.80–1.00) |
10 / 6 |
0 |
0 |
10 % (13/127) |
7 / 43 |
100 % |
61.2 s |
$0.09 |
2026-09-11 |
|
0.88 (0.00–1.00) |
9 / 6 |
1 |
0 |
21 % (45/211) |
18 / 48 |
92 % |
49.7 s |
$0.09 |
2026-09-11 (medium requested; grok-4.20-reasoning does not reason and is sent no level) |
|
1.00 |
12 / 6 |
0 |
0 |
0 % (0/41) |
3 / 32 |
100 % |
52.1 s |
$0.11 |
2026-09-11 |
|
0.98 (0.80–1.00) |
11 / 6 |
0 |
0 |
9 % (11/124) |
29 / 68 |
100 % |
132.0 s |
$0.13 † |
2026-09-11 |
|
0.67 (0.00–1.00) |
7 / 6 |
1 |
0 |
7 % (8/108) |
22 / 41 |
75 % |
61.1 s |
$0.13 |
2026-09-12 |
|
0.97 (0.80–1.00) |
10 / 6 |
0 |
0 |
5 % (2/40) |
0 / 14 |
100 % |
28.4 s |
$0.14 |
2026-09-11 |
|
0.92 (0.00–1.00) |
11 / 6 |
1 |
0 |
7 % (11/153) |
38 / 161 |
100 % |
80.5 s |
$0.16 |
2026-09-12 |
|
1.00 |
12 / 6 |
0 |
0 |
0 % (0/29) |
5 / 28 |
100 % |
32.2 s |
$0.17 |
2026-09-11 |
|
0.97 (0.80–1.00) |
10 / 6 |
0 |
0 |
0 % (0/101) |
17 / 46 |
100 % |
84.3 s |
$0.18 |
2026-09-11 |
|
1.00 |
12 / 6 |
0 |
0 |
2 % (1/56) |
19 / 86 |
100 % |
102.4 s |
$0.23 |
2026-09-11 (medium requested; gemini-3.1-pro-preview does not offer it and pi clamps to high) |
|
1.00 |
12 / 6 |
0 |
0 |
0 % (0/18) |
12 / 28 |
100 % |
31.0 s |
$0.24 |
2026-09-11 |
|
0.85 (0.00–1.00) |
7 / 6 |
0 |
0 |
25 % (121/483) |
32 / 77 |
83 % |
82.1 s |
$0.24 |
2026-09-12 (medium requested; grok-4.20-non-reasoning does not reason and is sent no level) |
|
1.00 |
12 / 6 |
0 |
0 |
6 % (8/140) |
43 / 173 |
100 % |
395.8 s |
$0.33 † |
2026-09-12 |
† introductory rate, through 2026-12-31: vertex-gemini/gemini-3.7-flash, vertex-gemini/gemini-3.8-flash. A cost argument from this table does not survive that date; the catalog’s standard rate is in its //cost-google note.
coder: experiments
Not read by the matrix: the same task under other knobs or another stack, and the task’s experiment arms.
| Permutation | Hidden tests | Solved | Load failures | Capped | Tool failures (non-bash) | Bash exits ≠ 0 | Ended by itself | Wall | Cost / case | Measured |
|---|---|---|---|---|---|---|---|---|---|---|
|
1.00 |
12 / 6 |
0 |
0 |
0 % (0/32) |
8 / 32 |
100 % |
42.0 s |
$0.25 |
2026-09-13 (judge ×4 $0.02) |
|
0.95 (0.80–1.00) |
9 / 6 |
0 |
0 |
9 % (10/116) |
13 / 57 |
100 % |
75.6 s |
$0.10 |
2026-09-13 (judge ×8 $0.04) |
|
1.00 |
12 / 6 |
0 |
0 |
8 % (10/128) |
48 / 153 |
100 % |
362.7 s |
$0.33 † |
2026-09-13 (judge ×129 $0.42) |
† introductory rate, through 2026-12-31: vertex-gemini/gemini-3.8-flash. A cost argument from this table does not survive that date; the catalog’s standard rate is in its //cost-google note.
integrity
| Model | Intact | Whole / completed | Capped | Errors | Wall | Measured |
|---|---|---|---|---|---|---|
|
100 % |
12 / 12 |
0 |
0 |
0.3 s |
2026-09-10 |
|
100 % |
12 / 12 |
0 |
0 |
0.4 s |
2026-09-10 |
|
100 % |
12 / 12 |
0 |
0 |
0.4 s |
2026-09-10 |
|
83 % |
10 / 12 |
0 |
0 |
0.6 s |
2026-09-10 |
|
100 % |
12 / 12 |
0 |
0 |
0.6 s |
2026-09-10 |
|
100 % |
12 / 12 |
0 |
0 |
1.0 s |
2026-09-10 |
|
100 % |
12 / 12 |
0 |
0 |
1.1 s |
2026-09-10 |
|
100 % |
12 / 12 |
0 |
0 |
1.1 s |
2026-09-10 |
|
100 % |
12 / 12 |
0 |
0 |
1.3 s |
2026-09-10 |
|
100 % |
12 / 12 |
0 |
0 |
1.4 s |
2026-09-10 |
|
100 % |
12 / 12 |
0 |
0 |
1.4 s |
2026-09-10 (off requested; claude-fable-5-1 does not offer it and thinks at its own discretion) |
|
100 % |
12 / 12 |
0 |
0 |
1.4 s |
2026-09-12 |
|
100 % |
12 / 12 |
0 |
0 |
1.5 s |
2026-09-10 |
|
100 % |
12 / 12 |
0 |
0 |
3.5 s |
2026-09-10 (off requested; pi-vertex floors a Gemini request to low (#78)) |
|
100 % |
12 / 12 |
0 |
0 |
3.6 s |
2026-09-10 (off requested; pi-vertex floors a Gemini request to low (#78)) |
|
100 % |
12 / 12 |
0 |
0 |
3.8 s |
2026-09-10 (off requested; pi-vertex floors a Gemini request to low (#78)) |
|
100 % |
12 / 12 |
0 |
0 |
4.7 s |
2026-09-10 (off requested; pi-vertex floors a Gemini request to low (#78)) |
|
100 % |
12 / 12 |
0 |
0 |
5.9 s |
2026-09-10 (off requested; pi-vertex floors a Gemini request to low (#78)) |
|
0 % |
0 / 5 |
6 |
1 |
13.5 s |
2026-09-11 |
The matrix
Derived from `tools/model-battery/battery.json’s rules over each role’s production cells; each cell shows the verdict and the numbers that decided it.
| Model | reviewer | planning | judge | explore | verify | coder |
|---|---|---|---|---|---|---|
|
fit |
usable |
fit |
fit |
fit |
fit ★ |
|
usable |
usable |
fit |
fit |
fit |
fit |
|
fit ★ |
fit ★ |
no |
fit |
fit |
fit |
|
usable |
no |
fit ★ |
fit |
fit |
fit |
|
no |
usable |
fit |
usable |
fit |
no |
|
no |
usable |
no |
fit ★ |
fit ★ |
usable |
|
no |
no |
fit |
fit |
fit |
usable |
|
no |
no |
fit |
usable |
usable |
fit |
|
usable |
usable |
fit |
usable |
usable |
fit |
|
no |
no |
no |
usable |
fit |
no |
|
no |
no |
no |
no |
no |
no |
|
usable |
no |
fit |
no |
usable |
no |
|
no |
no |
no |
no |
usable |
no |
|
no |
no |
fit |
usable |
no |
no |
|
usable |
no |
fit |
no |
fit |
usable |
|
no |
no |
no |
no |
no |
fit |
|
no (budget-starved: 2/16 capped) |
no |
no |
no (budget-starved: 17/48 capped) |
no (budget-starved: 3/12 capped) |
no (budget-starved: 6/12 capped) |
|
no |
no |
fit |
no (budget-starved: 1/48 capped) |
fit |
no |
|
no |
no |
no (budget-starved: 2/286 capped) |
no |
fit |
usable |
The shipped defaults, argued from the table
Argued on the retired 2026-09-08 run (its records live in git history under tools/model-battery/results/, removed when the store replaced it - #108); the cells above re-argue each default as the store fills, and a default whose cell reads not measured stands on that run’s numbers until then. reviewer: the pre-commit cold reader (chooseReviewer’s floor and tiers pick at or above the author); planning: the same reader on a plan; judge: the gatekeeper for auto and plan - moved from Haiku 4.5 low to Sonnet 4.6 low on 2026-09-12 (#102) on the store’s bake-off (366 verdicts each): within three of each other on exact (352 vs 349), neither unsafe, the same price per call, 1.5 s median against 4.2 s. On the shipped-stack coder cells the faster judge took Haiku’s lane from 91 s to 69 s and its cost level with bare; Opus 5, whose commands are almost all rule-allowed, made three gated calls in twelve lanes and did not move - its shipped overhead is the larger prompt, not the judge (#100); explore: the Explore helper - moved from Haiku 4.5 to a Gemini Flash on 2026-09-08 (#89) on the blind explore set: Haiku 0.89 worst-corpus recall with a repeatable wrong-subsystem pick on the large repository, 90 % yield, 3 % tool failures; the 3.x Flashes 0.97-1.00 on every tier of both corpora, 100 % yield, under 2 % tool failures, at the same standard price (Flash’s current rate is introductory through 2026-12-31 and was not the argument). 3.8 over 3.7 Flash: level on a fresh n=96 confirmation (0.991 vs 0.997, two partial misses vs one, both 100 % yield, wall 53 s vs 59 s), and the newer model carries the longer support runway; the coder-proxy slowness that first pointed at 3.7 did not appear in explore. Re-measured in the store 2026-09-11 (every packaged model, 48 lanes each, #108): 3.8 and 3.7 Flash 1.00 on both corpora at $0.05 a lane, the Claude flagships 0.98-1.00 at three to six times the price, Haiku 0.89 on the large repository again. The fit rule moved from 0.8 to 0.95 worst-corpus recall on that run (#101): at 0.8 thirteen models ranked fit, Grok 4.1 Fast, Qwen3 Coder and MiniMax among them at 0.81-0.85 - wrong on one question in five or six - and price would have decided among them; at 0.95 six rank fit and the cheapest is the default. verify: the Verify helper - moved from Haiku 4.5 to Gemini 3.8 Flash on 2026-09-12 (#117) on the verify task’s first full fill (run a check with a known outcome, report it faithfully; six cases, two runs, every packaged model; the matrix’s falsePassMax 0 admits no FALSE PASS and none occurred): 3.8 Flash 12/12 faithful with 100 % yield at $0.004 a lane against Haiku’s 12/12 at $0.012, and one model for both helper roles. Haiku had stayed on Verify when Explore moved (#89) because the coder proxy had shown 3.8 Flash weak with bash; the role’s own task does not show it. coder: the auto-mode session model. The plan-mode author (Fable) writes plans, which this battery does not score.
-
reviewer →
vertex-anthropic/claude-fable-5-1: fit — recall 1.00 ≥ 0.9 ✓; integrity 1.00 ≥ 0.98 ✓ -
planning →
vertex-anthropic/claude-fable-5-1: fit — recall 0.93 ≥ 0.9 ✓ -
judge →
vertex-anthropic/claude-sonnet-4-6: fit — unsafe 0.00 = 0 ✓; exact 0.97 ≥ 0.9 ✓; wallMsMedianMax 1563.00 ≤ 8000 ✓ -
explore →
vertex-gemini/gemini-3.8-flash: fit — pathRecallMin 1.00 ≥ 0.95 ✓; toolFailureRateMax 0.01 ≤ 0.05 ✓; integrity 1.00 ≥ 0.98 ✓ -
coder →
vertex-anthropic/claude-opus-5: fit — hiddenPassRate.mean 1.00 ≥ 0.8 ✓; toolFailureRateMax 0.00 ≤ 0.05 ✓; endedBySelf 1.00 ≥ 0.9 ✓; integrity 1.00 ≥ 0.98 ✓ -
verify →
vertex-gemini/gemini-3.8-flash: fit — faithful 1.00 ≥ 0.9 ✓; falsePassMax 0.00 ≤ 0 ✓; yieldRate 1.00 ≥ 0.9 ✓; integrity 1.00 ≥ 0.98 ✓