Model battery: every shipped model, per role, as permutations

On this page
NOTE

Store: 144 permutation(s), 10459 record(s), measured 2026-09-10 – 2026-09-13 · cost $154.37 · summary generated 2026-09-13

What this page is

Every model @gadhs/pi ships, measured by the battery one permutation at a time — a task (what is asked and how it is graded) × a subject (a model and its knobs: thinking level, thinking budget, output cap, temperature, sampling) × a stack (which extensions the lane runs under: bare pi, or the shipped @gadhs/pi distribution) — and scored mechanically. The numbers below are the battery’s store (tools/model-battery/store/), one record per call, accumulated across runs rather than produced whole; the page shows the permutations the manifest (tools/model-battery/battery.json) declares, and a declared permutation with no records reads not measured. The recommendation matrix at the end is derived from the production cells — the bare model ref under the bare stack, the settings production runs — by the rules in the same manifest, and every cell prints the numbers that decided it. Nothing on this page is editorial: change a threshold, re-render, and the matrix follows.

The shipped defaults are starred (★). They are argued from the same table, and changing one is its own issue, citing this page — never a battery run.

How to read it

  • Prompts are production’s. The reviewer and planning tasks send the shipped review prompts (imported from pi-workflow) with production’s output budget and no reasoning level — which is what the pre-commit review sends. The judge task runs the shipped runJudge with the shipped mode policy at its effort. The explore task runs the packaged Explore.md verbatim at its frontmatter level. The coder task runs pi’s own default system prompt at medium, what a developer session runs at. A subject’s knobs override; the record carries both what was asked and the level the request was built with (a Gemini asked for off is sent low, the floor pi-vertex applies; a Claude whose catalogue marks off unsupported is sent no thinking field and thinks at its own discretion — the Measured column’s cell says so where it applies), and each record’s request-knobs.json evidence holds the request fields the provider actually built.

  • Scores. Reviewer/planning: pooled recall over planted defects (the scorer was calibrated against hand scores; cases/reviewer/CALIBRATION.md), false flags on the control case, unscored extras. Judge: exact verdicts, UNSAFE (a critical case allowed against its label — the security failure), friction (an accepted allow answered otherwise). Explore: expected paths present in what the parent receives (the yield, else the final prose), per corpus — this repository (a small TypeScript monorepo) and ruff 0.16.6 (a widely used Rust workspace, ~800k lines across 52 crates) — with the matrix reading the worst corpus; yield rate beside it, tool-call failures with yield counted apart (a rejected yield is the contract’s retry ladder, reported as retries per yield). The twelve questions per corpus were written blind by a fresh-context model that read the repository and was told only the shape — what a new developer asks, naming no identifier from the repository’s code and stating no mechanism — and never saw a key; the keys (the files a correct answer cites, with alternatives where two are legitimate) were derived from the tree afterwards, no function names are scored, and a test forbids any key identifier in a question’s text. The author’s own difficulty label (grep / multi-hop / design) is reported per tier. Coder: hidden tests passed over total, tasks solved, tool-call failures (bash exits counted apart), whether the run ended by itself. Verify: the packaged Verify.md asked to run a check in a small project with a known outcome — a passing suite, a failing test, two failing tests, a type error, a syntax error, a command that does not exist — and scored on the object it yields against the case’s key: faithful when the verdict, whether it ran, the quoted evidence and the named failures all match; FALSE PASS counts a passed: true on a check that failed or never ran, the one failure that makes a verifier worse than none, and the matrix’s falsePassMax: 0 admits none (#91). Integrity: whole answers over completed answers on twelve tiny prompts — a truncated answer is a truncated tool call.

  • Spread. Every case runs N times (the manifest’s runs for the permutation); a mean is shown with its min-max where the task has one. The production subject adds no temperature: variance is measured, not hidden. A permutation with a temperature knob is an experiment and renders as one.

  • Capped is the number of calls that ended on the task’s output budget (length) rather than on the model’s own stop: production’s judge sends 300 tokens, the review 8k or 16k, and a reasoning model can spend those thinking before it answers. A low score with a high capped count is the budget’s verdict, not the model’s — the matrix marks such cells budget-starved (#81 is the per-model budgets issue).

  • Cost is what the wire reported at the catalog’s rates. Wall is the median wall time per call, including tool use for the agent-loop tasks.

  • Measured is the day the cell’s newest record was made, and it is all the page says about a cell’s age. Every record carries the fingerprint of what it was measured against - prompt, case, corpus pin, stack, the catalog entry, pi’s versions - as provenance, and nothing turns that into a to-do: no cell is flagged for its age, and no verb re-measures because the tree moved (#130). A re-measure is a person’s decision - after a prompt change meant to move a cell, after a model changes upstream (not detectable, pi-ai reports no served-model version), when a default is re-argued - made with battery fill --refresh <model>, a battery run --force on the slice, or a battery run --store DIR for a confirmation that leaves the store alone. A prompt that is being revised - a judge rubric, a helper’s agent file - is iterated on its cheap gate first (task eval for the judge, a battery run on one model for the rest) and the store re-measured once, when the wording has settled: the judge slice is 8,000 calls (#120’s second wording cost a second re-measure). Cost incomplete marks a lane whose spend the battery could not fully see — a helper run it could not price, or a mode-judge call without a price (pi-modes before 0.17 emitted none, #110; from 0.17 the judge’s calls are read priced from the lane’s review log and folded in).

  • Experiments sit beside the role they vary: the same task under a knob (a thinking budget, an output cap) or under the shipped stack, and the task’s experiment arms (explore-ripwire: explore plus a read-only ripwire call-graph tool and one paragraph on when to reach for it, #85; explore-ripwire-first inverts the priority, #88). The matrix ignores them. Under the shipped stack an agent-loop lane is a top-level session, not a subagent: pi-modes' own yield tells a parentless session to answer normally, so yielded is structurally false there and the answer is scored as prose — the stack axis measures the bolt-ons' effect on the work, not the subagent contract.

To measure: node tools/task.mjs battery fill fills the manifest’s gaps (--retry-errors remakes the rate-limited and timed-out records, at --jobs 1); battery run --task T --model M [knobs] measures one permutation, listed or not; then battery render. Every verb that calls a model prints a cost projection and refuses a run over --budget (default US$300) before any call, and a knob the model cannot honour is refused before any call (battery knobs <model>). Details: Local Development.

reviewer

Model Recall Caught / planted Per-case spread False flags (control) Extras Capped Errors Wall Cost / case Measured

vertex-maas/gpt-oss-120b

5 %

3 / 58

0.05 (0.00–0.75)

0

2

0

0

2.4 s

$0.0006

2026-09-12

vertex-maas/qwen3-235b

38 %

22 / 58

0.38 (0.25–0.75)

0

0

0

0

4.2 s

$0.0009

2026-09-12

vertex-maas/grok-4.1-fast-reasoning

66 %

38 / 58

0.67 (0.50–1.00)

0

6

0

0

29.0 s

$0.0009

2026-09-12

vertex-maas/qwen3-coder-480b

26 %

15 / 58

0.27 (0.00–0.67)

0

0

0

0

5.3 s

$0.0010

2026-09-12

vertex-maas/grok-4.20-non-reasoning

50 %

29 / 58

0.51 (0.25–1.00)

0

10

0

0

1.4 s

$0.0042

2026-09-12

vertex-maas/minimax-m2

45 %

26 / 58

0.41 (0.00–1.00)

0

10

0

0

23.1 s

$0.0054

2026-09-12

vertex-gemini/gemini-3.8-flash

74 %

43 / 58

0.78 (0.33–1.00)

0

1

0

0

19.1 s

$0.0062 †

2026-09-12 (off requested; pi-vertex floors a Gemini request to low (#78))

vertex-gemini/gemini-3.7-flash

74 %

43 / 58

0.78 (0.33–1.00)

0

1

0

0

15.0 s

$0.0066 †

2026-09-12 (off requested; pi-vertex floors a Gemini request to low (#78))

vertex-maas/grok-4.20-reasoning

78 %

45 / 58

0.80 (0.50–1.00)

0

6

0

0

26.6 s

$0.0074

2026-09-12

vertex-maas/qwen3-next-80b-thinking

34 %

20 / 58

0.33 (0.00–1.00)

0

13

2

1

30.6 s

$0.0088

2026-09-12

vertex-maas/kimi-k2-thinking

81 %

47 / 58

0.82 (0.50–1.00)

0

9

0

0

14.6 s

$0.01

2026-09-12

vertex-anthropic/claude-haiku-4-5

43 %

25 / 58

0.39 (0.00–1.00)

0

3

0

0

15.3 s

$0.02

2026-09-12

vertex-gemini/gemini-3.1-pro-preview

76 %

44 / 58

0.77 (0.50–1.00)

0

10

0

0

12.5 s

$0.02

2026-09-12 (off requested; pi-vertex floors a Gemini request to low (#78))

vertex-gemini/gemini-2.5-pro

62 %

36 / 58

0.65 (0.17–1.00)

0

17

0

0

19.4 s

$0.02

2026-09-12 (off requested; pi-vertex floors a Gemini request to low (#78))

vertex-gemini/gemini-3.5-flash

57 %

33 / 58

0.60 (0.25–1.00)

0

2

0

0

16.9 s

$0.02

2026-09-12 (off requested; pi-vertex floors a Gemini request to low (#78))

vertex-anthropic/claude-sonnet-4-6

88 %

51 / 58

0.89 (0.50–1.00)

0

4

0

0

24.6 s

$0.03

2026-09-12

vertex-anthropic/claude-opus-4-8

88 %

51 / 58

0.89 (0.67–1.00)

0

6

0

0

18.6 s

$0.06

2026-09-12

vertex-anthropic/claude-opus-5

97 %

56 / 58

0.96 (0.75–1.00)

0

1

0

0

22.0 s

$0.06

2026-09-12

vertex-anthropic/claude-fable-5-1

100 %

58 / 58

1.00

0

2

0

0

24.2 s

$0.12

2026-09-12 (off requested; claude-fable-5-1 does not offer it and thinks at its own discretion)

† introductory rate, through 2026-12-31: vertex-gemini/gemini-3.8-flash, vertex-gemini/gemini-3.7-flash. A cost argument from this table does not survive that date; the catalog’s standard rate is in its //cost-google note.

planning

Model Recall Caught / planted Per-case spread False flags (control) Extras Capped Errors Wall Cost / case Measured

vertex-maas/gpt-oss-120b

57 %

8 / 14

0.57

0

0

0

5.3 s

$0.0007

2026-09-12

vertex-maas/grok-4.1-fast-reasoning

71 %

10 / 14

0.71

0

0

0

24.1 s

$0.0008

2026-09-12

vertex-maas/qwen3-235b

64 %

9 / 14

0.64 (0.29–1.00)

0

0

0

10.8 s

$0.0014

2026-09-12

vertex-maas/qwen3-coder-480b

36 %

5 / 14

0.36 (0.29–0.43)

0

0

0

10.6 s

$0.0015

2026-09-12

vertex-maas/minimax-m2

29 %

4 / 14

0.29 (0.00–0.57)

0

0

0

11.3 s

$0.0017

2026-09-12

vertex-maas/grok-4.20-non-reasoning

43 %

6 / 14

0.43

0

0

0

1.0 s

$0.0040

2026-09-12

vertex-gemini/gemini-3.7-flash

71 %

10 / 14

0.71

0

0

0

10.9 s

$0.0045 †

2026-09-12 (off requested; pi-vertex floors a Gemini request to low (#78))

vertex-gemini/gemini-3.8-flash

79 %

11 / 14

0.79 (0.71–0.86)

0

0

0

11.1 s

$0.0045 †

2026-09-12 (off requested; pi-vertex floors a Gemini request to low (#78))

vertex-maas/grok-4.20-reasoning

64 %

9 / 14

0.64 (0.57–0.71)

0

0

0

11.2 s

$0.0053

2026-09-12

vertex-maas/qwen3-next-80b-thinking

71 %

10 / 14

0.71

0

0

0

29.6 s

$0.0057

2026-09-12

vertex-maas/kimi-k2-thinking

57 %

8 / 14

0.57 (0.43–0.71)

0

0

0

11.4 s

$0.0099

2026-09-12

vertex-gemini/gemini-3.1-pro-preview

79 %

11 / 14

0.79 (0.71–0.86)

0

0

0

9.6 s

$0.01

2026-09-12 (off requested; pi-vertex floors a Gemini request to low (#78))

vertex-gemini/gemini-3.5-flash

50 %

7 / 14

0.50 (0.43–0.57)

0

0

0

14.1 s

$0.02

2026-09-12 (off requested; pi-vertex floors a Gemini request to low (#78))

vertex-gemini/gemini-2.5-pro

57 %

8 / 14

0.57

0

0

0

20.6 s

$0.02

2026-09-12 (off requested; pi-vertex floors a Gemini request to low (#78))

vertex-anthropic/claude-haiku-4-5

79 %

11 / 14

0.79 (0.71–0.86)

0

0

0

17.4 s

$0.02

2026-09-12

vertex-anthropic/claude-sonnet-4-6

71 %

10 / 14

0.71

0

0

0

24.4 s

$0.04

2026-09-12

vertex-anthropic/claude-opus-5

86 %

12 / 14

0.86

0

0

0

25.4 s

$0.06

2026-09-12

vertex-anthropic/claude-opus-4-8

79 %

11 / 14

0.79 (0.71–0.86)

0

0

0

20.5 s

$0.07

2026-09-12

vertex-anthropic/claude-fable-5-1

93 %

13 / 14

0.93 (0.86–1.00)

0

0

0

24.0 s

$0.13

2026-09-12 (off requested; claude-fable-5-1 does not offer it and thinks at its own discretion)

† introductory rate, through 2026-12-31: vertex-gemini/gemini-3.7-flash, vertex-gemini/gemini-3.8-flash. A cost argument from this table does not survive that date; the catalog’s standard rate is in its //cost-google note.

judge

Model Exact UNSAFE Friction Deferred Capped Errors Wall Cost / case Measured

vertex-maas/gpt-oss-120b

85 %

3

43

41

0

0

0.5 s

$0.0001

2026-09-12

vertex-maas/grok-4.20-non-reasoning

90 %

14

20

35

0

0

0.5 s

$0.0003

2026-09-12 (low requested; grok-4.20-non-reasoning does not reason and is sent no level)

vertex-maas/qwen3-235b

97 %

0

20

18

0

0

0.7 s

$0.0000

2026-09-12 (low requested; qwen3-235b does not reason and is sent no level)

vertex-anthropic/claude-opus-4-8

96 %

0

26

30

0

0

1.3 s

$0.0029

2026-09-12

vertex-maas/grok-4.1-fast-reasoning

97 %

0

21

23

0

0

1.5 s

$0.0001

2026-09-12 (low requested; grok-4.1-fast-reasoning does not reason and is sent no level)

vertex-anthropic/claude-sonnet-4-6

97 %

0

24

27

0

0

1.6 s

$0.0025

2026-09-12

vertex-maas/qwen3-coder-480b

87 %

0

48

32

0

0

1.6 s

$0.0001

2026-09-12 (low requested; qwen3-coder-480b does not reason and is sent no level)

vertex-anthropic/claude-opus-5

93 %

0

32

32

0

1

1.8 s

$0.0024

2026-09-12

vertex-anthropic/claude-fable-5-1

64 %

0

41

139

0

104

1.8 s

$0.0027

2026-09-12

vertex-maas/kimi-k2-thinking

95 %

0

22

24

0

1

2.1 s

$0.0016

2026-09-12

vertex-maas/minimax-m2

90 %

16

23

31

2

0

2.8 s

$0.0006

2026-09-13

vertex-gemini/gemini-3.5-flash

92 %

0

35

27

0

0

4.1 s

$0.0027

2026-09-12

vertex-anthropic/claude-haiku-4-5

97 %

0

23

29

0

0

4.3 s

$0.0032

2026-09-12

vertex-gemini/gemini-3.7-flash

98 %

0

22

30

0

0

4.6 s

$0.0011 †

2026-09-12

vertex-maas/grok-4.20-reasoning

97 %

0

24

24

0

0

4.7 s

$0.0003

2026-09-12 (low requested; grok-4.20-reasoning does not reason and is sent no level)

vertex-gemini/gemini-3.1-pro-preview

96 %

0

25

23

0

0

4.9 s

$0.0032

2026-09-12

vertex-maas/qwen3-next-80b-thinking

90 %

2

39

49

0

0

5.1 s

$0.0013

2026-09-12

vertex-gemini/gemini-2.5-pro

73 %

0

94

92

0

0

8.1 s

$0.0065

2026-09-12

vertex-gemini/gemini-3.8-flash

94 %

0

32

30

0

0

10.1 s

$0.0013 †

2026-09-12

† introductory rate, through 2026-12-31: vertex-gemini/gemini-3.7-flash, vertex-gemini/gemini-3.8-flash. A cost argument from this table does not survive that date; the catalog’s standard rate is in its //cost-google note.

judge: experiments

Not read by the matrix: the same task under other knobs or another stack, and the task’s experiment arms.

Permutation Exact UNSAFE Friction Deferred Capped Errors Wall Cost / case Measured

vertex-anthropic/claude-haiku-4-5:off · bare

92 %

0

54

51

0

0

1.1 s

$0.0013

2026-09-12

vertex-anthropic/claude-haiku-4-5:minimal · bare

97 %

0

36

45

0

0

3.8 s

$0.0029

2026-09-12

vertex-anthropic/claude-haiku-4-5:low · bare

97 %

0

39

48

0

0

4.2 s

$0.0032

2026-09-12

vertex-anthropic/claude-haiku-4-5:low@budget=1024 · bare

96 %

0

37

48

0

0

4.0 s

$0.0029

2026-09-12

vertex-anthropic/claude-haiku-4-5:low@budget=2048 · bare

97 %

0

32

43

0

0

3.9 s

$0.0030

2026-09-12

vertex-anthropic/claude-sonnet-4-6:off · bare

95 %

0

44

41

0

0

1.9 s

$0.0023

2026-09-12

vertex-anthropic/claude-sonnet-4-6:low · bare

98 %

0

33

40

0

0

1.5 s

$0.0021

2026-09-12

explore

Model Path recall (worst corpus) By corpus By tier Yield rate Yield retries / yield Capped Tool failures (non-yield) Ended by itself Wall Cost / case Measured

vertex-maas/gpt-oss-120b

0.00

pi 0.00 · ruff 0.00

grep 0.00 · multi-hop 0.00 · design 0.00

0 %

0

0 % (0/26)

100 %

4.2 s

$0.0005

2026-09-11

vertex-maas/grok-4.1-fast-reasoning

0.85

pi 1.00 · ruff 0.85 (0.00–1.00)

grep 1.00 · multi-hop 0.94 · design 0.84

52 %

0.00

0

3 % (26/777)

100 %

25.3 s

$0.01

2026-09-11 (low requested; grok-4.1-fast-reasoning does not reason and is sent no level)

vertex-maas/qwen3-coder-480b

0.81

pi 0.82 (0.00–1.00) · ruff 0.81 (0.00–1.00)

grep 0.84 · multi-hop 0.79 · design 0.81

98 %

0.30

0

3 % (15/523)

98 %

23.2 s

$0.01

2026-09-11 (low requested; qwen3-coder-480b does not reason and is sent no level)

vertex-maas/qwen3-next-80b-thinking

0.00

pi 0.00 · ruff 0.06 (0.00–1.00)

grep 0.06 · multi-hop 0.03 · design 0.00

0 %

17

58 % (7/12)

65 %

54.9 s

$0.02

2026-09-11

vertex-maas/minimax-m2

0.66

pi 0.96 (0.00–1.00) · ruff 0.66 (0.00–1.00)

grep 0.81 · multi-hop 0.84 · design 0.77

98 %

0.17

0

1 % (10/782)

100 %

28.0 s

$0.02

2026-09-13

vertex-maas/qwen3-235b

0.46

pi 0.90 (0.00–1.00) · ruff 0.46 (0.00–1.00)

grep 0.94 · multi-hop 0.42 · design 0.69

38 %

0.00

1

2 % (32/1328)

73 %

28.7 s

$0.03

2026-09-11 (low requested; qwen3-235b does not reason and is sent no level)

vertex-gemini/gemini-3.8-flash

1.00

pi 1.00 · ruff 1.00

grep 1.00 · multi-hop 1.00 · design 1.00

100 %

0.00

0

1 % (3/515)

100 %

50.5 s

$0.05 †

2026-09-11

vertex-gemini/gemini-3.7-flash

1.00

pi 1.00 · ruff 1.00

grep 1.00 · multi-hop 1.00 · design 1.00

100 %

0.00

0

0 % (2/458)

100 %

54.8 s

$0.05 †

2026-09-11

vertex-gemini/gemini-2.5-pro

0.85

pi 0.85 (0.00–1.00) · ruff 0.91 (0.00–1.00)

grep 0.94 · multi-hop 0.81 · design 0.90

96 %

0.00

0

5 % (14/307)

100 %

48.7 s

$0.07

2026-09-11

vertex-maas/kimi-k2-thinking

0.69

pi 0.83 (0.00–1.00) · ruff 0.69 (0.00–1.00)

grep 0.94 · multi-hop 0.78 · design 0.57

54 %

0.08

0

2 % (13/696)

100 %

27.3 s

$0.08

2026-09-11

vertex-gemini/gemini-3.5-flash

0.90

pi 0.90 (0.00–1.00) · ruff 0.99 (0.67–1.00)

grep 1.00 · multi-hop 0.92 · design 0.92

96 %

0.00

0

0 % (0/482)

98 %

39.5 s

$0.10

2026-09-11

vertex-gemini/gemini-3.1-pro-preview

0.86

pi 0.94 (0.33–1.00) · ruff 0.86 (0.33–1.00)

grep 1.00 · multi-hop 0.82 · design 0.89

100 %

0.00

0

0 % (1/346)

100 %

40.9 s

$0.10

2026-09-11

vertex-anthropic/claude-haiku-4-5

0.89

pi 1.00 · ruff 0.89 (0.00–1.00)

grep 0.94 · multi-hop 1.00 · design 0.90

92 %

0.41

0

3 % (33/981)

94 %

52.0 s

$0.12

2026-09-11

vertex-maas/grok-4.20-reasoning

0.73

pi 0.73 (0.00–1.00) · ruff 0.84 (0.00–1.00)

grep 0.97 · multi-hop 0.72 · design 0.67

29 %

0.00

0

11 % (71/662)

100 %

21.8 s

$0.12

2026-09-11 (low requested; grok-4.20-reasoning does not reason and is sent no level)

vertex-anthropic/claude-sonnet-4-6

0.98

pi 0.98 (0.50–1.00) · ruff 0.99 (0.67–1.00)

grep 1.00 · multi-hop 0.97 · design 0.98

100 %

0.00

0

1 % (4/576)

100 %

44.7 s

$0.13

2026-09-11

vertex-maas/grok-4.20-non-reasoning

0.71

pi 0.71 (0.00–1.00) · ruff 0.74 (0.00–1.00)

grep 0.72 · multi-hop 0.81 · design 0.65

2 %

0.00

0

6 % (42/699)

100 %

17.6 s

$0.15

2026-09-11 (low requested; grok-4.20-non-reasoning does not reason and is sent no level)

vertex-anthropic/claude-opus-4-8

1.00

pi 1.00 · ruff 1.00

grep 1.00 · multi-hop 1.00 · design 1.00

96 %

0.43

0

1 % (4/371)

100 %

37.6 s

$0.21

2026-09-11

vertex-anthropic/claude-opus-5

1.00

pi 1.00 · ruff 1.00

grep 1.00 · multi-hop 1.00 · design 1.00

100 %

0.15

0

1 % (3/383)

100 %

32.2 s

$0.22

2026-09-11

vertex-anthropic/claude-fable-5-1

0.98

pi 1.00 · ruff 0.98 (0.50–1.00)

grep 1.00 · multi-hop 1.00 · design 0.97

100 %

0.04

0

0 % (2/417)

100 %

39.3 s

$0.32

2026-09-11

† introductory rate, through 2026-12-31: vertex-gemini/gemini-3.8-flash, vertex-gemini/gemini-3.7-flash. A cost argument from this table does not survive that date; the catalog’s standard rate is in its //cost-google note.

explore: experiments

Not read by the matrix: the same task under other knobs or another stack, and the task’s experiment arms.

Permutation Path recall (worst corpus) By corpus By tier Yield rate Yield retries / yield Capped Tool failures (non-yield) Ended by itself Wall Cost / case Measured

vertex-gemini/gemini-3.8-flash · shipped

1.00

pi 1.00 · ruff 1.00

grep 1.00 · multi-hop 1.00 · design 1.00

0 %

0

0 % (1/592)

100 %

66.6 s

$0.07 †

2026-09-13 (judge ×4 $0.02)

† introductory rate, through 2026-12-31: vertex-gemini/gemini-3.8-flash. A cost argument from this table does not survive that date; the catalog’s standard rate is in its //cost-google note.

verify

Model Faithful FALSE PASS Yield rate Bash calls Tool failures Ended by itself Errors Wall Cost / case Measured

vertex-maas/gpt-oss-120b

0 %

0

0 %

13

11

100 %

0

2.7 s

$0.0003

2026-09-11

vertex-maas/grok-4.1-fast-reasoning

58 %

0

75 %

12

10

100 %

0

5.5 s

$0.0006

2026-09-11 (low requested; grok-4.1-fast-reasoning does not reason and is sent no level)

vertex-maas/qwen3-235b

100 %

0

100 %

12

10

100 %

0

3.0 s

$0.0008

2026-09-11 (low requested; qwen3-235b does not reason and is sent no level)

vertex-maas/qwen3-coder-480b

67 %

0

100 %

12

10

100 %

0

3.6 s

$0.0008

2026-09-11 (low requested; qwen3-coder-480b does not reason and is sent no level)

vertex-maas/minimax-m2

92 %

0

100 %

12

11

100 %

0

9.2 s

$0.0018

2026-09-13

vertex-maas/kimi-k2-thinking

92 %

0

100 %

16

12

100 %

0

4.9 s

$0.0020

2026-09-11

vertex-gemini/gemini-3.8-flash

100 %

0

100 %

12

10

100 %

0

5.7 s

$0.0037 †

2026-09-11

vertex-gemini/gemini-3.7-flash

100 %

0

100 %

12

10

100 %

0

6.0 s

$0.0038 †

2026-09-11

vertex-maas/grok-4.20-non-reasoning

83 %

0

100 %

15

12

100 %

0

2.4 s

$0.0041

2026-09-11 (low requested; grok-4.20-non-reasoning does not reason and is sent no level)

vertex-maas/grok-4.20-reasoning

83 %

0

100 %

14

10

100 %

0

3.9 s

$0.0066

2026-09-11 (low requested; grok-4.20-reasoning does not reason and is sent no level)

vertex-gemini/gemini-2.5-pro

92 %

0

100 %

12

10

100 %

0

9.1 s

$0.0093

2026-09-11

vertex-gemini/gemini-3.1-pro-preview

75 %

0

100 %

14

11

100 %

0

6.8 s

$0.0097

2026-09-11

vertex-anthropic/claude-haiku-4-5

100 %

0

100 %

14

17

100 %

0

7.9 s

$0.01

2026-09-11

vertex-gemini/gemini-3.5-flash

75 %

0

100 %

15

13

100 %

0

8.5 s

$0.01

2026-09-11

vertex-maas/qwen3-next-80b-thinking

0 %

0

0 %

0

0

75 %

0

28.9 s

$0.01

2026-09-11

vertex-anthropic/claude-sonnet-4-6

100 %

0

100 %

12

6

100 %

0

7.5 s

$0.02

2026-09-11

vertex-anthropic/claude-opus-5

100 %

0

100 %

13

2

100 %

0

6.6 s

$0.03

2026-09-11

vertex-anthropic/claude-opus-4-8

100 %

0

100 %

14

11

100 %

0

11.4 s

$0.04

2026-09-11

vertex-anthropic/claude-fable-5-1

100 %

0

100 %

15

0

100 %

0

7.2 s

$0.05

2026-09-11

† introductory rate, through 2026-12-31: vertex-gemini/gemini-3.8-flash, vertex-gemini/gemini-3.7-flash. A cost argument from this table does not survive that date; the catalog’s standard rate is in its //cost-google note.

coder

Model Hidden tests Solved Load failures Capped Tool failures (non-bash) Bash exits ≠ 0 Ended by itself Wall Cost / case Measured

vertex-maas/gpt-oss-120b

0.03 (0.00–0.20)

0 / 6

0

0

0 % (0/2)

0 / 0

100 %

1.6 s

$0.0002

2026-09-11

vertex-maas/grok-4.1-fast-reasoning

0.52 (0.00–1.00)

3 / 6

4

0

3 % (2/73)

13 / 19

100 %

66.4 s

$0.0045

2026-09-11 (medium requested; grok-4.1-fast-reasoning does not reason and is sent no level)

vertex-maas/qwen3-235b

0.83 (0.00–1.00)

6 / 6

1

0

14 % (17/120)

15 / 27

100 %

34.7 s

$0.0098

2026-09-12 (medium requested; qwen3-235b does not reason and is sent no level)

vertex-maas/qwen3-coder-480b

0.92 (0.80–1.00)

7 / 6

0

0

3 % (4/117)

8 / 45

100 %

36.4 s

$0.02

2026-09-11 (medium requested; qwen3-coder-480b does not reason and is sent no level)

vertex-maas/minimax-m2

0.75 (0.00–1.00)

6 / 6

0

0

4 % (4/113)

10 / 47

100 %

67.1 s

$0.02

2026-09-13

vertex-maas/qwen3-next-80b-thinking

0.17 (0.00–0.80)

0 / 6

0

6

33 % (1/3)

0 / 0

50 %

183.1 s

$0.03

2026-09-12

vertex-anthropic/claude-haiku-4-5

0.97 (0.80–1.00)

10 / 6

0

0

10 % (13/127)

7 / 43

100 %

61.2 s

$0.09

2026-09-11

vertex-maas/grok-4.20-reasoning

0.88 (0.00–1.00)

9 / 6

1

0

21 % (45/211)

18 / 48

92 %

49.7 s

$0.09

2026-09-11 (medium requested; grok-4.20-reasoning does not reason and is sent no level)

vertex-anthropic/claude-sonnet-4-6

1.00

12 / 6

0

0

0 % (0/41)

3 / 32

100 %

52.1 s

$0.11

2026-09-11

vertex-gemini/gemini-3.7-flash

0.98 (0.80–1.00)

11 / 6

0

0

9 % (11/124)

29 / 68

100 %

132.0 s

$0.13 †

2026-09-11

vertex-gemini/gemini-2.5-pro

0.67 (0.00–1.00)

7 / 6

1

0

7 % (8/108)

22 / 41

75 %

61.1 s

$0.13

2026-09-12

vertex-anthropic/claude-opus-4-8

0.97 (0.80–1.00)

10 / 6

0

0

5 % (2/40)

0 / 14

100 %

28.4 s

$0.14

2026-09-11

vertex-maas/kimi-k2-thinking

0.92 (0.00–1.00)

11 / 6

1

0

7 % (11/153)

38 / 161

100 %

80.5 s

$0.16

2026-09-12

vertex-anthropic/claude-opus-5

1.00

12 / 6

0

0

0 % (0/29)

5 / 28

100 %

32.2 s

$0.17

2026-09-11

vertex-gemini/gemini-3.5-flash

0.97 (0.80–1.00)

10 / 6

0

0

0 % (0/101)

17 / 46

100 %

84.3 s

$0.18

2026-09-11

vertex-gemini/gemini-3.1-pro-preview

1.00

12 / 6

0

0

2 % (1/56)

19 / 86

100 %

102.4 s

$0.23

2026-09-11 (medium requested; gemini-3.1-pro-preview does not offer it and pi clamps to high)

vertex-anthropic/claude-fable-5-1

1.00

12 / 6

0

0

0 % (0/18)

12 / 28

100 %

31.0 s

$0.24

2026-09-11

vertex-maas/grok-4.20-non-reasoning

0.85 (0.00–1.00)

7 / 6

0

0

25 % (121/483)

32 / 77

83 %

82.1 s

$0.24

2026-09-12 (medium requested; grok-4.20-non-reasoning does not reason and is sent no level)

vertex-gemini/gemini-3.8-flash

1.00

12 / 6

0

0

6 % (8/140)

43 / 173

100 %

395.8 s

$0.33 †

2026-09-12

† introductory rate, through 2026-12-31: vertex-gemini/gemini-3.7-flash, vertex-gemini/gemini-3.8-flash. A cost argument from this table does not survive that date; the catalog’s standard rate is in its //cost-google note.

coder: experiments

Not read by the matrix: the same task under other knobs or another stack, and the task’s experiment arms.

Permutation Hidden tests Solved Load failures Capped Tool failures (non-bash) Bash exits ≠ 0 Ended by itself Wall Cost / case Measured

vertex-anthropic/claude-opus-5 · shipped

1.00

12 / 6

0

0

0 % (0/32)

8 / 32

100 %

42.0 s

$0.25

2026-09-13 (judge ×4 $0.02)

vertex-anthropic/claude-haiku-4-5 · shipped

0.95 (0.80–1.00)

9 / 6

0

0

9 % (10/116)

13 / 57

100 %

75.6 s

$0.10

2026-09-13 (judge ×8 $0.04)

vertex-gemini/gemini-3.8-flash · shipped

1.00

12 / 6

0

0

8 % (10/128)

48 / 153

100 %

362.7 s

$0.33 †

2026-09-13 (judge ×129 $0.42)

† introductory rate, through 2026-12-31: vertex-gemini/gemini-3.8-flash. A cost argument from this table does not survive that date; the catalog’s standard rate is in its //cost-google note.

integrity

Model Intact Whole / completed Capped Errors Wall Measured

vertex-maas/grok-4.20-non-reasoning

100 %

12 / 12

0

0

0.3 s

2026-09-10

vertex-anthropic/claude-haiku-4-5

100 %

12 / 12

0

0

0.4 s

2026-09-10

vertex-maas/qwen3-235b

100 %

12 / 12

0

0

0.4 s

2026-09-10

vertex-maas/gpt-oss-120b

83 %

10 / 12

0

0

0.6 s

2026-09-10

vertex-maas/kimi-k2-thinking

100 %

12 / 12

0

0

0.6 s

2026-09-10

vertex-anthropic/claude-sonnet-4-6

100 %

12 / 12

0

0

1.0 s

2026-09-10

vertex-maas/grok-4.20-reasoning

100 %

12 / 12

0

0

1.1 s

2026-09-10

vertex-anthropic/claude-opus-5

100 %

12 / 12

0

0

1.1 s

2026-09-10

vertex-maas/grok-4.1-fast-reasoning

100 %

12 / 12

0

0

1.3 s

2026-09-10

vertex-maas/qwen3-coder-480b

100 %

12 / 12

0

0

1.4 s

2026-09-10

vertex-anthropic/claude-fable-5-1

100 %

12 / 12

0

0

1.4 s

2026-09-10 (off requested; claude-fable-5-1 does not offer it and thinks at its own discretion)

vertex-maas/minimax-m2

100 %

12 / 12

0

0

1.4 s

2026-09-12

vertex-anthropic/claude-opus-4-8

100 %

12 / 12

0

0

1.5 s

2026-09-10

vertex-gemini/gemini-3.7-flash

100 %

12 / 12

0

0

3.5 s

2026-09-10 (off requested; pi-vertex floors a Gemini request to low (#78))

vertex-gemini/gemini-3.8-flash

100 %

12 / 12

0

0

3.6 s

2026-09-10 (off requested; pi-vertex floors a Gemini request to low (#78))

vertex-gemini/gemini-3.5-flash

100 %

12 / 12

0

0

3.8 s

2026-09-10 (off requested; pi-vertex floors a Gemini request to low (#78))

vertex-gemini/gemini-3.1-pro-preview

100 %

12 / 12

0

0

4.7 s

2026-09-10 (off requested; pi-vertex floors a Gemini request to low (#78))

vertex-gemini/gemini-2.5-pro

100 %

12 / 12

0

0

5.9 s

2026-09-10 (off requested; pi-vertex floors a Gemini request to low (#78))

vertex-maas/qwen3-next-80b-thinking

0 %

0 / 5

6

1

13.5 s

2026-09-11

The matrix

Derived from `tools/model-battery/battery.json’s rules over each role’s production cells; each cell shows the verdict and the numbers that decided it.

Model reviewer planning judge explore verify coder

vertex-anthropic/claude-opus-5

fit
recall 0.97 ≥ 0.9 ✓; integrity 1.00 ≥ 0.98 ✓

usable
recall 0.86 ≥ 0.75 ✓

fit
unsafe 0.00 = 0 ✓; exact 0.93 ≥ 0.9 ✓; wallMsMedianMax 1750.00 ≤ 8000 ✓

fit
pathRecallMin 1.00 ≥ 0.95 ✓; toolFailureRateMax 0.01 ≤ 0.05 ✓; integrity 1.00 ≥ 0.98 ✓

fit
faithful 1.00 ≥ 0.9 ✓; falsePassMax 0.00 ≤ 0 ✓; yieldRate 1.00 ≥ 0.9 ✓; integrity 1.00 ≥ 0.98 ✓

fit
hiddenPassRate.mean 1.00 ≥ 0.8 ✓; toolFailureRateMax 0.00 ≤ 0.05 ✓; endedBySelf 1.00 ≥ 0.9 ✓; integrity 1.00 ≥ 0.98 ✓

vertex-anthropic/claude-opus-4-8

usable
recall 0.88 ≥ 0.75 ✓; integrity 1.00 ≥ 0.98 ✓

usable
recall 0.79 ≥ 0.75 ✓

fit
unsafe 0.00 = 0 ✓; exact 0.96 ≥ 0.9 ✓; wallMsMedianMax 1318.00 ≤ 8000 ✓

fit
pathRecallMin 1.00 ≥ 0.95 ✓; toolFailureRateMax 0.01 ≤ 0.05 ✓; integrity 1.00 ≥ 0.98 ✓

fit
faithful 1.00 ≥ 0.9 ✓; falsePassMax 0.00 ≤ 0 ✓; yieldRate 1.00 ≥ 0.9 ✓; integrity 1.00 ≥ 0.98 ✓

fit
hiddenPassRate.mean 0.97 ≥ 0.8 ✓; toolFailureRateMax 0.05 ≤ 0.05 ✓; endedBySelf 1.00 ≥ 0.9 ✓; integrity 1.00 ≥ 0.98 ✓

vertex-anthropic/claude-fable-5-1

fit
recall 1.00 ≥ 0.9 ✓; integrity 1.00 ≥ 0.98 ✓

fit
recall 0.93 ≥ 0.9 ✓

no
unsafe 0.00 = 0 ✓; exact 0.64 ≥ 0.9 ✗; wallMsMedianMax 1784.00 ≤ 8000 ✓

fit
pathRecallMin 0.98 ≥ 0.95 ✓; toolFailureRateMax 0.00 ≤ 0.05 ✓; integrity 1.00 ≥ 0.98 ✓

fit
faithful 1.00 ≥ 0.9 ✓; falsePassMax 0.00 ≤ 0 ✓; yieldRate 1.00 ≥ 0.9 ✓; integrity 1.00 ≥ 0.98 ✓

fit
hiddenPassRate.mean 1.00 ≥ 0.8 ✓; toolFailureRateMax 0.00 ≤ 0.05 ✓; endedBySelf 1.00 ≥ 0.9 ✓; integrity 1.00 ≥ 0.98 ✓

vertex-anthropic/claude-sonnet-4-6

usable
recall 0.88 ≥ 0.75 ✓; integrity 1.00 ≥ 0.98 ✓

no
recall 0.71 ≥ 0.9 ✗

fit
unsafe 0.00 = 0 ✓; exact 0.97 ≥ 0.9 ✓; wallMsMedianMax 1563.00 ≤ 8000 ✓

fit
pathRecallMin 0.98 ≥ 0.95 ✓; toolFailureRateMax 0.01 ≤ 0.05 ✓; integrity 1.00 ≥ 0.98 ✓

fit
faithful 1.00 ≥ 0.9 ✓; falsePassMax 0.00 ≤ 0 ✓; yieldRate 1.00 ≥ 0.9 ✓; integrity 1.00 ≥ 0.98 ✓

fit
hiddenPassRate.mean 1.00 ≥ 0.8 ✓; toolFailureRateMax 0.00 ≤ 0.05 ✓; endedBySelf 1.00 ≥ 0.9 ✓; integrity 1.00 ≥ 0.98 ✓

vertex-anthropic/claude-haiku-4-5

no
recall 0.43 ≥ 0.9 ✗; integrity 1.00 ≥ 0.98 ✓

usable
recall 0.79 ≥ 0.75 ✓

fit
unsafe 0.00 = 0 ✓; exact 0.97 ≥ 0.9 ✓; wallMsMedianMax 4300.00 ≤ 8000 ✓

usable
pathRecallMin 0.89 ≥ 0.85 ✓; toolFailureRateMax 0.03 ≤ 0.1 ✓; integrity 1.00 ≥ 0.98 ✓

fit
faithful 1.00 ≥ 0.9 ✓; falsePassMax 0.00 ≤ 0 ✓; yieldRate 1.00 ≥ 0.9 ✓; integrity 1.00 ≥ 0.98 ✓

no
hiddenPassRate.mean 0.97 ≥ 0.8 ✓; toolFailureRateMax 0.10 ≤ 0.05 ✗; endedBySelf 1.00 ≥ 0.9 ✓; integrity 1.00 ≥ 0.98 ✓

vertex-gemini/gemini-3.8-flash

no
recall 0.74 ≥ 0.9 ✗; integrity 1.00 ≥ 0.98 ✓

usable
recall 0.79 ≥ 0.75 ✓

no
unsafe 0.00 = 0 ✓; exact 0.94 ≥ 0.9 ✓; wallMsMedianMax 10124.00 ≤ 8000 ✗

fit
pathRecallMin 1.00 ≥ 0.95 ✓; toolFailureRateMax 0.01 ≤ 0.05 ✓; integrity 1.00 ≥ 0.98 ✓

fit
faithful 1.00 ≥ 0.9 ✓; falsePassMax 0.00 ≤ 0 ✓; yieldRate 1.00 ≥ 0.9 ✓; integrity 1.00 ≥ 0.98 ✓

usable
hiddenPassRate.mean 1.00 ≥ 0.6 ✓; toolFailureRateMax 0.06 ≤ 0.1 ✓; endedBySelf 1.00 ≥ 0.8 ✓; integrity 1.00 ≥ 0.98 ✓

vertex-gemini/gemini-3.7-flash

no
recall 0.74 ≥ 0.9 ✗; integrity 1.00 ≥ 0.98 ✓

no
recall 0.71 ≥ 0.9 ✗

fit
unsafe 0.00 = 0 ✓; exact 0.98 ≥ 0.9 ✓; wallMsMedianMax 4632.00 ≤ 8000 ✓

fit
pathRecallMin 1.00 ≥ 0.95 ✓; toolFailureRateMax 0.00 ≤ 0.05 ✓; integrity 1.00 ≥ 0.98 ✓

fit
faithful 1.00 ≥ 0.9 ✓; falsePassMax 0.00 ≤ 0 ✓; yieldRate 1.00 ≥ 0.9 ✓; integrity 1.00 ≥ 0.98 ✓

usable
hiddenPassRate.mean 0.98 ≥ 0.6 ✓; toolFailureRateMax 0.09 ≤ 0.1 ✓; endedBySelf 1.00 ≥ 0.8 ✓; integrity 1.00 ≥ 0.98 ✓

vertex-gemini/gemini-3.5-flash

no
recall 0.57 ≥ 0.9 ✗; integrity 1.00 ≥ 0.98 ✓

no
recall 0.50 ≥ 0.9 ✗

fit
unsafe 0.00 = 0 ✓; exact 0.92 ≥ 0.9 ✓; wallMsMedianMax 4136.50 ≤ 8000 ✓

usable
pathRecallMin 0.90 ≥ 0.85 ✓; toolFailureRateMax 0.00 ≤ 0.1 ✓; integrity 1.00 ≥ 0.98 ✓

usable
faithful 0.75 ≥ 0.75 ✓; falsePassMax 0.00 ≤ 0 ✓; integrity 1.00 ≥ 0.98 ✓

fit
hiddenPassRate.mean 0.97 ≥ 0.8 ✓; toolFailureRateMax 0.00 ≤ 0.05 ✓; endedBySelf 1.00 ≥ 0.9 ✓; integrity 1.00 ≥ 0.98 ✓

vertex-gemini/gemini-3.1-pro-preview

usable
recall 0.76 ≥ 0.75 ✓; integrity 1.00 ≥ 0.98 ✓

usable
recall 0.79 ≥ 0.75 ✓

fit
unsafe 0.00 = 0 ✓; exact 0.96 ≥ 0.9 ✓; wallMsMedianMax 4946.00 ≤ 8000 ✓

usable
pathRecallMin 0.86 ≥ 0.85 ✓; toolFailureRateMax 0.00 ≤ 0.1 ✓; integrity 1.00 ≥ 0.98 ✓

usable
faithful 0.75 ≥ 0.75 ✓; falsePassMax 0.00 ≤ 0 ✓; integrity 1.00 ≥ 0.98 ✓

fit
hiddenPassRate.mean 1.00 ≥ 0.8 ✓; toolFailureRateMax 0.02 ≤ 0.05 ✓; endedBySelf 1.00 ≥ 0.9 ✓; integrity 1.00 ≥ 0.98 ✓

vertex-gemini/gemini-2.5-pro

no
recall 0.62 ≥ 0.9 ✗; integrity 1.00 ≥ 0.98 ✓

no
recall 0.57 ≥ 0.9 ✗

no
unsafe 0.00 = 0 ✓; exact 0.73 ≥ 0.9 ✗; wallMsMedianMax 8094.50 ≤ 8000 ✗

usable
pathRecallMin 0.85 ≥ 0.85 ✓; toolFailureRateMax 0.05 ≤ 0.1 ✓; integrity 1.00 ≥ 0.98 ✓

fit
faithful 0.92 ≥ 0.9 ✓; falsePassMax 0.00 ≤ 0 ✓; yieldRate 1.00 ≥ 0.9 ✓; integrity 1.00 ≥ 0.98 ✓

no
hiddenPassRate.mean 0.67 ≥ 0.8 ✗; toolFailureRateMax 0.07 ≤ 0.05 ✗; endedBySelf 0.75 ≥ 0.9 ✗; integrity 1.00 ≥ 0.98 ✓

vertex-maas/gpt-oss-120b

no
recall 0.05 ≥ 0.9 ✗; integrity 0.83 ≥ 0.98 ✗

no
recall 0.57 ≥ 0.9 ✗

no
unsafe 3.00 = 0 ✗; exact 0.85 ≥ 0.9 ✗; wallMsMedianMax 515.00 ≤ 8000 ✓

no
pathRecallMin 0.00 ≥ 0.95 ✗; toolFailureRateMax 0.00 ≤ 0.05 ✓; integrity 0.83 ≥ 0.98 ✗

no
faithful 0.00 ≥ 0.9 ✗; falsePassMax 0.00 ≤ 0 ✓; yieldRate 0.00 ≥ 0.9 ✗; integrity 0.83 ≥ 0.98 ✗

no
hiddenPassRate.mean 0.03 ≥ 0.8 ✗; toolFailureRateMax 0.00 ≤ 0.05 ✓; endedBySelf 1.00 ≥ 0.9 ✓; integrity 0.83 ≥ 0.98 ✗

vertex-maas/grok-4.20-reasoning

usable
recall 0.78 ≥ 0.75 ✓; integrity 1.00 ≥ 0.98 ✓

no
recall 0.64 ≥ 0.9 ✗

fit
unsafe 0.00 = 0 ✓; exact 0.97 ≥ 0.9 ✓; wallMsMedianMax 4724.50 ≤ 8000 ✓

no
pathRecallMin 0.73 ≥ 0.95 ✗; toolFailureRateMax 0.11 ≤ 0.05 ✗; integrity 1.00 ≥ 0.98 ✓

usable
faithful 0.83 ≥ 0.75 ✓; falsePassMax 0.00 ≤ 0 ✓; integrity 1.00 ≥ 0.98 ✓

no
hiddenPassRate.mean 0.88 ≥ 0.8 ✓; toolFailureRateMax 0.21 ≤ 0.05 ✗; endedBySelf 0.92 ≥ 0.9 ✓; integrity 1.00 ≥ 0.98 ✓

vertex-maas/grok-4.20-non-reasoning

no
recall 0.50 ≥ 0.9 ✗; integrity 1.00 ≥ 0.98 ✓

no
recall 0.43 ≥ 0.9 ✗

no
unsafe 14.00 = 0 ✗; exact 0.90 ≥ 0.9 ✗; wallMsMedianMax 520.50 ≤ 8000 ✓

no
pathRecallMin 0.71 ≥ 0.95 ✗; toolFailureRateMax 0.06 ≤ 0.05 ✗; integrity 1.00 ≥ 0.98 ✓

usable
faithful 0.83 ≥ 0.75 ✓; falsePassMax 0.00 ≤ 0 ✓; integrity 1.00 ≥ 0.98 ✓

no
hiddenPassRate.mean 0.85 ≥ 0.8 ✓; toolFailureRateMax 0.25 ≤ 0.05 ✗; endedBySelf 0.83 ≥ 0.9 ✗; integrity 1.00 ≥ 0.98 ✓

vertex-maas/grok-4.1-fast-reasoning

no
recall 0.66 ≥ 0.9 ✗; integrity 1.00 ≥ 0.98 ✓

no
recall 0.71 ≥ 0.9 ✗

fit
unsafe 0.00 = 0 ✓; exact 0.97 ≥ 0.9 ✓; wallMsMedianMax 1526.50 ≤ 8000 ✓

usable
pathRecallMin 0.85 ≥ 0.85 ✓; toolFailureRateMax 0.03 ≤ 0.1 ✓; integrity 1.00 ≥ 0.98 ✓

no
faithful 0.58 ≥ 0.9 ✗; falsePassMax 0.00 ≤ 0 ✓; yieldRate 0.75 ≥ 0.9 ✗; integrity 1.00 ≥ 0.98 ✓

no
hiddenPassRate.mean 0.52 ≥ 0.8 ✗; toolFailureRateMax 0.03 ≤ 0.05 ✓; endedBySelf 1.00 ≥ 0.9 ✓; integrity 1.00 ≥ 0.98 ✓

vertex-maas/kimi-k2-thinking

usable
recall 0.81 ≥ 0.75 ✓; integrity 1.00 ≥ 0.98 ✓

no
recall 0.57 ≥ 0.9 ✗

fit
unsafe 0.00 = 0 ✓; exact 0.95 ≥ 0.9 ✓; wallMsMedianMax 2102.00 ≤ 8000 ✓

no
pathRecallMin 0.69 ≥ 0.95 ✗; toolFailureRateMax 0.02 ≤ 0.05 ✓; integrity 1.00 ≥ 0.98 ✓

fit
faithful 0.92 ≥ 0.9 ✓; falsePassMax 0.00 ≤ 0 ✓; yieldRate 1.00 ≥ 0.9 ✓; integrity 1.00 ≥ 0.98 ✓

usable
hiddenPassRate.mean 0.92 ≥ 0.6 ✓; toolFailureRateMax 0.07 ≤ 0.1 ✓; endedBySelf 1.00 ≥ 0.8 ✓; integrity 1.00 ≥ 0.98 ✓

vertex-maas/qwen3-coder-480b

no
recall 0.26 ≥ 0.9 ✗; integrity 1.00 ≥ 0.98 ✓

no
recall 0.36 ≥ 0.9 ✗

no
unsafe 0.00 = 0 ✓; exact 0.87 ≥ 0.9 ✗; wallMsMedianMax 1624.00 ≤ 8000 ✓

no
pathRecallMin 0.81 ≥ 0.95 ✗; toolFailureRateMax 0.03 ≤ 0.05 ✓; integrity 1.00 ≥ 0.98 ✓

no
faithful 0.67 ≥ 0.9 ✗; falsePassMax 0.00 ≤ 0 ✓; yieldRate 1.00 ≥ 0.9 ✓; integrity 1.00 ≥ 0.98 ✓

fit
hiddenPassRate.mean 0.92 ≥ 0.8 ✓; toolFailureRateMax 0.03 ≤ 0.05 ✓; endedBySelf 1.00 ≥ 0.9 ✓; integrity 1.00 ≥ 0.98 ✓

vertex-maas/qwen3-next-80b-thinking

no (budget-starved: 2/16 capped)
recall 0.34 ≥ 0.9 ✗; integrity 0.00 ≥ 0.98 ✗

no
recall 0.71 ≥ 0.9 ✗

no
unsafe 2.00 = 0 ✗; exact 0.90 ≥ 0.9 ✗; wallMsMedianMax 5141.50 ≤ 8000 ✓

no (budget-starved: 17/48 capped)
pathRecallMin 0.00 ≥ 0.95 ✗; toolFailureRateMax 0.58 ≤ 0.05 ✗; integrity 0.00 ≥ 0.98 ✗

no (budget-starved: 3/12 capped)
faithful 0.00 ≥ 0.9 ✗; falsePassMax 0.00 ≤ 0 ✓; yieldRate 0.00 ≥ 0.9 ✗; integrity 0.00 ≥ 0.98 ✗

no (budget-starved: 6/12 capped)
hiddenPassRate.mean 0.17 ≥ 0.8 ✗; toolFailureRateMax 0.33 ≤ 0.05 ✗; endedBySelf 0.50 ≥ 0.9 ✗; integrity 0.00 ≥ 0.98 ✗

vertex-maas/qwen3-235b

no
recall 0.38 ≥ 0.9 ✗; integrity 1.00 ≥ 0.98 ✓

no
recall 0.64 ≥ 0.9 ✗

fit
unsafe 0.00 = 0 ✓; exact 0.97 ≥ 0.9 ✓; wallMsMedianMax 702.50 ≤ 8000 ✓

no (budget-starved: 1/48 capped)
pathRecallMin 0.46 ≥ 0.95 ✗; toolFailureRateMax 0.02 ≤ 0.05 ✓; integrity 1.00 ≥ 0.98 ✓

fit
faithful 1.00 ≥ 0.9 ✓; falsePassMax 0.00 ≤ 0 ✓; yieldRate 1.00 ≥ 0.9 ✓; integrity 1.00 ≥ 0.98 ✓

no
hiddenPassRate.mean 0.83 ≥ 0.8 ✓; toolFailureRateMax 0.14 ≤ 0.05 ✗; endedBySelf 1.00 ≥ 0.9 ✓; integrity 1.00 ≥ 0.98 ✓

vertex-maas/minimax-m2

no
recall 0.45 ≥ 0.9 ✗; integrity 1.00 ≥ 0.98 ✓

no
recall 0.29 ≥ 0.9 ✗

no (budget-starved: 2/286 capped)
unsafe 16.00 = 0 ✗; exact 0.90 ≥ 0.9 ✗; wallMsMedianMax 2756.00 ≤ 8000 ✓

no
pathRecallMin 0.66 ≥ 0.95 ✗; toolFailureRateMax 0.01 ≤ 0.05 ✓; integrity 1.00 ≥ 0.98 ✓

fit
faithful 0.92 ≥ 0.9 ✓; falsePassMax 0.00 ≤ 0 ✓; yieldRate 1.00 ≥ 0.9 ✓; integrity 1.00 ≥ 0.98 ✓

usable
hiddenPassRate.mean 0.75 ≥ 0.6 ✓; toolFailureRateMax 0.04 ≤ 0.1 ✓; endedBySelf 1.00 ≥ 0.8 ✓; integrity 1.00 ≥ 0.98 ✓

The shipped defaults, argued from the table

Argued on the retired 2026-09-08 run (its records live in git history under tools/model-battery/results/, removed when the store replaced it - #108); the cells above re-argue each default as the store fills, and a default whose cell reads not measured stands on that run’s numbers until then. reviewer: the pre-commit cold reader (chooseReviewer’s floor and tiers pick at or above the author); planning: the same reader on a plan; judge: the gatekeeper for auto and plan - moved from Haiku 4.5 low to Sonnet 4.6 low on 2026-09-12 (#102) on the store’s bake-off (366 verdicts each): within three of each other on exact (352 vs 349), neither unsafe, the same price per call, 1.5 s median against 4.2 s. On the shipped-stack coder cells the faster judge took Haiku’s lane from 91 s to 69 s and its cost level with bare; Opus 5, whose commands are almost all rule-allowed, made three gated calls in twelve lanes and did not move - its shipped overhead is the larger prompt, not the judge (#100); explore: the Explore helper - moved from Haiku 4.5 to a Gemini Flash on 2026-09-08 (#89) on the blind explore set: Haiku 0.89 worst-corpus recall with a repeatable wrong-subsystem pick on the large repository, 90 % yield, 3 % tool failures; the 3.x Flashes 0.97-1.00 on every tier of both corpora, 100 % yield, under 2 % tool failures, at the same standard price (Flash’s current rate is introductory through 2026-12-31 and was not the argument). 3.8 over 3.7 Flash: level on a fresh n=96 confirmation (0.991 vs 0.997, two partial misses vs one, both 100 % yield, wall 53 s vs 59 s), and the newer model carries the longer support runway; the coder-proxy slowness that first pointed at 3.7 did not appear in explore. Re-measured in the store 2026-09-11 (every packaged model, 48 lanes each, #108): 3.8 and 3.7 Flash 1.00 on both corpora at $0.05 a lane, the Claude flagships 0.98-1.00 at three to six times the price, Haiku 0.89 on the large repository again. The fit rule moved from 0.8 to 0.95 worst-corpus recall on that run (#101): at 0.8 thirteen models ranked fit, Grok 4.1 Fast, Qwen3 Coder and MiniMax among them at 0.81-0.85 - wrong on one question in five or six - and price would have decided among them; at 0.95 six rank fit and the cheapest is the default. verify: the Verify helper - moved from Haiku 4.5 to Gemini 3.8 Flash on 2026-09-12 (#117) on the verify task’s first full fill (run a check with a known outcome, report it faithfully; six cases, two runs, every packaged model; the matrix’s falsePassMax 0 admits no FALSE PASS and none occurred): 3.8 Flash 12/12 faithful with 100 % yield at $0.004 a lane against Haiku’s 12/12 at $0.012, and one model for both helper roles. Haiku had stayed on Verify when Explore moved (#89) because the coder proxy had shown 3.8 Flash weak with bash; the role’s own task does not show it. coder: the auto-mode session model. The plan-mode author (Fable) writes plans, which this battery does not score.

  • reviewervertex-anthropic/claude-fable-5-1: fit — recall 1.00 ≥ 0.9 ✓; integrity 1.00 ≥ 0.98 ✓

  • planningvertex-anthropic/claude-fable-5-1: fit — recall 0.93 ≥ 0.9 ✓

  • judgevertex-anthropic/claude-sonnet-4-6: fit — unsafe 0.00 = 0 ✓; exact 0.97 ≥ 0.9 ✓; wallMsMedianMax 1563.00 ≤ 8000 ✓

  • explorevertex-gemini/gemini-3.8-flash: fit — pathRecallMin 1.00 ≥ 0.95 ✓; toolFailureRateMax 0.01 ≤ 0.05 ✓; integrity 1.00 ≥ 0.98 ✓

  • codervertex-anthropic/claude-opus-5: fit — hiddenPassRate.mean 1.00 ≥ 0.8 ✓; toolFailureRateMax 0.00 ≤ 0.05 ✓; endedBySelf 1.00 ≥ 0.9 ✓; integrity 1.00 ≥ 0.98 ✓

  • verifyvertex-gemini/gemini-3.8-flash: fit — faithful 1.00 ≥ 0.9 ✓; falsePassMax 0.00 ≤ 0 ✓; yieldRate 1.00 ≥ 0.9 ✓; integrity 1.00 ≥ 0.98 ✓

Edit this page · latest