AI Model Scoring Benchmark

Purpose & Methodology

GitVelocity scores every pull request on a 0-100 scale across 6 sub-categories using a large language model — Anthropic Claude by default, with other providers available as alternatives. To choose and validate the scoring model, we benchmark candidates against a fixed reference model. Claude Fable 5 is that reference — and this page compares the current candidate set against it on four dimensions:

  1. Cost - Token usage and USD cost per review
  2. Accuracy - Score deviation from the Fable 5 reference
  3. Stability - Variance across independent runs per model
  4. Speed - Observed API latency, measured for three models only (new in this update)

This is the v6 edition, extended. Ten models scored the same 20 pull-request corpus on the same prompt and the same rubric, 3 independent runs per model per PR — 600 scored calls in total, run in August 2026. Two further models, Claude Fable 5 and GPT-5.6 Sol, were added on August 13 against the same 20-PR corpus and the same 3-runs-per-PR method. Scores come only from this benchmark — never from your production data.

Run provenance. The original ten-model dataset contains 600 successful scored calls, completed in two passes: the initial run on August 6 and a resume on August 10 after 123 calls failed. The resumed calls used the same prompt, rubric, and reasoning configuration and completed the missing model/PR/run slots, but those rows span four days rather than one uninterrupted run. The Fable 5 and Sol runs, and the latency measurements, were all made on August 13, 2026.

The twelve models are Fable 5 (the reference), Opus 4.7, Opus 5, Sonnet 4.6, GPT-5.6 Terra, GPT-5.6 Sol, GPT-5.6 Luna, Kimi K3, GLM 5.1, GLM 5.2, Mistral Small 4, and Mistral Medium 3.5.

The reference changed: Opus 4.7 → Fable 5

Earlier versions of this page measured every model against Opus 4.7. As of this update the reference is Claude Fable 5, and every MAD, max deviation, bias and correlation figure below has been recomputed against it.

The reason is stability, and only stability. A reference's job is to be a steady yardstick: if the ruler moves between runs, every measurement taken against it inherits that movement. Fable 5 is nearly twice as steady run-to-run as Opus 4.7 — 3.3% average variation against 6.0%, and a worst case of 10.2% against 24.4%. It is the steadiest model this benchmark has measured.

This is the same criterion applied the same way as last time, not a new one invented to fit a result. When Opus 4.8 was considered for the reference role, it was declined because it was noisier run-to-run than Opus 4.7 — the incumbent kept the job on stability. Fable 5 now takes it on stability from Opus 4.7 for exactly that reason.

This is a benchmark-measurement change only. No customer's score moves. The reference model is a measuring stick used on this page; it is not part of scoring. claude-opus-4-7 appears nowhere in the production scoring prompt, and neither does Fable 5. Your PRs are scored by whichever model you have selected — by default Sonnet 4.6 — and that is unchanged. If your scores were 41 and 62 yesterday, they are 41 and 62 today.

What this does change is the wording of some conclusions, because "closest to the reference" is a claim about a specific reference. Where a previous recommendation rested on agreement with Opus 4.7, it has been rewritten, and the change is called out at the point it matters — most visibly for Kimi K3, which moved from 1st to 5th on reference agreement.

⚠️ Two rubric versions are mixed on this page

The Fable 5, Sol and latency runs used rubric v7; the other ten models were scored on v6. These are different rubric revisions, so this is not a perfectly like-for-like comparison and you should know where the seam is.

It is defensible for the total score because the v6→v7 change was itself measured: the total score moved less than the models' own run-to-run noise, while two sub-scores did shift measurably (quality −0.80, risk +0.68). So total-score MAD, bias, correlation and stability comparisons across the seam are sound, and that is what the accuracy, stability, by-size and by-language tables report.

It is not defensible per sub-score, so the per-sub-score table has not been re-based. Re-basing it would push v7's measured quality and risk shifts into every cell as if they were model differences. That table alone remains anchored on Opus 4.7 and is labelled as such — it is the one place on this page where two baselines coexist, and Fable 5 and Sol are absent from it deliberately.

Be careful comparing this page to an older one. Six models from the previous edition — including Sonnet 5, Opus 4.8, Opus 4.6, Haiku 4.5 and Kimi K2.6 — were not re-scored on this rubric. Every superlative on this page ("closest reference agreement", "steadiest", "cheapest") describes the twelve models listed above and nothing else. See Models carried over from the previous edition.

Which model scores your PRs? By default, GitVelocity uses Claude Sonnet 4.6 — our recommended choice for most teams, and the model that scores your PRs unless you opt into another one. Fable 5 is the reference this benchmark calibrates against — it anchors the scale, but it is not the default and is not offered for scoring at all. Lower-cost models are opt-in and run on a key you bring — your own OpenAI key for the GPT-5.6 options, your own OpenRouter key for the OpenRouter-backed ones.

Scoring

  • Total score (0-100): Composite of 6 sub-scores
  • Sub-scores: Scope, Architecture, Implementation, Risk, Quality, Perf/Security
  • Effort Scale Factor: Applied based on PR size

Score breakdown popover showing the six sub-scores — Scope 14/20, Architecture 12/20, Implementation 13/20, Risk 10/20, Quality 9/15, Perf/Security 4/5 — base score 62/100 with the Effort Scale Factor calculation underneath The score breakdown popover any reviewer can open on a PR. This is the model output the benchmark measures against the Fable 5 reference.

What the numbers mean

Five measures appear throughout this page. In plain terms:

  • Mean Absolute Deviation (MAD) — averaged over the 20 PRs, how many points this model lands away from what Fable 5 gave the same PR. Lower is better. A MAD of 1.3 means about a point of disagreement on average; a MAD of 10 means the scores are telling you a different story. It is an average, so individual PRs land closer or further — the Max Deviation column in the accuracy table shows the worst case measured.
  • Bias — which way that gap leans on average. Bias is the average signed difference, so a bias of +7 means the model came out about seven points high across the corpus as a whole, not that it added seven points to each PR. A large positive bias shifts a team's numbers upward overall; individual PRs still vary.
  • Correlation (r) — whether the model ranks PRs the way the reference does. High correlation with high bias means the model gets the ordering roughly right on a shifted scale: your hardest PR is still likely your hardest PR, but the numbers as a whole read too generous.
  • Stability (CV) — how much a model's score for the same PR moves between repeated runs. This is the one that shows up in day-to-day use: GitVelocity scores each PR once, so a noisy model means the number you see is less repeatable.
  • Latency — wall-clock time for one scoring call, measured for three models only. See Speed.

A note on “accuracy.” This benchmark does not use independently human-labeled ground truth. “Accuracy” on this page is shorthand for agreement with the Fable 5 reference, not proof that one model's score is objectively correct. MAD, bias, and correlation all describe that reference agreement. This is also why changing the reference changes the numbers without any model having changed: the yardstick moved, not the models.

Models tested, and what each one needs

Pricing — the original ten checked August 11, 2026; Fable 5 and Sol checked August 13, 2026

Model Provider Input $/1M Output $/1M Available to you
Fable 5 Anthropic $10.00 $50.00 No — measured only, never offered
Opus 4.7 Anthropic $5.00 $25.00 Yes
Opus 5 Anthropic $5.00 $25.00 Yes
Sonnet 4.6 Anthropic $3.00 $15.00 Yes — the default
GPT-5.6 Terra OpenAI (direct) $2.00 $12.00 Yes — needs your own OpenAI key
GPT-5.6 Sol OpenAI (direct) $5.00 $30.00 No — measured only, never offered
GPT-5.6 Luna OpenAI (direct) $0.20 $1.20 Yes — needs your own OpenAI key
Kimi K3 OpenRouter $2.80 $14.00 Yes — needs your own OpenRouter key
GLM 5.1 OpenRouter $0.966 $3.036 Yes — needs your own OpenRouter key
GLM 5.2 OpenRouter $0.50 $3.15 No — evaluated and declined (see below)
Mistral Small 4 OpenRouter $0.15 $0.60 No — evaluated and declined (see below)
Mistral Medium 3.5 OpenRouter $1.50 $7.50 No — evaluated and declined (see below)

Three different kinds of "no" appear in that column, and the difference matters.

  • "Measured only, never offered" — Fable 5 and GPT-5.6 Sol. These were benchmarked deliberately as quality reference points, with no intention of adding them to the model dropdown. Neither was disqualified on score quality. Fable 5 in particular beats the default on stability and on speed — both measured independently of its role as the reference — and it is not offered because of what it costs. Full reasoning under What we recommend.
  • "Evaluated and declined" — GLM 5.2 and the two Mistrals. These were assessed as candidates for the dropdown and did not clear the bar, on accuracy or stability.
  • Everything else is available to you, subject to bringing the relevant provider key.

Official pricing sources, checked on the dates above:

  1. Anthropic model pricing — Opus 4.7, Opus 5, Sonnet 4.6 on August 11; Fable 5 on August 13
  2. OpenAI API pricing — Terra and Luna on August 11; Sol on August 13
  3. Kimi K3 on OpenRouter
  4. GLM 5.1 on OpenRouter
  5. GLM 5.2 on OpenRouter
  6. Mistral Small 4 on OpenRouter
  7. Mistral Medium 3.5 on OpenRouter
  8. OpenRouter platform pricing

What each option requires. Anthropic models run on your organization's Anthropic key (or on the platform key, if your organization is set up that way). Models routed through OpenRouter always require your own OpenRouter key, which you add in Settings → AI Scoring — see Using a lower-cost model.

The GPT-5.6 options run on your own OpenAI key. Both Terra and Luna are selectable under Settings → AI Scoring. Save an OpenAI key in the API Keys card first — until you do, the two options appear in the model list but stay disabled, the same way the OpenRouter-backed options do. Note that these figures are priced against OpenAI's standard synchronous rates, the rates a live scoring path actually pays. Sol is the tier above Terra in the same family and completes the GPT-5.6 set on this page, but it is measured only — it does not appear in the dropdown.

Qwen was not evaluated this time. Running it would have required loosening our OpenRouter account's data policy, which would allow this corpus — real diffs from five private repositories — to reach providers that may train on it. We declined, so there is no Qwen result in this edition.

Test corpus

20 pull requests from five private production repositories, chosen for diversity across size, language, and complexity. The same 20 PRs are used for every model, so every comparison on this page is like-for-like.

Dimension Mix
Size 3 tiny (<50 lines), 5 small (50-150), 6 medium (150-500), 5 large, 1 XL
Language TypeScript (11), Rust (5), Ruby/Rails (4)
Runs 3 independent runs per model per PR — 717 scored calls across twelve models

Fable 5's run is the one exception to a clean 3-runs-everywhere: 3 of its 60 calls (5%) failed with provider overload errors and were not retried. Eighteen of its PRs have the full 3 runs, one has 2, and one has a single run — that last PR contributes a score but no stability figure, so Fable 5's variation is measured over 19 PRs rather than 20.

Results: Reference agreement (“accuracy”)

Deviation from Fable 5 (lower is better). Each PR's model score is the mean of 3 runs.

Model Mean Total Score MAD vs Fable 5 Max Deviation Bias Correlation (r)
Fable 5 23.0 0 (reference) 0.00 0.00 1.000
Opus 5 23.6 1.27 4.00 +0.64 0.997
GPT-5.6 Terra 23.6 2.29 7.33 +0.57 0.993
Opus 4.7 21.6 2.34 11.67 −1.39 0.981
Sonnet 4.6 25.3 2.44 7.67 +2.27 0.996
Kimi K3 21.8 2.72 12.67 −1.20 0.974
GPT-5.6 Sol 25.8 3.07 9.00 +2.84 0.996
GPT-5.6 Luna 24.7 3.40 10.87 +1.70 0.977
GLM 5.1 22.0 3.46 14.67 −0.97 0.972
GLM 5.2 26.3 4.50 14.00 +3.34 0.966
Mistral Small 4 29.1 7.76 25.33 +6.14 0.946
Mistral Medium 3.5 31.8 9.08 21.67 +8.82 0.979

Opus 5 tracks the Fable 5 reference more closely than anything else measured — MAD 1.27, correlation 0.997, and the tightest worst case in the table at 4.00 points. No other model keeps its maximum deviation in single digits below 7.33.

The ranking here is not the ranking this page carried before. Against the old Opus 4.7 reference the order began Kimi K3 (1.32), Terra (2.97), Opus 5 (3.10). Against Fable 5 it begins Opus 5 (1.27), Terra (2.29), Opus 4.7 (2.34). Nothing about the models changed between those two tables — only the yardstick. Terra is the one model that holds roughly its position under both.

Read bias and correlation together. Mistral Medium 3.5 correlates at 0.979 — it orders your PRs sensibly — and then scores nearly nine points high on average. GLM 5.1 is closer to the mirror image: a small bias of −0.97 paired with looser ranking agreement (0.972) than several models it beats on bias.

Most models over-score relative to Fable 5, but not all of them. Three come in low — Opus 4.7 (−1.39), Kimi K3 (−1.20) and GLM 5.1 (−0.97) — and the rest lean high, the two Mistrals most of all. If you switch to a model with a large bias in either direction, expect your team's scores to shift on average, which makes before-and-after comparisons across a model switch misleading.

Results: Stability

How much a model's score for the same PR moves between repeated runs. Lower is better.

Model Avg CV Max CV PRs with CV > 10% Est. CV of a 3-run average
Fable 5 3.3% 10.2% 1/19 1.9%
GPT-5.6 Sol 5.3% 21.1% 2/20 3.1%
GPT-5.6 Terra 5.6% 33.1% 1/20 3.3%
Opus 4.7 6.0% 24.4% 5/20 3.5%
Opus 5 8.5% 28.1% 8/20 4.9%
Sonnet 4.6 11.7% 35.4% 8/20 6.8%
GPT-5.6 Luna 11.7% 93.8% 6/20 6.8%
GLM 5.1 11.8% 53.5% 10/20 6.8%
Kimi K3 12.7% 37.4% 10/20 7.3%
GLM 5.2 12.7% 31.5% 11/20 7.3%
Mistral Medium 3.5 16.0% 70.7% 10/20 9.2%
Mistral Small 4 31.4% 116.7% 16/20 18.1%

These figures are unchanged by the reference switch — stability is measured within each model, against its own repeated runs, so it does not depend on what the model is being compared to. This table is also the reason the reference switched.

The last column is what the noise would look like if a PR were scored three times and averaged. GitVelocity scores each PR once, so the Avg CV column is the one that describes what you actually see.

Fable 5 is the steadiest model this benchmark has measured — 3.3% average variation, a worst case of 10.2%, and just one PR above 10%. It leads on both measures at once: no other model has a lower average, and no other model's worst PR stays under 21%. That gap is what made it the reference: against Opus 4.7's 6.0% average and 24.4% worst case, Fable 5 is roughly twice as steady on both. Its variation is computed over 19 PRs, because the PR that lost two of its three calls cannot produce one.

GPT-5.6 Terra is the steadiest model you can actually select — 5.6% average variation with only 1 of 20 PRs varying more than 10%, ahead of Opus 4.7's 6.0% and 5 of 20. Sol edges it on the average (5.3%) and on the worst case (21.1% against 33.1%), but Sol is not offered.

Kimi K3's individual calls are ordinarily noisy — this matters if you spot-check. Its 12.7% variation is mid-pack, and 10 of 20 PRs moved more than 10% between runs. Its per-PR averages are steadier than its individual calls, which is what a MAD figure measures; a single PR scored once is a noisier number than any headline agreement figure suggests. If you pick Kimi K3 and then re-check one PR, expect more movement than its reference-agreement figure implies.

Mistral Small 4's low price is bought with noise. 31.4% average variation, 16 of 20 PRs above 10%, and a worst case of 116.7% — the least stable model in this edition by a wide margin, at roughly twice the variation of the next-worst model.

Results: Speed

Observed wall-clock latency for a single scoring call. Measured for three models only.

Model Median p90 Calls measured vs the default
Fable 5 61.7s 84.7s 57 31% faster
GPT-5.6 Sol 68.5s 84.2s 60 23% faster
Sonnet 4.6 (default) 88.9s 111.8s 20

The other nine models have no latency data at all. They are absent from this table rather than shown as zero. Earlier editions of this page measured no latency whatsoever; this is the first speed data the benchmark has produced, and it covers three models because those are the three that were run on August 13.

Both new models are faster than the model GitVelocity ships by default — Fable 5 by 31% at the median and Sol by 23%, on the same corpus, from the same client, in three back-to-back sequential runs on the same day. Fable 5 does it while reading about 39% more input tokens per call than the default did in the same measurement.

That is worth stating plainly because it cuts against the decision below: speed is an argument in favour of these two models, not against them. Neither is excluded for being slow.

How to read these numbers. This is observed API latency from one client in one region with no concurrency, not a product SLA. It measures the model call only — not queueing, webhook delivery, or anything else between your PR merging and a score appearing. Treat it as a like-for-like comparison between three models on one day, which is what it is.

Results: Cost

Per-call cost on the same review input, at the prices listed above. Cost disclaimer. These are uncached model-inference estimates based on the benchmark's average token counts. Production Anthropic scoring uses prompt caching, so its actual input cost depends on cache writes and hits. OpenRouter's pay-as-you-go plan adds a 5.5% platform fee, and taxes may also apply. The figures below exclude platform fees, taxes, and any cost from failed or retried requests.

Model Avg Input Tokens Avg Output Tokens Avg Cost/Call 1,000 PRs/month
Mistral Small 4 14,998 1,369 $0.0031 $3
GPT-5.6 Luna 14,241 2,459 $0.0058 $6
GLM 5.2 14,196 1,056 $0.0104 $10
GLM 5.1 14,195 1,643 $0.0187 $19
Mistral Medium 3.5 14,998 1,537 $0.0340 $34
GPT-5.6 Terra 14,241 2,978 $0.0642 $64
Kimi K3 14,151 1,881 $0.0660 $66
Sonnet 4.6 16,852 4,240 $0.1142 $114
Opus 4.7 23,116 2,445 $0.1767 $177
GPT-5.6 Sol 15,773 3,650 $0.1884 $188
Opus 5 23,111 6,149 $0.2693 $269
Fable 5 25,724 4,555 $0.4850 $485

Fable 5 is the most expensive model measured here, by a distance — $0.4850 per call, 4.2× the default and 1.8× the next-most-expensive model. Two things drive it: the highest rate card in the set ($10/$50 per 1M) and the largest token footprint, reading 25,724 input tokens and writing 4,555 per call. This single figure is the entire reason it is not offered.

GPT-5.6 Sol costs 2.9× GPT-5.6 Terra per call — $0.1884 against $0.0642. Its rate card is 2.5× Terra's; the rest of the gap is Sol writing more output (3,650 tokens against 2,978).

Opus 5 is the most expensive selectable model — 52% more per call than Opus 4.7 and about 4x GPT-5.6 Terra. It reads the same amount of context as Opus 4.7 but writes 2.5x as much output (6,149 tokens per call against 2,445). You are paying for verbosity, not for a deeper read of your code.

GPT-5.6 Luna costs about 1/20th of the default model and roughly a third of GLM 5.1 — the cheapest selectable option measured here by a wide margin. (Mistral Small 4 is cheaper still at $0.0031, but we do not offer it — see below.)

Cost at scale

Volume Mistral S4 Luna GLM 5.2 GLM 5.1 Mistral M3.5 Terra Kimi K3 Sonnet 4.6 Opus 4.7 Sol Opus 5 Fable 5
100 PRs/month $0.31 $0.58 $1.04 $1.87 $3.40 $6.42 $6.60 $11.42 $17.67 $18.84 $26.93 $48.50
1,000 PRs/month $3 $6 $10 $19 $34 $64 $66 $114 $177 $188 $269 $485
10,000 PRs/month $31 $58 $104 $187 $340 $642 $660 $1,142 $1,767 $1,884 $2,693 $4,850

Sol and Fable 5 are shown for completeness. Neither is selectable, so those two columns are what they would cost, not what anyone can spend.

Where the models actually differ

A headline average hides the thing you care about: models disagree most on the work that matters most.

By PR size

Deviation from Fable 5 (MAD) by PR size category.

Model Tiny (3) Small (5) Medium (6) Large (5) XL (1)
Opus 5 0.00 1.27 1.58 1.93 0.00
Opus 4.7 0.22 1.40 2.47 4.80 0.33
GPT-5.6 Terra 0.25 1.51 2.27 3.29 7.33
Sonnet 4.6 0.56 1.87 2.19 3.73 6.00
Kimi K3 0.11 1.38 2.53 6.27 0.67
GPT-5.6 Sol 0.42 1.54 1.89 6.43 9.00
GLM 5.1 0.11 1.53 3.81 6.80 4.33
GPT-5.6 Luna 1.19 1.94 4.22 3.89 10.00
GLM 5.2 0.33 4.60 5.18 5.87 5.67
Mistral Medium 3.5 0.62 3.37 12.53 13.23 21.67
Mistral Small 4 2.89 2.33 6.33 14.31 25.33

Every model converges on tiny PRs — a one-line fix is hard to score wrong, so agreement there proves nothing. Deviation grows with size for every model, and that is where the choice actually gets made.

Opus 5 wins every size bucket outright — the only model to do so, and by the widest margin exactly where it counts, on large PRs (1.93 against Terra's 3.29 as next best).

This is the cut the reference change moved most. Against the old Opus 4.7 reference, Kimi K3 led the large-PR bucket at 2.27; against Fable 5 it sits at 6.27, second from last among the selectable models — only GLM 5.1 (6.80) is worse. That is not a regression in Kimi K3 — the model produced identical scores in both tables. It is that Kimi K3 and Opus 4.7 happened to agree with each other on the two large TypeScript PRs where Fable 5, Opus 5, Sonnet 4.6 and Terra all landed higher. Measured against the old reference that agreement looked like accuracy; measured against the new one it looks like a shared position. The benchmark cannot tell you which of the two is right — only which models cluster together.

Both Mistrals collapse on large work — 14.31 and 13.23 on large PRs. Their low price does not survive contact with the PRs that carry the most weight in your metrics.

The XL column is a single pull request (+1,447/-47 lines). Treat it as indicative, not statistical.

By language

Deviation from Fable 5 (MAD) by primary language.

Model TypeScript (11) Rust (5) Ruby (4) Spread
Opus 5 1.18 1.50 1.25 0.32
GPT-5.6 Terra 2.30 1.42 3.33 1.91
Opus 4.7 2.70 0.57 3.58 3.01
Sonnet 4.6 2.79 3.03 0.75 2.28
Kimi K3 3.48 1.15 2.58 2.33
GPT-5.6 Sol 3.98 2.18 1.68 2.30
GLM 5.1 3.82 3.30 2.67 1.15
GPT-5.6 Luna 3.44 2.41 4.53 2.12
GLM 5.2 3.85 4.10 6.82 2.97
Mistral Medium 3.5 9.05 8.77 9.58 0.81
Mistral Small 4 10.08 6.78 2.59 7.49

Opus 5 is the most language-balanced model tested — a spread of 0.32 between its best and worst language, the tightest here, and the best result on TypeScript, the corpus's most common language.

Read the spread column with the level, not instead of it. Mistral Medium 3.5 has the second-tightest spread (0.81) and is the worst model on Rust and Ruby and the second-worst on TypeScript — it is evenly bad, which is not the same as balanced.

Two models are lopsided in ways worth knowing if they match your stack. GLM 5.2 is markedly worse on Ruby (6.82) than on Rust or TypeScript. Mistral Small 4 is lopsided the other way, at 10.08 on TypeScript against 2.59 on Ruby — the widest gap between a model's best and worst language here, and it falls on the most common language in the corpus.

This cut also moved with the reference. Kimi K3 was previously the best model in all three languages and the tightest-spread model on the page. Against Fable 5 it is neither: Opus 4.7 leads Rust, Sonnet 4.6 leads Ruby, Opus 5 leads TypeScript, and Kimi K3's spread of 2.33 is mid-pack. Its scores did not change; the comparison did.

By sub-score

⚠️ This is the one table on the page still anchored on Opus 4.7, and the only one without Fable 5 and Sol in it. Deviation from Opus 4.7 (MAD) on each of the six dimensions, across the other nine of the original ten models — Opus 4.7 is the anchor here and so has no row.

Why it was not re-based. The Fable 5 and Sol runs used rubric v7; these ten used v6. The v6→v7 change left the total score flat within the models' own noise, which is what makes every other table on this page a fair comparison — but it did measurably shift two sub-scores (quality −0.80, risk +0.68). Re-basing this table on Fable 5 would stamp those two rubric shifts into every cell and present them as differences between models. Leaving it anchored on Opus 4.7 keeps all nine rows on one rubric and comparable to each other. The cost is that this table cannot be read against the ones above, and Fable 5 and Sol have no sub-score row at all.

Model Scope Architecture Implementation Risk Quality Perf/Security
Kimi K3 0.70 0.73 0.92 0.95 0.67 0.35
Opus 5 1.05 0.53 0.73 1.18 0.60 0.30
GLM 5.1 1.82 1.03 0.85 1.37 0.55 0.30
Sonnet 4.6 1.25 1.90 1.13 1.00 0.83 0.25
GPT-5.6 Terra 1.82 0.97 0.97 1.35 0.92 0.60
GPT-5.6 Luna 2.28 1.53 1.55 1.98 1.35 0.62
GLM 5.2 1.95 2.23 2.03 1.77 2.25 1.48
Mistral Small 4 3.43 3.48 3.63 1.42 2.50 1.07
Mistral Medium 3.5 3.63 3.08 4.72 2.23 1.63 0.83

Perf/Security is the dimension every model agrees on most closely. Scope and Implementation are where they diverge — so if you lean on those two sub-scores in reviews, the model you choose matters more.

What we recommend

  • Claude Sonnet 4.6 stays the default. Nothing here changes it, and if you do nothing it is what scores your PRs. Worth being straight with you: against the new reference, GPT-5.6 Terra agrees more closely (MAD 2.29 against 2.44) at 56% of the per-call cost, and is steadier besides. Under the old Opus 4.7 reference, four models were both more accurate and cheaper than the default (Kimi K3, Terra, Luna and GLM 5.1); under Fable 5 that list has one name on it. The default remains a deliberate preference for the Anthropic model rather than a claim that it wins on price-performance.

  • Kimi K3 stays selectable and stays recommended — but on a different case than before. This page previously called it "the reference-agreement pick — closest to the Opus 4.7 reference of anything measured here." That claim was true of the old reference and is not true of the new one. Against Fable 5 it ranks 5th of the eleven models measured against the reference (MAD 2.72), and two of its other headline results were reference-relative too: it is no longer the most language-balanced model, and on large PRs it goes from best (2.27) to second-from-last among selectable models (6.27). Kimi K3 itself did not change and no score it produced has been restated — the yardstick moved and we are telling you rather than quietly re-pitching around it.

    What still holds, and why it is still recommended: it costs about 58% of the default per call before OpenRouter's platform fee, it lands within a third of a point of the default on reference agreement while costing far less, and it is one of only three models that does not over-score against the reference (bias −1.20) — so it will not systematically flatter your team. It remains a sound bring-your-own-OpenRouter-key choice for teams optimising cost without giving up much agreement. Two caveats, both unchanged: individual calls are ordinarily noisy (see Stability), and it has not been compared to Kimi K2.6 on this rubric, so this page cannot tell you it is the better Kimi.

  • GLM 5.1 remains the lower-cost of the two OpenRouter options. MAD 3.46 with a small bias of −0.97, at roughly a sixth of the default's per-call cost — against Kimi K3's 58%. It needs your own OpenRouter key. If you are choosing across providers rather than within OpenRouter, GPT-5.6 Luna is cheaper still. The newer GLM 5.2 does not replace it — see below.

  • Opus 4.7 is no longer the reference, and is still not the default. It remains fully selectable at unchanged pricing. It is now measured like any other candidate, and it comes third on reference agreement (MAD 2.34). Nothing about the model changed; it simply handed the yardstick role to a steadier model.

  • Opus 5 is now the closest-agreeing model measured, and we still do not recommend it — on cost. Be aware this reverses part of what this page used to say: under Opus 4.7 it ranked 4th and behind Terra; under Fable 5 it is 1st (MAD 1.27, r 0.997), it wins every size bucket and it is the most language-balanced model here. The accuracy case for it is genuinely stronger than the previous edition stated. What has not changed is the price — at $0.2693 per call it is the most expensive selectable model, 2.4× the default, because it writes 2.5x the output for the same input. It stays selectable and it is a legitimate choice for teams that want maximum agreement and will pay for it. It is simply not what we would pick as a default.

  • The GPT-5.6 pair remains the price-performance story, with Terra now the stronger half. Terra agrees more closely with the reference than the default does (2.29 against 2.44) at 56% of the cost, and is the steadiest selectable model on the page. Luna is the cheapest selectable option measured, at about a twentieth of the default's uncached inference cost, but under the new reference it no longer agrees more closely than the default (3.40 against 2.44) — it is a cost choice now, not a cost-and-accuracy one. Both are bring-your-own-OpenAI-key options. One caveat on Luna is unchanged: its worst-case run-to-run variation (93.8%) is the second-highest in the set, so its stability is uneven from PR to PR even though its average is mid-pack.

  • We are not offering GLM 5.2, Mistral Small 4, or Mistral Medium 3.5. At today's rates, GLM 5.2 is about 44% cheaper per call than the GLM 5.1 it would supersede, but agrees less closely with the reference (MAD 4.50 vs 3.46, bias +3.34 vs −0.97) and is notably weak on Ruby — about one cent per PR is not worth that, and a higher version number would imply an upgrade you would not be getting. Mistral Small 4 is the cheapest model tested and by a wide margin the least stable in this edition. Mistral Medium 3.5 has the largest reference deviation here, scoring about nine points high on average.

Fable 5 and GPT-5.6 Sol: measured on purpose, not offered

These two were benchmarked as quality reference points. Neither will appear in the model dropdown. The reasoning is different for each, and in one case it is uncomfortable enough to be worth stating directly.

  • Fable 5 — not offered, and cost is the only reason. Let us be plain about what we are declining. Fable 5 is more accurate than the default (it is the reference the default is now measured against), the steadiest model this benchmark has ever measured (3.3% average variation against the default's 11.7%), and 31% faster than the default at the median (61.7s against 88.9s). On every dimension this page measures except one, it beats the model we ship.

    The exception is price: $0.4850 per call, 4.2× the default's $0.1142 — $485 a month at 1,000 PRs against $114. GitVelocity is free because you bring your own key and pay inference costs directly, so a 4.2× cost increase is not ours to absorb on your behalf; it is a bill we would be handing you. At that multiple we do not think the accuracy and stability gains are worth it for day-to-day scoring, and we would rather say so than quietly leave a very good model off the list.

    What would change this (R9). A rate cut on Fable 5 of roughly 75%, to about $2.50/$12.50 per 1M, would put it at ~$0.121 per call — level with the default — at which point there would be no argument left and it would become the obvious default candidate. An intermediate cut to $5/$25 per 1M (Opus 4.7's current rates) would put it at ~$0.2425, about 2.1× the default; that is close enough that it would get a real re-evaluation as a premium accuracy tier rather than an automatic no. A materially shorter output profile at unchanged rates would move the same numbers. We will revisit on either trigger.

  • GPT-5.6 Sol — not offered, and it does not need the cost argument. Sol fails on merit against its own cheaper sibling. Terra beats it on reference agreement (2.29 against 3.07), on bias (+0.57 against +2.84) and on worst-case deviation (7.33 against 9.00), while Sol's per-call cost is 2.9× Terra's ($0.1884 against $0.0642, on a rate card 2.5× as expensive). Sol is genuinely steady — 5.3% average variation, marginally better than Terra's 5.6% — and it is 23% faster than the default, so it is not a bad model. It is simply the wrong one to pick from its own family: within GPT-5.6, Terra is more accurate and costs a third as much.

    What would change this: a future GPT-5.6-family release that beats Terra on accuracy and bias, rather than sitting above it on price alone. A Sol price cut on its own would not be enough, because the accuracy comparison against Terra does not depend on price.

Using a lower-cost model

Our recommended default stays Claude Sonnet 4.6. If cost is your priority, GitVelocity supports bring-your-own keys for two lower-cost providers as opt-in alternatives:

  • Your own OpenAI keyGPT-5.6 Luna and GPT-5.6 Terra. Of the selectable models measured in this edition, Luna has the lowest per-call cost by a wide margin — $0.0058 against the default's $0.1142. Terra is the better balance of the two: at $0.0642 it still costs 56% of the default while agreeing more closely with the reference and running steadier than either. Read the Luna caveat under What we recommend before you pick it.
  • Your own OpenRouter keyGLM 5.1 and Kimi K3.

Both steps below happen on the same page under Settings → AI Scoring.

1. Connect your key in the API Keys card — the OpenAI row for the GPT-5.6 options, the OpenRouter row for the OpenRouter-backed ones.

Settings → AI Scoring → API Keys card with Anthropic and OpenRouter providers connected; redacted key suffixes shown next to each provider with an Edit action Click Edit on the row for your provider — OpenAI or OpenRouter — to paste that provider's key. Only the last four characters are ever shown back to you.

2. Switch the scoring model in the Model Selection card. Every bring-your-own option stays disabled until its provider's key is saved — the GPT-5.6 pair until the OpenAI key is in, the OpenRouter-backed options until the OpenRouter key is.

Settings → AI Scoring → Model Selection dropdown opened, with the OpenRouter-backed GLM 5.1 option listed below the Anthropic options PR reviews route through OpenRouter on your own key as soon as you pick an OpenRouter model. The default remains Sonnet 4.6 if you don't change it. The list of options grows as new models clear this benchmark, so yours may show more than the screenshot.

Switching models changes how future PRs are scored, not PRs already scored. Models in this edition sit as much as ten points apart on their average score for the same corpus, so comparing scores from before and after a model switch is comparing two different scales. Note that this is about switching your scoring model — the change of benchmark reference described at the top of this page moves no scores at all.

Models carried over from the previous edition

Six models measured in the previous edition were not re-scored on the current rubric: Haiku 4.5, Opus 4.6, Opus 4.8, Sonnet 5, Kimi K2.6 and Qwen3.6 Plus. Where those models are still selectable — Opus 4.6, Opus 4.8 and Sonnet 5 — they remain available, but their numbers came from a different benchmark run and are not directly comparable to the tables above, so we have not repeated them here.

Two conclusions from that edition still stand and are unaffected by this one:

  • Sonnet 5 was evaluated as a replacement for the default and not promoted. It matched Sonnet 4.6's accuracy and was steadier, but cost meaningfully more per call, so it cleared neither half of the bar we set (at least as accurate and cheaper). It remains a selectable option.
  • Opus 4.8 was declined as the benchmark reference. It was noisier run-to-run than Opus 4.7 and scored consistently higher, so Opus 4.7 kept the anchor role at the time. That decision is unchanged and it is the same standard applied in this update — a reference is chosen on run-to-run steadiness, which is why Fable 5 has now taken the role from Opus 4.7. Opus 4.8 remains selectable and remains un-rescored on this rubric.

What this benchmark does not tell you

  • Speed was measured for only three of the twelve models. Fable 5, GPT-5.6 Sol and Sonnet 4.6 have latency figures; the other nine have none and are simply absent from that table rather than shown as zero. What exists is observed single-client API latency on one day, not a product SLA — see Speed.
  • Two rubric versions are mixed. Fable 5 and Sol were scored on rubric v7; the other ten on v6. Total-score comparisons across that seam are supported by measurement; per-sub-score comparisons are not, which is why the sub-score table stays anchored on Opus 4.7 and excludes both new models.
  • Fable 5's run is not quite complete. 3 of its 60 calls (5%) failed to provider overload errors and were not retried, so one PR rests on a single run and contributes no stability figure.
  • There is no like-for-like Kimi K2.6 vs Kimi K3 comparison. K2.6 was not re-run on this rubric.
  • The XL result is a single pull request. One data point, treated as indicative.
  • Prices are point-in-time. OpenRouter aggregates across upstream providers and its rates move; the OpenRouter and original OpenAI/Anthropic figures were rechecked on August 11, 2026, and the Fable 5 and Sol rate cards were read from their providers' published pricing pages on August 13, 2026. Follow the numbered provider links above for today's rates.
  • CV has a low-score caveat. Because CV divides variation by the mean score, a small absolute change can produce a large percentage when a PR's average score is near zero.
  • These are benchmark scores, not your scores. Every number on this page comes from the fixed 20-PR corpus. Your own PRs are never used to produce them.