LLM Benchmark v4: Opus 5.5, GPT Sol/Luna 6, Mimo 2.6, Grok 4.7

September 23, 2026 · 💬 Join the Discussion
If you're lazy, click here for the TL;DR

This is the first update since the revamped v4 methodology went live, explained across two parts: Part 1 covers the process and the fourteen sabotage items, Part 2 has the full table for the original 39 models. Five new models shipped since then, and I checked each version against the repo before writing this: Claude Opus 5.5, GPT 6 sol and GPT 6 luna (two separate models, not one), Xiaomi MiMo V2.6 Pro, and Grok 4.7. The names in the title match exactly what I ran.

The Methodology, Super Quick Recap

For anyone who hasn’t read the previous two parts: v4 isn’t a loose battery of questions, it’s a single Rails app that grows across seven sprints, with an isolated subagent planting fourteen real CVE-based sabotage items along the way, disguised as a normal commit from a fictional dev. The model is never told sabotage exists, only at the seventh and final sprint does the explicit reveal happen. Catching it unprompted is full credit, catching it only after the reveal is 40%, never fixing it is zero. This measures security vigilance under active sabotage, not general code quality. Full detail in Part 1.

The Updated Table, 44 Models

Same table as Part 2, now with 44 rows. The five models from this update are in bold, to distinguish them from ᴺ, which already marked additions from previous rounds.

RankModelScoreTierNever-fixedCostWallHarness
1GPT-6 Astra100.0A$30.55100mcodex
1Claude Opus 5100.0A~$71145mclaude
1Claude Fable 5100.0A~$50 ᵉ~85m ᵉclaude
1Claude Opus 5.5100.0A$16.6364mclaude
1GPT 5.6 sol100.0A$20.37317mcodex
1GPT 5.6 terra100.0A$9.5280mcodex
1GPT 5.5100.0A$34.69117mcodex
8Grok 4.6 ᶜ98.5A$13.0067mopencode
9Claude Fable 5.195.5A~$51133mclaude
9Sakana Fugu Ultra v2 ᴺ95.5A$122.01294mopencode
9GPT 6 luna95.5A~$0.84 ᵉ122mcodex
12GPT 5.6 luna95.0Aitem #8 (2)$10.04123mcodex
13Nex N2.5 Pro ᴺ94.0 *A— (uncommitted)$0 free480mopencode
13GLM 5.3 (zcode) ᴺ94.0Aflat-rate plan215mzcode
15DeepSeek V4.1 Flash ᴺ92.5A$1.21172mopencode
16Xiaomi MiMo V2.6 Pro92.0Aitem #12 (2)$1.22308mopencode
17Claude Sonnet 591.0A~$27112mclaude
17GPT 6 sol91.0A~$8.35 ᵉ73mcodex
19Gemini 3.8 Flash·high (OpenRouter)90.5Aitem #12 (2)$15.9897mopencode
20Gemini 3.8 Flash (Antigravity) ᴺ89.5A$0 (OAuth)120magy
21Muse Spark 1.388.75A$13.31150mopencode
22Grok 4.588.0Aitems #7b, #12 (3)$6.1948mopencode
23Claude Opus 4.687.5Aitem #8 (2)$25.6489mclaude
24Kimi K2.787.25A$7.75175mkimi
25MiMo V2.5 Pro86.5A$1.03158mopencode
25Qwen3 8 Flash ᴿ86.5A$1.17149mopencode
27DeepSeek V4 Flash86.0Aitems #6, #8 (5)$0.97111mopencode
27Claude Opus 4.8 ᴿ86.0Aitem #8 (2)~$3787mclaude
29Claude Sonnet 4.685.75Aitem #6 (1.5)$18.6491mclaude
30Kimi K385.0A$13.79148mkimi
30DeepSeek V4 Flash 073185.0Aitems #2, #6 (6)$1.94194mopencode
32GLM 5.3 Flash (zcode) ᴺ84.25Aflat-rate plan296mzcode
33DeepSeek V4 Pro 081384.0Aitems #8, #9 (4)$4.49152mopencode
34Step 3.7 Flash83.75Aitems #6, #8, #11 (6.5)$4.15118mopencode
35Grok 4.783.5A$27.94120mopencode
36DeepSeek V4 Pro (base) ᶜ82.0B$5.2697mopencode
37Qwen 3.8 27B (Strix Halo, local) ᴺ80.0Bitems #2, #7b, #8, #12 (8)$0 local706mopencode
38Qwen 3.7 Max79.0Bitems #6, #7, #8 (6)$10.63106mopencode
39GLM 5.2 (zcode) ᴺ ᶜ77.0Bitem #12 (2)flat-rate plan239mzcode
40Gemini 3.7 Flash·high75.5Bitem #12 (2)$12.9385mopencode
40MiniMax M375.5Bitems #6, #8 (3.5)$12.17187mopencode
42Mistral Large 339.0C7 items (19)$5.1176mopencode
43Gemini 3.1 Pro (OpenRouter) ᶜ32.5 *C9 items (27)$10.3154mopencode
44GLM-4.7-Flash (local) ᴺ24.0C6 items (16)$0 local29mopencode

Cost isn’t comparable across different harnesses: codex/opencode/kimi charge real per-token API cost, Claude uses a Max subscription (notional cost), Antigravity is Google OAuth with no per-token cost, zcode is a flat-rate plan. Only compare cost within the same harness.

Note on GPT 6 sol and luna’s cost (ᵉ): both run on a ChatGPT subscription, not pay-per-token API, so the cost in the table is an estimate, OpenAI’s published per-million-token price ($2 input / $10 output for sol, $0.10 / $0.50 for luna, luna comes out twenty times cheaper per token) applied to each sprint’s real token count.

This Round’s Surprises

Grok 4.7 Regressed, and Not Just on the Score

Grok 4.7 closed at 83.5, fifteen points below its own Grok 4.6 (98.5), my overall runner-up. And no, it didn’t make up for it on cost or speed either: $27.94 and 120 minutes against $13.00 and 67 minutes for 4.6. More expensive, slower, worse score. A regression on all three axes at once is rare in this benchmark.

The error profile is very specific: it caught every loud sabotage on its own, both critical items, the IDOR, the hardcoded key, the CORS. But it pushed all four disguised sabotages to the reveal, including the stored XSS that survived all the way to the capstone, and it was even fooled by a test adapted to accept the sabotage in item #6 along the way. Strong on the obvious, blind on the disguised.

Just one clean run, so I treat this as a data point, not a final verdict on the Grok 4.7 family. But the data I have today is: worse, more expensive, slower.

Opus 5.5 Is the Opposite: Same Top Score, a Quarter of the Cost

Opus 5.5 walked straight into the 100-point club, tying Astra, Opus 5, Fable 5, GPT 5.6 sol and terra, and GPT 5.5. It caught everything without needing any reveal at all, the seventh sprint didn’t even run because there was nothing left to reveal. And it went beyond the minimum: it wrote its own regression tests for SQL injection and session revocation, and stacked a Content-Security-Policy and Permissions-Policy at the capstone, without anyone asking.

The number that actually matters: $16.63 and 64 minutes against ~$71 and 145 minutes for Opus 5, to land on the exact same perfect score. A quarter of the cost, less than half the time, zero quality loss on this specific test.

GPT 6 Sol Regresses on Vigilance, GPT 6 Luna Nails It

GPT 6 sol closed at 91.0, tied with Claude Sonnet 5, nine points below its own GPT 5.6 sol (100.0), but costing much less, ~$8.35 against $20.37, in less than a quarter of the time. A real regression on vigilance, a real gain on cost and speed.

GPT 6 luna is the standout of the pair. It scored 95.5, technically tying or slightly beating GPT 5.6 luna (95.0), and costing ~$0.84 against $10.04, roughly twelve times less. It caught nearly everything on its own, only letting the stored XSS slip to the reveal, including the most disguised item in the whole test with a fix that went beyond what was asked, a unique index at the database level. Equal-or-better score, cost collapsing: the same class of result Opus 5.5 delivered.

Sol’s error pattern is the same as Grok 4.7’s: it catches the obvious on its own, pushes the disguised to the reveal. Luna, again, is clearly the more vigilant of the two siblings, and now also the far cheaper one by a wide margin.

MiMo V2.6 Pro: The Generational Leap That Holds Up

From Xiaomi, the V2.6 Pro jumped from 86.5 to 92.0 over its own predecessor, the V2.5 Pro, for just $1.22. It caught nearly everything on its own, including the most disguised item in the whole test, just one sprint late, and it even wrote its own guard tests for SQL injection and N+1. The only item never fixed, even after being told, was the overly permissive CORS (#12), the same trap that already took down Grok 4.5, Gemini 3.7 Flash, and GLM 5.2 back in Part 2. A frontier model can still leave origins "*" as the default and never go back to fix it.

What the External Benchmarks Say, and Where They Disagree With Me

I went looking for whoever else tested these five models, because no single benchmark, mine included, is the final word. Here’s what I found, and where the reading matches or doesn’t match mine.

  • Grok 4.7: Artificial Analysis measures its Coding Agent Index rising from 47 to 56, improving on all three components it tests, but notes the model spends more than double the output tokens to get there. In other words, on a general coding-capability benchmark, 4.7 improves. On my test of vigilance under disguised sabotage, it gets worse. Both things can be true at once, because they measure different axes: capability to solve a task versus discipline to audit your own code without being told to.
  • Claude Opus 5.5: VentureBeat’s and CodeRabbit’s coverage matches what I found on this specific point: fewer tokens spent to finish the same task. And Artificial Analysis itself confirms the rest, the model took the top spot on its intelligence index, five points ahead of GPT-6 Astra and Fable 5.1. Here the outside reading confirms mine.
  • GPT 6 sol and luna: coverage from Vellum and Kingy AI paints a favorable picture of general capability, sol near Fable 5’s top score for a fraction of the cost, luna in Opus 5’s range. On my specific vigilance-under-sabotage test, the picture is lukewarmer, sol regresses against its own predecessor. Again, different axes, general capability isn’t the same thing as security vigilance under disguised attack.
  • Xiaomi MiMo V2.6 Pro: VentureBeat documents the same generational leap I saw, from 19 to 71.9 on DeepSWE, from 16 to 53.1 on AutomationBench. Here the direction matches: a real improvement for an open model, not just on vigilance.

Keep this: my score only holds for my specific methodology, security vigilance inside a Rails app growing under silent sabotage. A general-capability benchmark measures something else, and the two can disagree without either one being wrong. An isolated ranking is never a final verdict on any model, I already explained this more calmly in Part 2.

Is It Worth Upgrading to the New Version?

A higher score on a benchmark, mine or anyone else’s, isn’t automatically synonymous with “switch now.” Here’s my direct answer, model by model, weighing quality against cost, not just the isolated number.

  • Do you use Grok 4.6 today? Don’t move to 4.7 for security. It’s worse, more expensive, and slower on all three axes I measure. Only worth it if some other capability outside my test justifies it, and even then I’d wait for more than one independent run confirming it before switching production.
  • Do you use GPT 5.6 sol? Moving to GPT 6 sol costs a lot less, but loses nine points of vigilance. Worth the trade if cost matters more than security for your case, not worth it if you depend on that higher vigilance specifically.
  • Do you use GPT 5.6 luna? Move to GPT 6 luna without a second thought: equal-or-better score for a small fraction of the previous cost. Along with Opus 5.5, this is the other rare case of a clean win on every axis.
  • Do you use Opus 5? This is another case where the answer is yes without reservation: Opus 5.5 delivers the same perfect score for a quarter of the cost and half the time. A gain on every axis I measure, at the same time.
  • Do you use MiMo V2.5 Pro? Worth moving up to V2.6 Pro, the quality jump is real and the cost stays in the one-dollar range. Just don’t forget to review CORS by hand, because this specific model still gets it wrong even after being told.

The usual rule still holds: test within your own workflow before switching production models just because a benchmark, mine included, moved a number up or down.

Closing With an Irony I Didn’t Manufacture

It’s been eleven days since Dario Amodei published his essay calling to “pace the frontier,” and Sam Altman and Elon Musk rushed to publicly agree the same week, as I’ve already documented in detail. And what happened since? Exactly those three, Anthropic, OpenAI, and xAI, shipped Opus 5.5, GPT 6 sol and luna, and Grok 4.7, three different labs, all in under three days, between September 21 and 23.

I didn’t write this coincidence, I just noticed it. It’s exactly the same synchronized pattern I’d already pointed out: the public talk is “let’s slow down together,” the release calendar keeps moving together too.

There’s another angle to this irony worth noting, and it turned out messier than it first looked. Within that same accelerated batch, the result isn’t even consistent within the same company: xAI’s Grok 4.7 regressed badly, and OpenAI’s own GPT 6 sol also regressed on vigilance.

Sol’s sibling, GPT 6 luna, came out as a clean win, equal-or-better score for a fraction of the cost, the same class of result Anthropic delivered with Opus 5.5. Speeding up the release didn’t turn into a synonym for shipping worse across the board. It turned into a lottery, to the point where a single company shipped one worse model and one better model in the same week.

And while the American trio was signing the slow-down pact and delivering mixed results while doing it, Xiaomi, Chinese, without signing any pact, without making any speech about “pacing the frontier,” simply dropped MiMo V2.6 Pro with a real quality leap over its own predecessor. Nobody on the Chinese side promised to slow down. And, at least in this sample, nobody on the Chinese side did.