Comparison Head to Head

Opus 5.5 vs Grok 4.7 vs MiMo V2.6: The 'Cheap' Model Cost Twice as Much

Three launches in 48 hours. The one with the lowest sticker price ran the biggest bill, the open-weights model matched it for a twentieth of the cost, and Opus finally got its comeback episode.

The Anthropic, xAI, and Xiaomi logos lit side by side on a dark control panel
Illustration generated for Run the Eval
The receipts
  • Claude Opus 5.5 ($4 in / $20 out) takes the top spot on the independent Artificial Analysis Intelligence Index at 58 at max effort, and costs $1.34 per task at its default setting.
  • Grok 4.7 lists at $2 / $6, a third of Opus 5.5's output price, but cost $2.73 per task on the same independent test because it burned about 2.6x the output tokens. Independent Terminal-Bench 4.0: 26%.
  • Xiaomi's MiMo V2.6 Pro is open weights under MIT, scores 46 on the index (tied with Grok 4.7), and costs $0.13 per task. Flash runs $0.14 / $0.28 with 15B active parameters.
  • The Opus 5.5 catch: most cybersecurity tasks silently reroute to Opus 4.8, and thinking can no longer be switched off.
Short answer

Claude Opus 5.5 is the strongest of the three, scoring 58 on the Artificial Analysis Intelligence Index at max effort and costing $1.34 per task at default. Grok 4.7 lists cheaper at $2/$6 per million tokens but cost $2.73 per task. MiMo V2.6 Pro, an open-weights model, matches Grok's index score of 46 for $0.13 per task.

Monday night I told the timeline I’d pulled Opus off some of my client projects. Tuesday afternoon Anthropic shipped the Opus I’d been waiting for.

Three frontier launches landed inside 48 hours. xAI shipped Grok 4.7 and Xiaomi shipped MiMo V2.6 on Monday, September 21. Anthropic shipped Claude Opus 5.5 on Tuesday. I posted about all three in real time, and I’m going to hold myself to those posts, because the receipts are the whole point of this desk.

Here’s the part the launch posts didn’t put on the first slide: the model with the cheapest sticker price ran the biggest bill. Let me show you why, then tell you which one to run for which job.

The only number that matters: cost per task

Price per token is the sticker on the window. Cost per task is the gas bill after a month of driving. Everybody quotes the sticker. The gas bill is what hits your card.

Artificial Analysis runs every model through the same ten-benchmark suite and reports what each one actually cost to finish it. Same test, same rules, nobody grading their own homework. Here’s how this week’s launches landed on version 4.3.2 of their index:

ModelSettingIntelligence IndexCost per taskOutput tokens per task
Claude Opus 5.5max effort58n/an/a
Claude Opus 5.5medium (default)51.2$1.3425.7k
Grok 4.7high (default)46.3$2.7365.9k
MiMo V2.6 Prodefault46$0.13n/a
Claude Fable 5.1for reference53n/an/a
GPT-6 Astrafor reference53n/an/a

Read that twice. Grok 4.7 costs $6 per million output tokens. Opus 5.5 costs $20. On paper Grok is a third of the price. On the actual job, Grok cost $2.73 and Opus cost $1.34, because Grok generated about 2.6 times as many output tokens to get there. Wilder still: $2.33 of Grok’s $2.73 was input. The cheap output price was never the problem. Most of the bill came from what it read, not what it wrote.

Cheap sticker. Drinks premium on every trip.

Now look at the MiMo row. MiMo V2.6 Pro lands the same index score as Grok 4.7, 46, for thirteen cents a task. Same grade. About a twentieth of the bill.

Claude Opus 5.5: the comeback episode

I’m going to quote myself, because I was loud about this.

“Claude 5.5… You outdid yourself my friend. I would cuss out Claude 5 religiously, and have even gone back to GPT because of it… Or generally have Fable instruct Opus… but now Opus finally behaves.

Opus 5 was one of @AnthropicAI’s worst models EVER.”

@MicahBerkley, September 22

That’s not a hot take I’m walking back. Opus 5 was the employee I had to supervise. Half the time I had Fable instruct Opus instead of trusting Opus on its own. “Finally behaves” is the whole review.

What you pay. Anthropic’s launch post puts it at $4 per million input tokens and $20 per million output, down from $5 and $25 on Opus 5. Cache reads fell from $0.50 to $0.20 per million. Cache writes fell from $6.25 to $5. There’s a fast mode at $8 and $40 if you need speed more than savings. Context is 1 million tokens with up to 128K of output. Anthropic claims typical workloads come out about 40% cheaper than Opus 5, because the new model also finishes in fewer tokens and writes output more than 30% faster.

What it scored. Anthropic’s own table, against Opus 5:

  • Terminal-Bench 4.0: 66.4% vs 52.3%
  • CursorBench 4.0: 57.8% vs 46.6%
  • FrontierCode: 54.4% vs 48.0%
  • AutomationBench: 40.0% vs 26.9%

Those are Anthropic’s numbers. Vendor tables are home-court refs, so I weigh them less than the independent stuff. The independent stuff agrees anyway. Artificial Analysis put Opus 5.5 at 58 at max effort, five points clear of Fable 5.1 and GPT-6 Astra, which are tied at 53. Vals AI measured it on ProgramBench at $42.66 per task against Opus 5’s $60.29. When the home ref and the away ref make the same call, it’s a real call.

Effort is a dial now. Low, medium, high, xhigh, max. Medium is the default, and the default is where the $1.34-per-task number comes from. Most people should leave it there and only crank it for the hard jobs.

The catches Anthropic didn’t put in the headline

Two things you need to know before you point an agent fleet at this.

One: cybersecurity work gets rerouted to an older model. Anthropic says most cybersecurity tasks will go to Opus 4.8 instead. Routine work like finding and fixing bugs in your own code stays on 5.5. The problem is visibility: reporting on the launch says there’s no visible sign when the switch happens. If you run a security shop, you could be getting Opus 4.8 answers on work you thought was running on 5.5, and never see the swap. Anthropic says its Cyber Verification Program extends to 5.5 in the coming weeks. Until then, test your security workflows on purpose instead of assuming.

Two: thinking is always on now. You can’t switch thinking mode off anymore. And API accounts created on or after August 31, 2026 get “preserved thinking,” an anti-distillation safeguard. If your harness was built around thinking-off responses on Opus 5, budget a day to retest it.

Neither one is a dealbreaker. Both are the kind of thing that turns into a 2am incident if nobody told you.

Grok 4.7: the highlight reel vs the game film

Here’s what I posted the night before Opus 5.5 landed:

“Grok 4.7 is still behind Muse lol. Yeah the company of Instagram and Facebook. Top 3 Models: Fable, Astro, Muse.

I wouldn’t polymarket this sector if you paid me.”

@MicahBerkley, September 21

The numbers back the joke. On the same independent index, Meta’s Muse Spark 1.3 at max effort scored 48.1 for $1.60 a task. Grok 4.7 scored 46.3 for $2.73.

What xAI actually shipped. A new, larger base model rather than a tune of Grok 4.6, plus a longer reinforcement-learning run on harder, multi-hour tasks. Same price as 4.6: $2 in, $6 out, with a faster variant at double. A 500,000-token context window, text and image in, text out, and a knowledge cutoff of May 2026. It’s live in the xAI API, in Cursor on all plans (it’s the default in Grok Build), and on OpenRouter, Vercel, and Cloudflare. xAI calls it its most capable model yet for coding and knowledge work.

Here’s the game film. xAI’s own table puts Grok 4.7 at 38.0% on Terminal-Bench 4.0. Independent testing reported by The Decoder has it at 26%, next to 60% for GPT-6 Astra and 55% for Fable 5.1. DeepSeek V4.1 Flash scored 27% on the same test. A mid-tier flash model edged the new Grok flagship on agentic coding.

That’s a free agent whose highlight reel came from his own camera crew. The reel says 38. The game film says 26.

Where Grok 4.7 claims real wins. In fairness, xAI’s table has a couple of genuine flexes: 64.0% on EEBench against 56.4% for Fable 5.1 Max, and 19.6% on Harvey’s legal agent benchmark against 6.7% for Fable 5.1 Max. On DeepSWE v1.1 it’s close to the top at 71.0% vs 72.7% for GPT-5.6 Sol Max. Those are xAI’s numbers, not independent ones, so treat them as leads to test, not verdicts. If you do legal document work, it’s worth an afternoon of your own evals. If you’re building coding agents, the independent film is the one to trust.

MiMo V2.6: the out-of-town hire who outworks the payroll

Monday night I posted this:

“Listen… Mimo-v2.6 is a beast.

Btw if you don’t have @cline your bugged out. Swapped it out for come of my client projects that were using Opus… It’s delivering.”

@MicahBerkley, September 21

Some of that was Opus 5 frustration talking. But the independent numbers back it: Artificial Analysis had already clocked MiMo V2.6 Pro as the top open-weights model on their index, out of 114 open-weight models they track. Its predecessor, V2.5 Pro, scored 26. V2.6 Pro scored 46. That’s a twenty-point jump in one version, at the same price.

The two models.

  • MiMo V2.6 Pro: 1.02 trillion total parameters, 42 billion active per token. $0.435 per million input tokens, $0.87 output, and cached input drops to $0.004. About $0.13 per index task.
  • MiMo V2.6 Flash: 309 billion total, 15 billion active. $0.14 input, $0.28 output, cached input $0.0028.
  • Pro UltraSpeed: roughly 20x Pro’s output speed for 10x the price, for when latency is the whole game.

Both are mixture-of-experts, which means only a slice of the model lights up for each token. Both take text, images, audio, and video in and write text out. Both carry a 1-million-token context window. And both are MIT-licensed open weights on Hugging Face, which means commercial use and self-hosting are allowed. You can also reach them through Xiaomi’s own API, MiMo Code, and OpenRouter. According to VentureBeat, Xiaomi says the RL training took 30 large steps covering about 750,000 trajectories in under six days, at a reported $2.62 million.

Pro vs Flash, per Xiaomi’s own table:

BenchmarkProFlash
DeepSWE v1.171.967.9
AutomationBench53.152.3
Terminal Bench 2.189.987.6
CyberGym94.095.1

Flash gives up four points on coding and basically nothing on automation, for about a third of Pro’s price. It even edges Pro on CyberGym. For triage, code review, and document work at volume, Flash is the one I’d start with.

Three honest caveats. First, those are Xiaomi’s numbers, and Flash isn’t on the Artificial Analysis index yet as of September 22. Second, Terminal Bench 2.1 is not Terminal-Bench 4.0, so don’t hold 89.9 up against Opus 5.5’s 66.4. Different test, different season. Third, Xiaomi hasn’t published hardware guidance for self-hosting. “15 billion active” makes Flash fast per token, but all 309 billion parameters still have to sit in memory somewhere. The MIT license is real. So is the GPU bill.

One more thing for anyone running client data. The Xiaomi API is Xiaomi-hosted. If data residency matters to your clients, that’s exactly what the open weights solve: run it on infrastructure you control, or route through a provider you’ve already vetted.

Which one runs which job

I stopped thinking about models as a single hire a long time ago. They’re a staff. Here’s how I’d staff this week’s roster:

  1. Long-running coding agents and client deliverables: Opus 5.5 at medium effort. Best independent score, and cheaper per finished task than the model that looks cheaper. This is your senior engineer.
  2. High-volume loops, code review, triage, and document work: MiMo V2.6 Flash. Start here, and step up to Pro when Flash misses. This is the crew that does ten times the reps for a fraction of the payroll.
  3. Anything with open-weights or self-hosting requirements: MiMo V2.6 Pro. Nothing else in this comparison lets you take the model home.
  4. Security research: Opus 5.5, but test it knowing about the Opus 4.8 reroute, and get into the Cyber Verification Program when it opens.
  5. Grok 4.7: only if you have a specific legal or EEBench-style workload, and only after you’ve run your own evals. For agentic coding, skip it until the independent numbers move.

And stack them. One model to plan, one model to grind, one model to check the work. That’s how you get Opus-grade output at MiMo-grade prices on the bulk of your traffic.

What I’d do Monday

If you’re paying for Opus 5 right now, switch the model string to claude-opus-5-5 and retest anything that depended on thinking being off. That’s the cheapest upgrade you’ll make all year: better scores, lower price, fewer tokens.

If your agents do a lot of sorting and reviewing at volume, put MiMo V2.6 Flash in front of that traffic this week and watch your bill. I already moved client work onto it, and the independent index backs the call.

And the next time a launch post leads with price per token, ask for the gas bill. The sticker price told you Grok was the budget pick. The receipt said otherwise.

For more on how the open-weights gap keeps shrinking, see Open Weights vs the Frontier. For the two models Opus 5.5 just passed, see Claude Fable 5.1 vs GPT-6 Astra. For the Muse model I ranked above Grok, see Muse Spark 1.3 vs Gemini 3.8 Flash.

Welp… Opus behaves now. Epic son.

#TheAIMogul

Bottom lineOpus 5.5 is the new default for serious agent work: the best independent score, and a lower bill per task than the model that looks cheaper on the price sheet. MiMo V2.6 is the value play and the open-weights pick. Grok 4.7 is a skip for agentic coding until independent numbers catch up to xAI's own.

Frequently asked

Which is best: Claude Opus 5.5, Grok 4.7, or MiMo V2.6?
For long-running coding agents and serious knowledge work, Claude Opus 5.5. It scores 58 on the Artificial Analysis Intelligence Index at max effort, the top score on the board, and costs $1.34 per task at default effort. For high-volume work on a budget, or if you need open weights you can self-host, MiMo V2.6 Pro scores 46 for $0.13 per task. Grok 4.7 also scores 46 but costs $2.73 per task and trails badly on independent agentic coding tests.
Is Grok 4.7 cheaper than Claude Opus 5.5?
Per token, yes. Grok 4.7 costs $2 per million input tokens and $6 per million output, versus $4 and $20 for Opus 5.5. Per finished task, no. On Artificial Analysis's independent testing at each model's default effort, Grok 4.7 cost $2.73 per task and Opus 5.5 cost $1.34, because Grok generated about 2.6 times as many output tokens and most of its cost came from input.
What is the difference between MiMo V2.6 Pro and MiMo V2.6 Flash?
Pro is the bigger model: 1.02 trillion total parameters with 42 billion active per token, priced at $0.435 per million input tokens and $0.87 per million output. Flash has 309 billion total and 15 billion active, priced at $0.14 and $0.28. On Xiaomi's own benchmarks Flash gives up a few points (67.9 vs 71.9 on DeepSWE v1.1) and actually edges Pro on CyberGym. Both are MIT-licensed, omnimodal, and support a 1-million-token context window.
Does Claude Opus 5.5 route requests to an older model?
For cybersecurity work, yes. Anthropic says most cybersecurity tasks will be re-routed to Opus 4.8, while routine work like finding and fixing bugs in your own code stays on Opus 5.5. Reporting on the launch says there is no visible sign when it happens. Anthropic says its Cyber Verification Program will extend to Opus 5.5 in the coming weeks for vetted security teams.
Can I self-host MiMo V2.6?
The license allows it. Both MiMo V2.6 Pro and Flash are MIT-licensed with weights on Hugging Face, so commercial use and self-hosting are permitted. Xiaomi has not published hardware guidance or throughput figures, and even Flash's 15 billion active parameters sit inside a 309-billion-parameter model that has to be held in memory, so plan real GPU capacity before you commit.
How much does Claude Opus 5.5 cost?
$4 per million input tokens and $20 per million output tokens, 20% below Opus 5. Cache reads dropped to $0.20 per million and cache writes to $5. A fast mode runs $8 input and $40 output. Anthropic says typical workloads cost about 40% less than Opus 5 overall because the model also finishes tasks in fewer tokens.