Comparison Head to Head

Muse Spark 1.3 vs Gemini 3.8 Flash: Two Points and Three Cents

Meta and Google shipped budget-tier flagships on the same day. One scores higher and costs less per task. The other has a price cliff on its calendar.

An antique brass balance scale with a blue Meta infinity loop on one pan and a purple Gemini star on the other, tipped toward Meta.
Illustration generated for Run the Eval
The receipts
  • Meta and Google both shipped on September 2, 2026. Muse Spark 1.3 (xhigh) scores 61 on the Artificial Analysis Intelligence Index; Gemini 3.8 Flash scores 59.
  • Cost per Intelligence Index task: $0.55 for Muse Spark, $0.58 for Gemini 3.8 Flash. Muse wins on both axes, barely.
  • Gemini 3.8 Flash's headline price doubles on January 1, 2027 — $0.75/$3.75 becomes $1.50/$7.50 per million tokens.
  • Muse Spark's cache hits land at $0.15 per million. If your workload re-reads a big prompt, that is the number that decides your bill.
Short answer

Muse Spark 1.3 (xhigh) scores 61 on the Artificial Analysis Intelligence Index at $0.55 per task; Gemini 3.8 Flash scores 59 at $0.58 per task. Muse Spark's API is $1.25 per million input and $4.25 per million output tokens. Gemini 3.8 Flash is $0.75 input and $3.75 output through December 2026, doubling to $1.50 and $7.50 on January 1, 2027.

Two labs shipped budget-tier flagships on September 2. Same day. That is not a coincidence, and it is not a coincidence that both of them are priced to be somebody’s default.

Here is the whole comparison in one table, and then I will tell you what the table leaves out.

Muse Spark 1.3 (xhigh)Gemini 3.8 Flash (high)
Intelligence Index6159
Input / 1M tokens$1.25$0.75 (through Dec 31, 2026)
Output / 1M tokens$4.25$3.75 (through Dec 31, 2026)
Cached input / 1M$0.1590% off list ($0.075 through 2026)
Price on Jan 1, 2027unchanged$1.50 / $7.50
Cost per Index task$0.55$0.58
Terminal-Bench 2.185%90.8%
Context window1M tokens1M tokens
Where it losesSticker price per token; terminal workToken appetite per task; the 2027 price cliff

Six days later Meta put this model to work in a consumer product: Meta Muse, the personal AI agent that runs on Muse Spark and is free for most people.

The two-point lead is real and it is small

Meta’s Muse Spark 1.3 scores 61 on the Artificial Analysis Intelligence Index. That ties GPT-5.6 Sol at max effort and Grok 4.6 at high. A limited-preview max variant gets to 62. Google’s Gemini 3.8 Flash scores 59, three points up from 3.7 Flash, level with GPT-5.6 Sol at xhigh and Grok 4.6 at medium.

Two points. I have watched people re-architect a stack over two points and regret it. That gap is inside the range where your prompt quality matters more than your model choice.

Don’t move for two points.

What the Index is actually scoring

Know what the number is before you argue over it. Intelligence Index v4.2 is a weighted average across four categories: Agents 30%, Coding 20%, Scientific Reasoning 20%, General 30%. Ten evaluations feed it. Terminal-Bench 2.1 is one of them, filed under Coding at a 10% weight.

The weighting leans agentic, and both models earned their gains there. Artificial Analysis says Gemini’s three points came mostly from tool-use and agent evals, with the biggest jump on τ³-Banking, up 12 points to 45%. Muse Spark’s xhigh went from 35% to 47% on the same banking eval versus 1.2, and from 80% to 85% on Terminal-Bench 2.1.

Same story from two labs. The budget tier is getting better at tool use faster than it is getting better at anything else.

The interesting number is per-task, not per-token

Gemini is cheaper per token and more expensive per job. Muse Spark costs $0.55 per Intelligence Index task; Gemini 3.8 Flash costs $0.58.

Here is why that metric exists. Artificial Analysis prices the whole Index run using the token counts each provider’s API reports, combined with live measurements of the model’s typical cache hit rate. So cost per task is the invoice for finishing the work, not the rate card. It captures how many tokens a model burns thinking, how many turns it takes, how much of its context comes back from cache.

The Gemini receipt makes this concrete. Per-token pricing did not change from 3.7 Flash to 3.8 Flash. Cost per task still rose about 40%, from $0.40 to $0.58, because average output tokens per task went up 30% to 48k and the model takes more turns on agentic evals. Same sticker. Bigger bill.

Meta went the other direction. 1.3 is quieter than 1.2, roughly 20% fewer tool calls and 25% fewer tokens on the same work, at the same price as the older model. That is the whole reason a $1.25 input model undercuts a $0.75 input model on cost per job.

That is the metric I actually trust. Token price is what a lab charges you. Cost per task is what the work costs. A model that thinks in fewer moves beats a cheaper model that flails, every time — the same dynamic I walked through in what AI coding agents actually cost.

And for context: Muse Spark’s Index peers, GPT-5.6 Sol (max) and Grok 4.6 (high), cost $0.95 and $0.94 per task. Same 61 score, a 70%-plus premium. That is the actual headline of the Muse release and it is not about Gemini at all.

The worked example

Say your agent runs 10 million input tokens and 2 million output tokens a month, and 8 million of those input tokens are cache hits because you keep re-reading the same codebase or policy set. List prices from the table, arithmetic only.

Muse Spark 1.3: 8M cached at $0.15 is $1.20. 2M fresh input at $1.25 is $2.50. 2M output at $4.25 is $8.50. Total $12.20.

Gemini 3.8 Flash, 2026 pricing: Artificial Analysis notes cached input keeps a 90% discount, so 8M at $0.075 is $0.60. 2M fresh input at $0.75 is $1.50. 2M output at $3.75 is $7.50. Total $9.60.

Gemini 3.8 Flash, January 1, 2027 pricing: 8M cached at $0.15 is $1.20. 2M fresh input at $1.50 is $3.00. 2M output at $7.50 is $15.00. Total $19.20.

Read that twice. On the rate card Gemini wins today by $2.60 and loses in four months by $7.00. And the rate card assumes both models need the same tokens to finish the job, which $0.55 versus $0.58 says they do not.

The date nobody put on the slide

Gemini 3.8 Flash is $0.75 and $3.75 per million through the end of 2026. On January 1, 2027, it becomes $1.50 and $7.50. Double. Batch pricing today is $0.375 and $1.875, half of list.

Double every token price, hold token usage flat, and cost per task doubles with it: $0.58 becomes roughly $1.16. That is more than GPT-5.6 Sol (max) costs today, for a lower score.

So the honest read: Gemini’s price advantage has a four-month shelf life. If you are picking a model for something that ships in Q1, you are not choosing between $0.75 and $1.25. You are choosing between $1.50 and $1.25, and Muse Spark just won that on points too.

Google is moving fast on this line… Artificial Analysis calls 3.8 Flash its fourth Flash model in under four months. That cadence is exactly why I model the standard price and not the promo. The promo is a ramp. The standard price is the product.

Terminal-Bench is the tiebreaker, and it cuts Google’s way

Gemini 3.8 Flash hits 90.8% on Terminal-Bench 2.1, up from 81.6% for 3.7 Flash. Muse Spark 1.3 lands at 85% on xhigh and 86% on max.

Terminal-Bench is the agent-in-a-shell test. Per the Index methodology it is 89 terminal-based tasks, scored by test suite pass or fail at pass@1, run three times. No partial credit for a pretty diff. Either the task resolved or it did not.

If your workload is a coding agent living in a terminal, that is the benchmark that measures your job. A near six-point lead there outweighs a two-point lead on an aggregate where this test is one-tenth of the weight.

Speed cuts the same way. On high reasoning Gemini 3.8 Flash averages about 300 output tokens per second with a Time per Task of 2.5 minutes. Drop it to low reasoning and Time per Task falls to 0.8 minutes at a 52 score and $0.24 per task. Medium is 57 at $0.41. That is three price points on one model, and the bottom one is a real product number.

Where I would actually put each one

Coding agents in a terminal, CI fixers, anything shell-driven: Gemini 3.8 Flash, through December. Re-price it before January.

Long-context agents over a fixed corpus, high-volume extraction, classification, anything where cost per job is the KPI: Muse Spark 1.3. Stable price, 1M context, $0.55 a task, and a cache line that already matches Gemini’s 2027 rate.

Latency-bound features where a 52 is good enough: Gemini 3.8 Flash on low reasoning. Under a minute per task and a quarter per task. Nothing at Muse’s tier is competing for that slot.

Anything with a budget that has to survive Q1 2027: Muse Spark, unless Google extends the promo… and I do not build forecasts on unless.

If you are choosing at the top of the stack instead, I ran the same exercise on the big models in GPT-5.6 Sol vs Gemini 3.1 Pro. Different tier, same lesson.

The broader pattern holds, and it is the same one I traced in open weights versus the frontier: the interesting competition stopped being at the top. It is down here, where two points and three cents decide it.

#TheAIMogul

Bottom lineMuse Spark 1.3 is the better buy today on points and on price. But Gemini 3.8 Flash is 90.8% on Terminal-Bench 2.1, and if terminal work is your job, two Index points don't outrank that.

Frequently asked

Which is smarter, Muse Spark 1.3 or Gemini 3.8 Flash?
Muse Spark 1.3, by two points. On the Artificial Analysis Intelligence Index the xhigh variant scores 61 and Gemini 3.8 Flash scores 59. A limited-preview max variant of Muse Spark scores 62. Two points at this tier is a real but narrow lead — both models sit in the same band as sub-maximum reasoning efforts of GPT-5.6 Sol and Grok 4.6.
Which one is cheaper?
It depends on what you measure. Per token, Gemini 3.8 Flash is cheaper right now at $0.75 input and $3.75 output per million versus Muse Spark's $1.25 and $4.25. Per unit of work, Muse Spark is cheaper at $0.55 per Intelligence Index task against Gemini's $0.58, because it needs fewer tokens to finish the job.
What happens to Gemini 3.8 Flash pricing in 2027?
It doubles. Google is holding $0.75 per million input and $3.75 per million output through the end of 2026, then moving to $1.50 and $7.50 on January 1, 2027. If you are costing out a build that ships in Q1, model the 2027 number, not the one on the pricing page today.
Is Muse Spark 1.3 better than Muse Spark 1.2?
Meaningfully, and at the same price. Meta says 1.3 uses roughly 20% fewer tool calls and 25% fewer tokens than 1.2 on the same work. That efficiency is the whole reason its cost per task beats a model with a lower sticker price.
Which should I use for coding and terminal work?
Look at Terminal-Bench 2.1 before you look at the Index. Gemini 3.8 Flash scores 90.8% there, up from 81.6% for 3.7 Flash. That is a specific, large jump on exactly the kind of tool-driven shell work coding agents do all day, and it is a better signal for that job than a two-point general-intelligence gap.