Four Flash Models in Four Months. The Bill Per Task Went Up 40%.

Google shipped Gemini 3.8 Flash on September 2, three weeks after 3.7 Flash, at exactly the same price per token. The same day, Artificial Analysis measured the cost of running its benchmark suite through the new model at $0.58 per task, up from $0.40 on 3.7 Flash. Both facts are true at once, and the gap between them is the whole story. The price of a token did not move. The number of tokens the model decides to spend did, by 30%, and Google says that is the design. This is what "works harder" means on an invoice, why the workhorse tier is where budgets actually move, and what happens on January 1 when the introductory price doubles.


A cadence, not a launch

The interesting fact about Gemini 3.8 Flash is not any single benchmark. It is the calendar. Gemini 3.7 Flash shipped on August 13. Gemini 3.8 Flash shipped on September 2, which 9to5Google counts as the third Flash update in three months and Artificial Analysis counts as the fourth Flash model in under four months. Whichever way you count, the model underneath a Flash-tier agent now changes roughly every three weeks, and the version number moves by a tenth each time.

Google's own framing, quoted by 9to5Google, is "our most intelligent workhorse model, delivering significant improvements from 3.7 Flash across software engineering, agentic tasks, and critical, multi-step reasoning in specialized domains," and it adds that the model is "often approaching the performance of higher-cost frontier models." The numbers Google publishes are Google's: 54.9% on HLE-Verified, a claim of outperforming most larger frontier models on DeepSWE v1.1 "at a fraction of the cost," and wins on the Vals Finance Agent V2 and Harvey legal agent benchmarks. Treat them as a vendor's scoreboard until someone independent reproduces them.

Availability is broad from day one: the Gemini app for AI Pro and Ultra subscribers, AI Mode, Gemini in Sheets, and for developers Antigravity, AI Studio, and the Gemini API. There is also a restricted twin, Gemini 3.8 Flash Cyber, offered to trusted testers through Google's new Fairwind Program. Google's claims for it are again Google's own: the Chrome Security team reporting 2.6 times more correct vulnerability patches than much larger commercial models, and Wiz reporting 7.5 to 9.7 points higher recall on an internal penetration-testing benchmark at 2.3 to 5.2 times lower cost. Google, Anthropic, and OpenAI all shipped or announced a gated cyber variant beside a public model in the same week. That pattern deserves its own essay. This one is about the public model and its bill.

What "works harder" means on the invoice

Google says the gains "stem from a core design choice: 3.8 Flash works harder." The sentence that follows is the one to read slowly. On complex tasks the model shows "greater diligence," executing extra reasoning steps and calling tools iteratively, and "at times, the model might use more tokens to maximize performance, especially at higher effort levels."

Now put Artificial Analysis's measurements next to that sentence. At high reasoning effort, Gemini 3.8 Flash scores 59 on their Intelligence Index, up 3 points from 3.7 Flash's 56. Average output tokens per task rose 30%, to about 48,000. Time per task rose from 2.2 minutes to 2.5, and the model streams at roughly 300 output tokens per second. And cost per task rose from $0.40 to $0.58, which they round to "about 40%." At medium effort the index is 57 for $0.41 per task. At low effort it is 52 for $0.24, and the task takes 0.8 minutes.

Read as a mechanism rather than a headline: the model got three points smarter by talking to itself 30% longer, and that extra talking is billed at the same rate as before. A model that "works harder" is a model that has been trained to spend more of your tokens before answering. The per-token price is a constant in that equation. The per-task cost is the output. If your dashboards track the first and your budget is set by the second, this release will look free and cost more, which is exactly the failure I wrote about in The Inference Bill: inference is priced per token and consumed per task, and the two drift apart the moment a model changes its own verbosity.

Fig. 1 · the diligence dial

Pick a reasoning effort. Every bar is an Artificial Analysis measurement of Gemini 3.8 Flash; the grey ghost is 3.7 Flash at high effort where they published it. The verdict does the division.

3.8 Flash at this effort3.7 Flash, high effort3.7 Flash marker

Index, cost per task, time per task, and output tokens per task from Artificial Analysis, September 2, 2026. Medium-effort time per task and low or medium token counts were not published, and are shown as such rather than guessed.

The dial gives back what the vendor took. At low effort the model buys 52 index points for $0.24, which is 217 points per dollar. At high effort it buys 59 points for $0.58, which is 102 points per dollar. The last seven points cost 2.4 times the money and 3.1 times the wall-clock time. None of those are judgments, they are quotients of published numbers, and they say that the effort setting is not a quality knob so much as a routing decision. A harness that sends every request at high effort because the launch post said "most intelligent" is paying the frontier premium on a workhorse. The harness, as I argued in The Harness, Not the Model, decides the bill more than the model does, and the effort parameter is the cheapest place in the harness to prove it.

Price times volume

A bill has two factors and only one of them is printed on the pricing page. Cost per task equals price per token times tokens per task. Google held the first factor flat and moved the second by 30%. Multiply them and you expect a 30% rise in cost per task. Artificial Analysis measured about 45% on the raw numbers ($0.58 against $0.40), which they describe as roughly 40%. The residual between 1.30 and 1.45 is not broken out in their article. Two ordinary things would produce it: a shift in the mix toward output tokens, which cost five times what input tokens do, or longer inputs as the model reads more before it writes. Either way, the arithmetic below is the only tool you need to catch it next time.

Fig. 2 · the bill decomposer

Two factors, one product. Drag the price of a token and the tokens per task. The baseline is 3.7 Flash at $0.40 per task; the red marker is what Artificial Analysis measured for 3.8 Flash.

+0%
+30%
price factor
1.00x
per token
×
volume factor
1.30x
tokens per task
=
cost per task
1.30x
$0.52
3.7 Flash
measured
$0.40
3.8 Flash
modeled
$0.52
3.8 Flash
measured
$0.58

Baseline and measured values from Artificial Analysis. The modeled value is the product of the two sliders applied to the $0.40 baseline; the residual is the ratio of measured to modeled and is reported, not explained.

Two things fall out of the decomposer. First, "no price change" is not the same as "no cost change," and a release note that says only the first is not lying, it is just answering a different question. Second, the volume factor is the one you can steer. Price is set in Mountain View. Tokens per task are set by your effort parameter, your context discipline, your tool schemas, and how many turns your loop takes before it stops. A 30% rise in volume you did not ask for can be undone by a 30% cut in context you did not need, and that trade is available today, at zero cost, in your own code.

The cliff on January 1

Gemini 3.8 Flash launched at $0.75 per million input tokens and $3.75 per million output tokens. Google calls it an introductory price and, per 9to5Google, it runs until December 31. Artificial Analysis lists the standard price as $1.50 and $7.50, which is exactly double on both sides. Cached input tokens keep a 90% discount, and the context window stays at one million tokens, unchanged from 3.7 Flash.

Introductory prices are not new, but two vendors handled theirs differently this summer and the contrast matters if you are budgeting. Anthropic launched Claude Sonnet 5 at $2 and $10 per million tokens with a planned rise to $3 and $15 on September 1, then made the introductory price permanent in an August 10 edit to the launch post. Google has said the opposite for Flash: the low price has an end date. The day before the Fairwind announcement, Anthropic also cut cache-read pricing on Claude Fable 5.1 by 75%, to $0.25 per million, which is a different lever again: same list price, cheaper repeats. Three vendors, three different places to move the number, and only one of them is the number on the pricing page.

Fig. 3 · the price-cliff calendar

Set a monthly token volume and how much of the input hits the cache. The strip prices the same volume through the introductory period and then on the standard rate from January 1.

400M
40M
0%

Rates: $0.75 / $3.75 per million input / output tokens through December 31, 2026; $1.50 / $7.50 standard; cached input at a 90% discount. Prices per Artificial Analysis and 9to5Google, September 2, 2026. Your actual bill also depends on batch, tier, and region pricing that this strip ignores.

The calendar makes the cache discount concrete. At the introductory rate a 90% discount on cached input is a nice-to-have. At the standard rate it is the difference between a bill that doubled and a bill that went up by a third, because on an agent workload most of the input is the same system prompt, the same tool schemas, and the same conversation prefix sent again and again. Cache discipline is a September problem with a January payoff.

The workhorse tier is where the money is

Frontier launches get the headlines. Workhorse launches get the volume. The model that runs your classification, your extraction, your first-pass code review, and your agent's inner loop is almost never the most expensive one on the price list, and it is the one whose per-task cost compounds across millions of calls. That is why a 40% rise in cost per task on a Flash model is a bigger budget event than a 10% change on a frontier model that handles a hundredth of the traffic.

Model (vendor)Input $/MOutput $/MNote
Gemini 3.8 Flash (Google), introductory0.753.75through December 31, 2026; cached input 90% off
Gemini 3.8 Flash (Google), standard1.507.50from January 1, 2027, per Artificial Analysis
Claude Haiku 4.5 (Anthropic)1.005.00cache read 0.10, per claude.com/pricing
Claude Sonnet 5 (Anthropic)2.0010.00introductory price made permanent August 10
Claude Fable 5.1 (Anthropic), frontier tier for contrast10.0050.00cache read 0.25 since September 1

List prices per million tokens as published on September 2 and 3, 2026. Prices for other vendors' workhorse tiers were not verified for this essay and are left out rather than guessed.

The cadence adds an operational tax that the price table does not show. When the model under your agent changes every three weeks, your evals have to run every three weeks, because the thing that changed is not only quality but verbosity, and verbosity is cost. There is a second, quieter reason to re-run them. Google notes, per 9to5Google, that the knowledge cutoff is March 2026 for some domains while in others the model's knowledge "is limited to January 2025." A model that got smarter and stayed a year and a half stale in some corners is a model whose retrieval layer matters more, not less, and retrieval means more input tokens, which brings us back to the cache discount.

The extra 0.3 minutes per task is the answer getting longer, not the engine getting slower, and Artificial Analysis's own numbers show it. Divide their 48,000 output tokens by 2.5 minutes and you get about 320 tokens per second; take 30% off those tokens for 3.7 Flash, about 37,000, divide by its 2.2 minutes, and you get about 280. Those are my quotients, not their measurement, but they bracket the roughly 300 tokens per second they report, and if you have read Tokens per Second Is a Memory Bandwidth Number you know why a hosted Flash model would land in the same place twice: the serving side did not change shape between versions, only the length of the answer did. The model chose to run longer.

What to do

Gemini 3.8 Flash is, by the independent numbers, a better model than the one it replaced, and at the same list price. It is also a more expensive model to run, by about 40% per task, and Google told you so in the launch post if you read it as an accountant. The job now is to read every launch post that way.

rg
Rohit Ghumare

CNCF Ambassador and Google Developer Expert. I build agent infrastructure and write about the fundamentals underneath the AI stack. Measurements here are Artificial Analysis's, Google's quotes come via 9to5Google's launch coverage, and Anthropic prices come from claude.com and the Sonnet 5 and Fable 5.1 announcements, all read on September 3, 2026. Introductory pricing and effort-level behavior change between releases, so check the current price list before you budget from these numbers.

Related: The Inference Bill · The Harness, Not the Model · More posts · X