Skip to content
TrustList
FR
Artificial Intelligence News

Five releases in four days, and not one of them was really about capability

Editorial

By TrustList Editorial

OpenAI, Sakana AI and DeepSeek all shipped in the same four days. Every release was a pricing decision wearing a model announcement.

About Five releases in four days, and not one of them was really about capability

Five releases in four days, and not one of them was really about capability

Between 8 and 11 September 2026, three companies shipped five models. Read the announcements next to each other and something odd emerges: almost none of the interesting content is about what the models can do.

OpenAI put out two image models on the same day, distinguished from each other not by architecture but by speed and cost. Sakana AI shipped a pair that share one system, split by whether you want the best answer or the cheapest adequate one. DeepSeek shipped a single model with two prices, and which one you pay depends on what time it is.

Three companies, three different products, one decision. Each of them stopped selling a model and started selling a position on a price-performance curve.

What actually shipped

GPT Image 2.5 Sunburst and GPT Image 2.5 Flare. OpenAI's model reference lists both, and the entire published difference between them is in their descriptions. Sunburst is "our most capable model for image generation and editing". Flare is "fast, high-quality everyday image generation". That is the whole positioning: one for when it matters, one for when it is routine. The reference gives no pricing and no release date for either; a third-party release timeline lists both as 8 September.

Fugu Max and Fugu Ultra v2.0, from Sakana AI, on 11 September. These share what the company calls "the same core orchestration architecture optimized for two distinct missions". Fugu Max answers the question "what is the best possible output we can deliver at the lowest possible cost?" and is priced at $2 per million input tokens and $6 per million output. Fugu Ultra v2 goes the other way, for complex work. Sakana reports 48.3 on Chartography and 74.3 on DeepSWE for Ultra v2, and says it is top two on seven of eight benchmarks it reports. It does not publish a price for Ultra v2 at all.

DeepSeek V4.1 Flash, on 10 September, served under the alias deepseek-flash. A 1M-token context window, and a price that moves with the clock: $0.30 per million input tokens and $1.20 output at peak, $0.15 and $0.60 off-peak. Cached input runs at $0.006 peak, $0.003 off-peak. Peak is Monday to Friday, 01:00 to 04:00 and 06:00 to 10:00 UTC. Everything else bills at half.

Five releases. Two of them are a capability tier and a volume tier of the same thing. Two more are a capability tier and a cost tier of the same architecture. The fifth is one model sold at two prices depending on the hour.

The same decision, three times, by companies with nothing in common

It would be easy to read this as coincidence. Three labs, three product strategies, a slow news week. But the companies involved have almost nothing in common — different countries, different funding, different customers, radically different positions in the market — and they arrived at the same structural answer inside four days.

That answer is: the unit you sell is no longer the model. It is an access tier. The model underneath may be identical, near-identical, or an entire routing system pretending to be one endpoint. What varies, and what the customer actually chooses between, is the price and the latency.

Sakana states this outright rather than leaving it to be inferred. Its own announcement says: "Fugu is an architecture, not a single model." The system "dynamically routes tasks to the leanest model capable of solving them", drawing on open-weights and specialist models, including NVIDIA Nemotron models through a partnership. What a Fugu customer buys is not a set of weights. It is a dispatcher, and a promise about what it will cost to run.

Once you see it in Sakana's framing, the other two read the same way. Sunburst and Flare are a dispatcher that the customer operates by hand: OpenAI has not told you what is different inside, only which one to reach for. DeepSeek's peak/off-peak table is a dispatcher operating on the time axis — the same capability, metered like electricity.

Pricing by the clock is the genuinely new thing here

Tiered pricing is not new. Fast-and-cheap alongside slow-and-capable has been standard for two years. What DeepSeek has done is different in kind, and it is worth separating from the rest.

Every other pricing decision in this market has been about what you are buying: a smaller model, a shorter context, a cached prefix, a batch queue that trades latency for cost. DeepSeek's peak/off-peak split is about when you are buying. The model is the same. The context window is the same. The output is the same. You pay double between 01:00 and 04:00 UTC on a Tuesday, and half at the weekend.

This is how electricity is sold, and how cloud spot capacity is sold, and it tells you something about the underlying constraint. A company prices by time of day when the scarce resource is not the product but the capacity to deliver it at a particular moment. Batch discounts already hinted at this; an explicit peak window states it.

For a buyer, this changes a familiar calculation in an unfamiliar way. The question "what does this model cost?" no longer has a single answer, and the honest response to a procurement spreadsheet is a distribution rather than a figure. A workload that can wait — overnight document processing, weekly enrichment runs, anything batch-shaped — is now materially cheaper than the same tokens consumed during an interactive session, on the same model, from the same vendor.

It also quietly penalises a particular kind of user: the European or UK team whose working day sits close to the peak window. The 06:00–10:00 UTC block is the first half of a London morning.

The cache rate is the number nobody quotes

There is a second figure on that same pricing page which is arguably more consequential than the peak/off-peak split, and it gets almost no attention because it does not fit in a comparison table.

DeepSeek V4.1 Flash charges $0.15 per million input tokens off-peak when the input is a cache miss. When it is a cache hit, it charges $0.003. That is a factor of fifty on the same tokens, going into the same model, at the same hour.

Cache economics of this shape change what is worth building. A long system prompt, a large retrieved document set, a fixed corpus that every request reasons over — these are the things that make prompts expensive, and they are exactly the things that cache well because they do not change between calls. At a fifty-fold differential, the engineering effort of structuring requests so the stable part comes first and stays byte-identical stops being a micro-optimisation and becomes one of the larger cost levers available.

It also makes published price comparisons close to meaningless without a stated cache-hit rate. Two teams running identical workloads against this model can see bills that differ by an order of magnitude, entirely because one of them restructured its prompts. When a vendor quotes "input tokens from $0.003" and a competitor quotes a flat $0.20, the two numbers are not describing the same thing, and neither is wrong.

We record the cache-miss peak rate on this directory for exactly that reason: it is the number a team pays before they have optimised anything, which is the only figure that is comparable across vendors without a footnote.

What this looks like from the buying side

If you are choosing between models rather than writing them, three things follow.

A model name has stopped being a specification. "We use Fugu" no longer tells a colleague what runs when a request comes in — that is the point of an orchestrator. "We use GPT Image 2.5" does not say whether you are on the capable tier or the everyday one. The name identifies a vendor relationship, not a system. Anything that matters for reproducibility, cost or risk now lives one level down, in the tier and the routing configuration.

Published prices need a timestamp and a condition attached. We record DeepSeek V4.1 Flash on this directory at its peak rate, not its off-peak rate, because the peak rate is what a buyer pays during a normal working day and recording the discount as the price would understate real cost. That is a judgement call, and it is the sort of judgement call every price comparison in this market now requires. A single number in a comparison table is doing more concealing than revealing.

Benchmark claims and price claims are drifting apart. Sakana names six benchmarks it says Fugu Max is best overall on, and publishes a figure for none of them in that announcement. It publishes two numbers for Ultra v2 — the capability tier — and no price. OpenAI publishes descriptions for its image pair and neither prices nor benchmarks. This is not dishonesty; each company is leading with the axis it is competing on. But it means the two halves of a purchase decision increasingly come from different documents, and sometimes from different companies.

This is the second time this pattern has appeared in a month

A week ago we wrote about a different split running through the same market: capability separated from permission to use it. Gemini 3.8 Flash Cyber behind Google's Fairwind Program, GPT-6 Astra's staged rollout through the gated Daybreak programme, Claude Mythos 5.1 offered to vetted organisations inside Project Glasswing. Three labs, one week, the same structural move — the model exists, and whether you may use it is a separate question with its own answer.

That was framed as a safety story, and it is one. But it is also the same dissolution happening on a different axis. In the access-gating case, "the model" split into capability and permission. In this week's case, it split into capability and price. Either way, the thing that used to be a single purchasable object — a model, with a name, a score and a rate card — has become a set of coordinates.

Our own launch board shows how fast this is now moving. In the 30 days to 12 September it holds 19 AI models — a new one every day and a half, from a board that only records releases it can verify against a named source, and five of those nineteen are the ones described above. At that rate, the question "which model is best" is becoming close to unanswerable in the form it is usually asked, and "which tier of which system, at what hour, for which workload" is the question that actually has an answer.

What we are not claiming

Three limits are worth stating plainly, because this is a directory rather than a newsletter and the distinction between what is known and what is inferred is the product.

Five releases in four days is a pattern, not a proof. It is a small sample from a fast-moving month, and a fortnight of single-model releases with flat pricing would weaken the reading considerably. We think the Sakana framing makes it more than coincidence, because a company describing its own product as "an architecture, not a single model" is not something you read into the data.

We do not know what is different inside the OpenAI pair. Sunburst and Flare may be separate models, the same model at different inference settings, or something else entirely. OpenAI's reference says which to reach for and does not say why. We have recorded exactly what it says.

Fugu Max's benchmark claim is a claim. Best overall on six named benchmarks, with no published figure for any of them at the time of writing. It may well be true. It is not currently checkable, which is why no score for Fugu Max appears on this board and why the sentence you just read exists.

Everything above traces to a page you can open yourself. Where we could not verify something — Fugu Ultra v2's price, the OpenAI models' pricing, licences and context windows for either pair — the field is blank on our listings rather than filled with a plausible guess. Blank is a worse-looking listing and a more honest one.

Sources