Skip to content
TrustList
News

A faster model tier at two and a half times the price, with no claim of more capability

Editorial

By TrustList Editorial

A hosted speed variant released on 18 September charges $0.37 and $1.25 per million tokens against $0.15 and $0.50 for the model it is built from — same context window, same output ceiling, same modalities, and a vendor peak of 200 tokens a second.

About A faster model tier at two and a half times the price, with no claim of more capability

A faster model tier at two and a half times the price, with no claim of more capability

20 September 2026 — A hosted variant released on 18 September is the clearest example this month of a category that is easy to misread on a price list: an infrastructure tier sold beside a model, at a model-sized premium.

The numbers, from the vendor's own pages

The vendor's documentation gives the new tier the same 1M-token context window as the model it is built from, the same 128K maximum output, and the same inputs — video, image, text and file in, text out. The only difference it claims is throughput: inference at up to 200 tokens a second.

The pricing table on the same site charges $0.37 per million input tokens, $1.25 per million output and $0.075 per million on cached input. The base model, on the same page, is $0.15, $0.50 and $0.03. That is roughly two and a half times the price for latency alone.

No benchmark score is worth quoting here, and that is the point rather than an omission: the vendor's own claim is that this is the same model served faster, so any figure lifted from the base model and printed under the new name would be a measurement nobody made.

There is one more line in the documentation that a buyer should read as pricing rather than packaging: the base model is now fully available on the vendor's coding subscription, and the faster tier is not yet on it. The premium is therefore a pay-as-you-go API cost rather than something a subscription absorbs.

What it means for a buyer

Decide first whether your workload is latency-bound or throughput-bound, because the premium only buys something in the first case. An interactive assistant where a person waits on the first token is a reasonable candidate. A nightly batch that processes a queue is not: it will finish at the same hour and cost two and a half times as much.

Second, treat 200 tokens a second as an advertisement rather than a measurement. It is a peak, quoted by the seller, on unstated conditions. If latency is the reason you are considering this, measure it on your own prompts — the ones with your real context length, which is usually the variable that decides the answer.

Third, watch for the same pattern elsewhere. Speed tiers priced as capability upgrades are becoming common, and the tell is exactly what it is here: identical specifications on the product page, a different number on the price page.

Sources