A faster model tier at two and a half times the price, with no claim of more capability
EditorialBy TrustList Editorial
A hosted speed variant released on 18 September charges $0.37 and $1.25 per million tokens against $0.15 and $0.50 for the model it is built from — same context window, same output ceiling, same modalities, and a vendor peak of 200 tokens a second.
About A faster model tier at two and a half times the price, with no claim of more capability
A faster model tier at two and a half times the price, with no claim of more capability
20 September 2026 — A hosted variant released on 18 September is the clearest example this month of a category that is easy to misread on a price list: an infrastructure tier sold beside a model, at a model-sized premium.
The numbers, from the vendor's own pages
The vendor's documentation gives the new tier the same 1M-token context window as the model it is built from, the same 128K maximum output, and the same inputs — video, image, text and file in, text out. The only difference it claims is throughput: inference at up to 200 tokens a second.
The pricing table on the same site charges $0.37 per million input tokens, $1.25 per million output and $0.075 per million on cached input. The base model, on the same page, is $0.15, $0.50 and $0.03. That is roughly two and a half times the price for latency alone.
No benchmark score is worth quoting here, and that is the point rather than an omission: the vendor's own claim is that this is the same model served faster, so any figure lifted from the base model and printed under the new name would be a measurement nobody made.
There is one more line in the documentation that a buyer should read as pricing rather than packaging: the base model is now fully available on the vendor's coding subscription, and the faster tier is not yet on it. The premium is therefore a pay-as-you-go API cost rather than something a subscription absorbs.
What it means for a buyer
Decide first whether your workload is latency-bound or throughput-bound, because the premium only buys something in the first case. An interactive assistant where a person waits on the first token is a reasonable candidate. A nightly batch that processes a queue is not: it will finish at the same hour and cost two and a half times as much.
Second, treat 200 tokens a second as an advertisement rather than a measurement. It is a peak, quoted by the seller, on unstated conditions. If latency is the reason you are considering this, measure it on your own prompts — the ones with your real context length, which is usually the variable that decides the answer.
Third, watch for the same pattern elsewhere. Speed tiers priced as capability upgrades are becoming common, and the tell is exactly what it is here: identical specifications on the product page, a different number on the price page.
Sources
- Z.ai model documentation for the GLM-5.3 family (read 20 September 2026) — 20 September 2026
- Z.ai pricing table (read 20 September 2026) — 20 September 2026
More on TrustList
Everything here links back to the same verified catalogue. Pick your next stop.
- CompaniesAgencies, consultancies and IT service providers, ranked by verified reviews.
- ProductsSoftware and SaaS with pricing, features, integrations and alternatives.
- AwardsAnnual recognition decided by verified reviews and an independent jury.
- LaunchesNew products and releases, voted up by the community every day.
- AI ModelsBenchmark scores and community ratings for every major model.
- RequestsBuyers describe what they need; vendors respond directly.
- PeopleReviewers, authors and makers with public profiles.
- ComparePut up to four listings side by side before you shortlist.