Cloudflare cuts Clef-flash price to $0.038 and its context to 24k tokens
EditorialBy TrustList Editorial
The hosted decision model drops from $0.09 to $0.038 per million input tokens and its window from 64k to 24k tokens. Clef stays at $0.24 with 64k, and Clef-omni launches at $0.15.
About Cloudflare cuts Clef-flash price to $0.038 and its context to 24k tokens
Cloudflare cuts Clef-flash price to $0.038 and its context to 24k tokens
9 October 2026: Cloudflare has cut the price of Clef-flash, the smallest model in its Clef family of decision models on Workers AI, from $0.09 to $0.038 per million input tokens. The hosted model's context window shrinks at the same time, from the 64,000 tokens Cloudflare advertised at launch to 24,000, so a workflow that sends Clef-flash longer inputs now has to move to a larger model.
Not yet independently verified. Single source: Cloudflare's own post. The post says its developer documentation carries the current prices, and we have not compared the two. We will update this when it can be confirmed, and remove this note.
The change was announced in a Cloudflare blog post on 9 October, the same post that launched Clef-omni, a version of the model that also takes audio and video, at $0.15 per million input tokens. Clef, the larger model, stays at $0.24 per million input tokens and keeps its 64,000-token window.
Cloudflare ties the smaller window directly to the lower price. By its own count only 0.24% of requests to Clef-flash exceed 24,000 input tokens, and it points customers with longer inputs to Clef. The limit applies to the hosted service only: the post says the weights on Hugging Face are unchanged and were trained to support a 256,000-token window for anyone who runs the model themselves.
Clef also got faster without new weights. Cloudflare moved serving to SGLang and gives median latencies falling from 262 to 152 milliseconds for inputs of about 800 tokens, from 616 to 305 milliseconds at about 3,400 tokens and from 2,721 to 1,635 milliseconds at about 16,000 tokens.
For teams already calling Clef-flash, the practical check is input size. Images and audio are converted to input tokens for pricing and limits, which the developer documentation explains, so a request that looks short can still pass 24,000 tokens. Routing long requests to Clef, or capping them before the call, avoids failures under the new limit.
Our AI model category lists Clef, Clef-flash and the new Clef-omni among its models.
Cloudflare says its developer documentation is the place for the most up-to-date prices.
Sources
Categories & features
TrustList Weekly
The week in software and IT, in one email
The news that matters to buyers, new rankings and our own research. Every Thursday, free, and easy to leave.
More on TrustList
Everything here links back to the same verified catalogue. Pick your next stop.
- More Artificial Intelligence SoftwareThe ranking for this subject
- CompaniesAgencies, consultancies and IT service providers, ranked by verified reviews.
- ProductsSoftware and SaaS with pricing, features, integrations and alternatives.
- AwardsAnnual recognition decided by verified reviews and an independent jury.
- LaunchesNew products and releases, voted up by the community every day.
- AI ModelsBenchmark scores and community ratings for every major model.
- RequestsBuyers describe what they need; vendors respond directly.
- PeopleReviewers, authors and makers with public profiles.
- ComparePut up to four listings side by side before you shortlist.