Skip to content
TrustList
News

Cloudflare cuts Clef-flash price to $0.038 and its context to 24k tokens

Editorial

By TrustList Editorial

The hosted decision model drops from $0.09 to $0.038 per million input tokens and its window from 64k to 24k tokens. Clef stays at $0.24 with 64k, and Clef-omni launches at $0.15.

About Cloudflare cuts Clef-flash price to $0.038 and its context to 24k tokens

Cloudflare cuts Clef-flash price to $0.038 and its context to 24k tokens

9 October 2026: Cloudflare has cut the price of Clef-flash, the smallest model in its Clef family of decision models on Workers AI, from $0.09 to $0.038 per million input tokens. The hosted model's context window shrinks at the same time, from the 64,000 tokens Cloudflare advertised at launch to 24,000, so a workflow that sends Clef-flash longer inputs now has to move to a larger model.

Not yet independently verified. Single source: Cloudflare's own post. The post says its developer documentation carries the current prices, and we have not compared the two. We will update this when it can be confirmed, and remove this note.

The change was announced in a Cloudflare blog post on 9 October, the same post that launched Clef-omni, a version of the model that also takes audio and video, at $0.15 per million input tokens. Clef, the larger model, stays at $0.24 per million input tokens and keeps its 64,000-token window.

Cloudflare ties the smaller window directly to the lower price. By its own count only 0.24% of requests to Clef-flash exceed 24,000 input tokens, and it points customers with longer inputs to Clef. The limit applies to the hosted service only: the post says the weights on Hugging Face are unchanged and were trained to support a 256,000-token window for anyone who runs the model themselves.

Clef also got faster without new weights. Cloudflare moved serving to SGLang and gives median latencies falling from 262 to 152 milliseconds for inputs of about 800 tokens, from 616 to 305 milliseconds at about 3,400 tokens and from 2,721 to 1,635 milliseconds at about 16,000 tokens.

For teams already calling Clef-flash, the practical check is input size. Images and audio are converted to input tokens for pricing and limits, which the developer documentation explains, so a request that looks short can still pass 24,000 tokens. Routing long requests to Clef, or capping them before the call, avoids failures under the new limit.

Our AI model category lists Clef, Clef-flash and the new Clef-omni among its models.

Cloudflare says its developer documentation is the place for the most up-to-date prices.

Sources

Categories & features

TrustList Weekly

The week in software and IT, in one email

The news that matters to buyers, new rankings and our own research. Every Thursday, free, and easy to leave.

We will email you to confirm. Unsubscribe with one click in any issue. Privacy policy