25 Sept 2026
Will social media survive the AI change? What the evidence says
Our overnight sample had no social referrers, while Meta's crawlers made 5.9% of page requests. Usage, labelling rules and the EU AI Act…
Alibaba's streaming audio-and-video conversation model with spoken replies, tool calling and remote MCP, announced 18 September 2026; audio input costs $0.93 and audio output $1.87 per million tokens in Singapore.
Not yet independently verified. Alibaba gives two dates: the Qwen blog announced the model on 18 September 2026, and GIGAZINE (18 September) and ProPakistani (19 September) reported it then, but Model Studio's lifecycle page lists its release as 21 September. The announcement date is used here. We will update this when it can be confirmed, and remove this note.
Qwen3.8-Omni-Flash-Realtime is the low-latency, streaming version of Alibaba's Qwen3.8-Omni-Flash omnimodal model. The Qwen team's announcement, dated 18 September 2026, introduced it alongside Qwen3.8-Omni-Flash and included API sample code for it. Alibaba Cloud Model Studio's model lifecycle page lists its release as 21 September 2026 in the Singapore and China (Beijing) regions.
It takes streaming audio and images, including video frames at a recommended one frame per second, and replies with text and speech. Connections use WebSocket, WebRTC or Alibaba's AOQ protocol. Alibaba lists custom function calling, remote MCP tools (approval required by default), two- or four-channel audio input, video aggregation to reduce computation, voice cloning, and spatial-audio perception that it says can locate a sound source. Speech recognition covers 113 languages and dialects and speech generation 36. Total input is capped at 196,608 tokens, a WebSocket session lasts at most 120 minutes, and context keeps up to 100 audio turns or 600 seconds of audio. Web search and tool calling cannot be used together. In its own measurements, Alibaba reports a time to first audio packet of about one second for short audio inputs.
Singapore pricing per million tokens is $0.23 for text, image and video input, $0.93 for audio input, $0.70 for text output and $1.87 for audio output; spoken replies are billed for both the audio and its transcript. Input audio is counted at 7 tokens per second and output audio at 12.5. A one-million-token free quota applies in Singapore only.
The model is proprietary and available only through the API. Alibaba has open-sourced a companion runtime, Qwen-Live Harness, but not the model weights. Suggested uses include live customer service, speaking practice and camera-based assistants.
| Benchmark | Official | Community avg |
|---|---|---|
| No benchmark scores yet. Be the first to add one. | ||
“Official” values are editor-approved and feed the ranking. “Community avg” is the mean of member submissions (shown for transparency; it never affects the ranking until an editor approves a value).
Sign in to add a benchmark score for this model.
An earned signal from verification, reviews, awards, transparency and engagement — the vendor can't buy it.
Updated 9/22/2026
Used it? Your experience helps other buyers decide.
Write a reviewNo questions yet. Be the first to ask about Qwen3.8-Omni-Flash-Realtime.
Everything here links back to the same verified catalogue. Pick your next stop.