Skip to content
TrustList
All launches
AI model launch

Gemini 3.8 Live

Google's low-latency speech-to-speech model for voice agents — sees video as it talks, 97 languages, $3/$12 per million audio tokens.

Launched 2026-09-15AI model

About this launch

Gemini 3.8 Live is a real-time dialogue model that Google began rolling out on 15 September 2026, built, in Google's words, for scale and cost efficiency. It takes audio, images, video and text in and answers in audio or text, processing visual input in near real time during a conversation, and it detects and switches between 97 supported languages mid-conversation. The Gemini API model page (model id `gemini-3.8-live`, listed as a stable version) gives an input limit of 131,072 tokens and an output limit of 65,536, with function calling, Search grounding and thinking supported, and caching, structured outputs, code execution and batch not supported. On the Gemini API pricing page it shares a price row with 3.8 Live Extended Thinking: per million tokens, $0.75 for text input, $3.00 for audio input (or $0.005 a minute) and $1.00 for image or video input, and $4.50 for text output and $12.00 for audio output (or $0.018 a minute). Google reports that it placed second in the Speech Agent Arena, a user-preference ranking; that is Google's own account. It is available to developers in the Gemini API and Google AI Studio, in private preview in Gemini Enterprise, and to consumers in Search Live. Google's model card says the Gemini 3.8 audio models are based on Gemini 3 Pro, and all audio they generate carries a SynthID watermark. For more reasoning during a conversation, Google offers the separate 3.8 Live Extended Thinking model.

Discussion

Questions & answers about Gemini 3.8 Live.

No questions yet. Be the first to ask about Gemini 3.8 Live.