Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking are Google’s newest real-time audio models for voice agents, multimodal conversations, and tool-using workflows. Google announced both on September 15, 2026, expanded the rollout details on September 17, and on September 24 made Gemini 3.8 Live with Live Avatar generally available in Gemini Enterprise.
The important change is not just better speech quality. Google is separating two voice-agent workloads that used to be forced into one design: low-latency conversation and deeper multi-step reasoning. The standard Live model prioritizes fast interaction, while Extended Thinking can keep talking, narrate progress, and continue reasoning or calling tools in the background.
If your application needs generated speech rather than a live conversational session, see our Gemini 3.8 TTS API guide for the separate TTS model IDs, pricing, voices and API behavior.
This guide focuses on what developers and buyers actually need to know: model IDs, pricing, availability, migration behavior, technical limits, and the practical difference between the two models. It is based on Google’s official launch post and Gemini API documentation; it is not a hands-on benchmark.
Gemini 3.8 Live at a glance
| Model | Model ID | Best fit | Reasoning behavior |
|---|---|---|---|
| Gemini 3.8 Live | gemini-3.8-live |
Low-latency voice agents, direct commands, fast tool calls | Interleaved reasoning with a fixed latency profile |
| Gemini 3.8 Live Extended Thinking | gemini-3.8-live-extended-thinking |
Complex support, booking, diagnostics, multi-tool workflows | Configurable background reasoning with low, medium, or high thinking levels |
Both models accept text, images, audio, and video. Google’s model documentation lists a 131,072-token input limit and a 65,536-token output limit for both. Both support the Live API, function calling, Search grounding, and audio generation.
What changed from Gemini 3.1 Flash Live?
Gemini 3.8 Live is the direct migration target for many applications currently using gemini-3.1-flash-live-preview. The basic move is to change the model string to gemini-3.8-live, but there are behavioral differences worth planning for.
- Asynchronous tool use is now central. Gemini 3.8 Live supports non-blocking function execution by default, so a voice agent can keep the conversation moving while a tool runs.
- Extended Thinking introduces a different session lifecycle. The client must track
interaction_statusrather than assuming thatturnComplete: truemeans all reasoning and tool work has finished. - Thinking configuration differs by model. Standard Gemini 3.8 Live does not accept
thinking_level. Extended Thinking supports low, medium, and high background reasoning levels. - Proactive audio is permanently enabled on the 3.8 Live models. That matters for both user experience and cost because the API can continue processing audio while it is listening.
Gemini 3.8 Live API pricing
Google currently prices Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking at the same published rates in the Gemini Developer API.
| Usage | Paid API rate | Approximate audio rate |
|---|---|---|
| Text input | $0.75 per 1M tokens | — |
| Audio input | $3.00 per 1M tokens | $0.005/minute |
| Image/video input | $1.00 per 1M tokens | $0.002/minute |
| Text output, including thinking tokens | $4.50 per 1M tokens | — |
| Audio output | $12.00 per 1M tokens | $0.018/minute |
Google also lists a free tier for these models. For paid accounts, Google Search grounding includes 5,000 search requests per month shared across Gemini 3.x models, then $14 per 1,000 requests.
A pricing table alone does not capture the cost of a long live session. Google’s Live API documentation says billing is based on token usage inside the active session context. As conversation history grows, previous context can be processed again on later turns. Developers should use context-window compression where appropriate instead of assuming that a 30-minute conversation costs exactly 30 times a one-minute exchange.
There is another operational detail: because proactive audio is always enabled on Gemini 3.8 Live and Extended Thinking, input usage can continue while the system is listening. If transcription is enabled, transcription output can add text-token charges on top of audio usage.
Gemini 3.8 Live vs Extended Thinking
The practical choice is less about which model is “better” and more about whether the task benefits from reasoning that continues behind the conversation.
Choose Gemini 3.8 Live when latency matters most
Use the standard model for customer-service triage, voice search, language practice, interactive assistants, smart-device commands, and other cases where the next spoken response should arrive quickly. It supports both blocking and non-blocking tools, which also makes it simpler for applications migrating from older Live API designs.
Choose Extended Thinking when the task is genuinely multi-step
Extended Thinking is designed for workflows such as technical diagnostics, travel planning, multi-source data retrieval, and complex tutoring. It can speak intermediate progress updates while it continues reasoning or waiting on asynchronous tools.
The tradeoff is more client complexity. Extended Thinking requires non-blocking tool declarations, and the application must keep listening after an utterance completes until interaction_status reports IDLE.
What the models can and cannot do
Google’s current model cards list several important limits. Both models support real-time multimodal input, function calling, Search grounding, and audio output. They do not currently support context caching, code execution, file search, image generation, structured outputs, URL context, or Batch API execution.
For Extended Thinking specifically, function calls must be asynchronous and non-blocking. This means an existing integration that assumes synchronous tools may need architectural changes rather than a one-line model swap.
Availability: API, Gemini app, Search, Workspace, and enterprise
Google’s September 17 rollout details separate developer, consumer, and enterprise access:
- Developers: both models are available through the Gemini API and Google AI Studio.
- Gemini app: Gemini 3.8 Live Extended Thinking is rolling out in Gemini Live.
- Search: Gemini 3.8 Live powers Search Live experiences.
- Google Workspace: Extended Thinking is available to Google AI Pro and Ultra subscribers in Docs, and to Google AI subscribers in Gmail and Keep, according to Google’s launch post.
- Enterprise: Gemini 3.8 Live with Live Avatar is generally available in Gemini Enterprise for production use, with U.S. and EU endpoints and provisioned throughput options. Gemini 3.8 Live Extended Thinking remains in private preview.
If your main question is which consumer Google AI subscription includes the surrounding Gemini features, see our Gemini pricing guide. For the app-by-app business experience in Gmail, Docs, Sheets, Meet, and Drive, see Gemini in Google Workspace.
Why this matters for voice agents
Voice AI has traditionally forced developers to choose between responsiveness and deeper reasoning. A cascaded stack might combine speech recognition, a text model, tools, and speech synthesis, but every extra stage adds latency and state-management overhead.
Gemini 3.8 Live is Google’s attempt to make that stack more native. The standard model is built for fast multimodal dialogue, while Extended Thinking adds a background reasoning loop that can remain conversational during longer operations.
That fits the broader shift described in our AI agents guide: models are becoming orchestration layers that can reason, call tools, and maintain state rather than simply answer one prompt at a time. For production systems, however, the usual agent controls still matter—permission boundaries, tool allowlists, cost ceilings, logging, and human approval for high-risk actions.
Migration checklist from Gemini 3.1 Flash Live
- Change the model string to
gemini-3.8-livefor the lowest-risk upgrade path. - Remove
thinking_levelorthinking_configwhen using standard Gemini 3.8 Live. - Test asynchronous function calling instead of assuming every tool blocks the conversation.
- If moving to Extended Thinking, change client state handling to use
interaction_status. - Declare Extended Thinking tools with
NON_BLOCKINGbehavior. - Re-test interruption handling, long-session context behavior, transcription cost, and idle detection before production rollout.
Bottom line
Gemini 3.8 Live is the safer default for developers who want a modern low-latency voice model without redesigning every part of the session lifecycle. Gemini 3.8 Live Extended Thinking is the more interesting choice when the agent must keep a natural conversation going while it plans, reasons, or waits on multiple tools.
The headline API rates are competitive, but production cost depends on more than minutes of audio. Persistent context, proactive listening, transcription, Search grounding, and multi-step tool use can all change the real bill. Treat the published per-minute figures as a starting point for capacity planning, not a complete voice-agent TCO model.
Primary sources
Source check: September 24, 2026.
- Google: Introducing Gemini 3.8 Live and 3.8 Live Extended Thinking
- Google AI for Developers: Gemini 3.8 Live model documentation
- Google AI for Developers: Gemini 3.8 Live Extended Thinking model documentation
- Google AI for Developers: Thinking in the Live API
- Google AI for Developers: Live API best practices
- Google AI for Developers: Gemini API pricing
- Google Cloud: Gemini 3.8 Live with Live Avatar is generally available
- Google Cloud: Gemini 3.8 Live model and enterprise availability
- Google: Introducing Gemini 3.8 Live with Live Avatar
AI-XBlog Weekly Brief
Keep up with AI that actually works
Join the AI-XBlog Weekly Brief for major AI updates, practical workflows, useful tools, and editor’s picks. No daily noise.
Reader discussion
Join the discussion
Have you tried this tool or workflow? Share your experience, corrections, or questions. Useful reader feedback may help us improve this article.
All comments are reviewed before publication. Your email address will not be published. Promotional links and low-value spam are removed.
