Gemini 3.8 Live API in 2026: Pricing, Model IDs, Extended Thinking, and Availability

By

· Published

· Updated

·

, ,
Gemini 3.8 Live API in 2026 with pricing, model IDs, extended thinking, and availability

Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking are Google’s newest real-time audio models for voice agents, multimodal conversations, and tool-using workflows. Google announced both on September 15, 2026, expanded the rollout details on September 17, and on September 24 made Gemini 3.8 Live with Live Avatar generally available in Gemini Enterprise.

The important change is not just better speech quality. Google is separating two voice-agent workloads that used to be forced into one design: low-latency conversation and deeper multi-step reasoning. The standard Live model prioritizes fast interaction, while Extended Thinking can keep talking, narrate progress, and continue reasoning or calling tools in the background.

If your application needs generated speech rather than a live conversational session, see our Gemini 3.8 TTS API guide for the separate TTS model IDs, pricing, voices and API behavior.

This guide focuses on what developers and buyers actually need to know: model IDs, pricing, availability, migration behavior, technical limits, and the practical difference between the two models. It is based on Google’s official launch post and Gemini API documentation; it is not a hands-on benchmark.

Gemini 3.8 Live at a glance

Model Model ID Best fit Reasoning behavior
Gemini 3.8 Live gemini-3.8-live Low-latency voice agents, direct commands, fast tool calls Interleaved reasoning with a fixed latency profile
Gemini 3.8 Live Extended Thinking gemini-3.8-live-extended-thinking Complex support, booking, diagnostics, multi-tool workflows Configurable background reasoning with low, medium, or high thinking levels

Both models accept text, images, audio, and video. Google’s model documentation lists a 131,072-token input limit and a 65,536-token output limit for both. Both support the Live API, function calling, Search grounding, and audio generation.

What changed from Gemini 3.1 Flash Live?

Gemini 3.8 Live is the direct migration target for many applications currently using gemini-3.1-flash-live-preview. The basic move is to change the model string to gemini-3.8-live, but there are behavioral differences worth planning for.

  • Asynchronous tool use is now central. Gemini 3.8 Live supports non-blocking function execution by default, so a voice agent can keep the conversation moving while a tool runs.
  • Extended Thinking introduces a different session lifecycle. The client must track interaction_status rather than assuming that turnComplete: true means all reasoning and tool work has finished.
  • Thinking configuration differs by model. Standard Gemini 3.8 Live does not accept thinking_level. Extended Thinking supports low, medium, and high background reasoning levels.
  • Proactive audio is permanently enabled on the 3.8 Live models. That matters for both user experience and cost because the API can continue processing audio while it is listening.

Gemini 3.8 Live API pricing

Google currently prices Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking at the same published rates in the Gemini Developer API.

Usage Paid API rate Approximate audio rate
Text input $0.75 per 1M tokens —
Audio input $3.00 per 1M tokens $0.005/minute
Image/video input $1.00 per 1M tokens $0.002/minute
Text output, including thinking tokens $4.50 per 1M tokens —
Audio output $12.00 per 1M tokens $0.018/minute

Google also lists a free tier for these models. For paid accounts, Google Search grounding includes 5,000 search requests per month shared across Gemini 3.x models, then $14 per 1,000 requests.

A pricing table alone does not capture the cost of a long live session. Google’s Live API documentation says billing is based on token usage inside the active session context. As conversation history grows, previous context can be processed again on later turns. Developers should use context-window compression where appropriate instead of assuming that a 30-minute conversation costs exactly 30 times a one-minute exchange.

There is another operational detail: because proactive audio is always enabled on Gemini 3.8 Live and Extended Thinking, input usage can continue while the system is listening. If transcription is enabled, transcription output can add text-token charges on top of audio usage.

Gemini 3.8 Live vs Extended Thinking

The practical choice is less about which model is “better” and more about whether the task benefits from reasoning that continues behind the conversation.

Choose Gemini 3.8 Live when latency matters most

Use the standard model for customer-service triage, voice search, language practice, interactive assistants, smart-device commands, and other cases where the next spoken response should arrive quickly. It supports both blocking and non-blocking tools, which also makes it simpler for applications migrating from older Live API designs.

Choose Extended Thinking when the task is genuinely multi-step

Extended Thinking is designed for workflows such as technical diagnostics, travel planning, multi-source data retrieval, and complex tutoring. It can speak intermediate progress updates while it continues reasoning or waiting on asynchronous tools.

The tradeoff is more client complexity. Extended Thinking requires non-blocking tool declarations, and the application must keep listening after an utterance completes until interaction_status reports IDLE.

What the models can and cannot do

Google’s current model cards list several important limits. Both models support real-time multimodal input, function calling, Search grounding, and audio output. They do not currently support context caching, code execution, file search, image generation, structured outputs, URL context, or Batch API execution.

For Extended Thinking specifically, function calls must be asynchronous and non-blocking. This means an existing integration that assumes synchronous tools may need architectural changes rather than a one-line model swap.

Availability: API, Gemini app, Search, Workspace, and enterprise

Google’s September 17 rollout details separate developer, consumer, and enterprise access:

  • Developers: both models are available through the Gemini API and Google AI Studio.
  • Gemini app: Gemini 3.8 Live Extended Thinking is rolling out in Gemini Live.
  • Search: Gemini 3.8 Live powers Search Live experiences.
  • Google Workspace: Extended Thinking is available to Google AI Pro and Ultra subscribers in Docs, and to Google AI subscribers in Gmail and Keep, according to Google’s launch post.
  • Enterprise: Gemini 3.8 Live with Live Avatar is generally available in Gemini Enterprise for production use, with U.S. and EU endpoints and provisioned throughput options. Gemini 3.8 Live Extended Thinking remains in private preview.

If your main question is which consumer Google AI subscription includes the surrounding Gemini features, see our Gemini pricing guide. For the app-by-app business experience in Gmail, Docs, Sheets, Meet, and Drive, see Gemini in Google Workspace.

Why this matters for voice agents

Voice AI has traditionally forced developers to choose between responsiveness and deeper reasoning. A cascaded stack might combine speech recognition, a text model, tools, and speech synthesis, but every extra stage adds latency and state-management overhead.

Gemini 3.8 Live is Google’s attempt to make that stack more native. The standard model is built for fast multimodal dialogue, while Extended Thinking adds a background reasoning loop that can remain conversational during longer operations.

That fits the broader shift described in our AI agents guide: models are becoming orchestration layers that can reason, call tools, and maintain state rather than simply answer one prompt at a time. For production systems, however, the usual agent controls still matter—permission boundaries, tool allowlists, cost ceilings, logging, and human approval for high-risk actions.

Migration checklist from Gemini 3.1 Flash Live

  1. Change the model string to gemini-3.8-live for the lowest-risk upgrade path.
  2. Remove thinking_level or thinking_config when using standard Gemini 3.8 Live.
  3. Test asynchronous function calling instead of assuming every tool blocks the conversation.
  4. If moving to Extended Thinking, change client state handling to use interaction_status.
  5. Declare Extended Thinking tools with NON_BLOCKING behavior.
  6. Re-test interruption handling, long-session context behavior, transcription cost, and idle detection before production rollout.

Bottom line

Gemini 3.8 Live is the safer default for developers who want a modern low-latency voice model without redesigning every part of the session lifecycle. Gemini 3.8 Live Extended Thinking is the more interesting choice when the agent must keep a natural conversation going while it plans, reasons, or waits on multiple tools.

The headline API rates are competitive, but production cost depends on more than minutes of audio. Persistent context, proactive listening, transcription, Search grounding, and multi-step tool use can all change the real bill. Treat the published per-minute figures as a starting point for capacity planning, not a complete voice-agent TCO model.

Primary sources

Source check: September 24, 2026.

AI-XBlog Weekly Brief

Keep up with AI that actually works

Join the AI-XBlog Weekly Brief for major AI updates, practical workflows, useful tools, and editor’s picks. No daily noise.

Double opt-in. Unsubscribe anytime. See our Privacy Policy.

Reader discussion

Join the discussion

Have you tried this tool or workflow? Share your experience, corrections, or questions. Useful reader feedback may help us improve this article.

All comments are reviewed before publication. Your email address will not be published. Promotional links and low-value spam are removed.

Add a comment

Comments are moderated to keep the discussion useful and trustworthy.

About the author

AI-XBlog Editorial Team researches and maintains practical coverage of AI tools, automation, agents and applied artificial intelligence. We prioritize primary sources, clear evidence and useful real-world guidance.

Editorial Policy · Review Methodology · Corrections Policy