Gemini 3.8 TTS: Pricing, Model IDs, Voice Replication & API Guide

By

· Published

· Updated

·

, ,
Gemini 3.8 Flash TTS and Flash-Lite TTS editorial cover with custom voice, API, and multilingual speech visualization

Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS are Google’s new production text-to-speech models for expressive narration, custom voice design, dubbing, and voice-agent pipelines. Google announced both on September 23, 2026 and made them available to developers through the Gemini API and Google AI Studio.

The important change is bigger than a quality refresh. Google is splitting TTS into two clear tiers: Flash for maximum voice fidelity and creative control, and Flash-Lite for high-throughput, lower-cost generation. The release also adds persistent custom voices, 30-second voice replication with consent verification, structured per-turn direction, and materially lower audio-output pricing than the older Gemini 3.1 Flash TTS Preview.

This guide focuses on the details that matter for developers and buyers: model IDs, API pricing, approximate per-minute output cost, supported languages, voice replication limits, migration changes, availability, and when to choose Flash versus Flash-Lite. It is based on Google’s official launch post and developer documentation; it is not a hands-on listening benchmark.

Gemini 3.8 TTS at a glance

Model Model ID Best fit Languages
Gemini 3.8 Flash TTS gemini-3.8-flash-tts Studio narration, audiobooks, complex dialogue, difficult pronunciation, creative voice design 130
Gemini 3.8 Flash-Lite TTS gemini-3.8-flash-lite-tts High-volume generation, dubbing, read-aloud features, voice-agent cascades 101

Both models accept text and output audio. Google’s model documentation lists an 8,192-token input limit and a 16,384-token output limit. Both support caching, Batch API, Flex inference, and Priority inference. They do not support function calling, code execution, Search grounding, structured outputs, or the Live API because they are dedicated speech-generation models rather than general-purpose conversational models.

Gemini 3.8 TTS pricing

Google’s current Gemini Developer API pricing has two date windows: lower rates through December 31, 2026, followed by scheduled higher rates beginning January 1, 2027.

Model Text input Audio output Approx. output cost/min*
Gemini 3.8 Flash TTS $0.50 / 1M tokens $9.00 / 1M audio tokens ~$0.0135
Gemini 3.8 Flash-Lite TTS $0.50 / 1M tokens $6.00 / 1M audio tokens ~$0.0090

*Google states that audio output uses 25 audio tokens per second. That is roughly 1,500 audio tokens per minute, so the figures above estimate output cost only and exclude text input, caching, storage, or other workload-specific charges.

Starting January 1, 2027, Google’s published standard rates rise to $1.00 / 1M text tokens for both models, $18 / 1M audio tokens for Flash TTS, and $12 / 1M audio tokens for Flash-Lite TTS.

Batch and Flex inference are cheaper. Through December 31, 2026, both use $0.25 / 1M text input; audio output is $4.50 / 1M for Flash and $3.00 / 1M for Flash-Lite. Google also lists a free tier for standard and priority access, while Batch and Flex do not have a free tier.

How much cheaper is this than Gemini 3.1 Flash TTS Preview?

The older gemini-3.1-flash-tts-preview is currently listed at $1.00 / 1M text input and $20 / 1M audio output for standard paid usage. At the current 2026 rates, Gemini 3.8 Flash TTS cuts the published audio-output rate from $20 to $9, while Flash-Lite cuts it to $6. Google explicitly positions Flash-Lite as the recommended replacement for the 3.1 preview model.

Do not treat that as a permanent 2027 discount, however. Google has already published higher rates that begin January 1, 2027, so production budgets should model both periods.

What changed beyond price

1. Voice design moves from presets to custom personas

Gemini 3.8 Flash TTS can create new voice personas from natural-language descriptions. Google’s launch materials say developers can control role, accent, and vocal characteristics, while the developer docs expose a Voice design workflow that can save a created voice and return a reusable voice_... ID.

Google also says Flash TTS can access more than 2,000 production-ready voices, while custom voices can be saved for consistent reuse across longer projects.

2. Voice replication uses a 30-second reference sample

Google says the new TTS system can recreate a consistent vocal profile from a 30-second audio sample when the user has the right to use that voice. The process includes consent verification, and generated audio is protected with SynthID watermarking; Google also references C2PA credentials for voice-replication workflows.

This feature is not available everywhere. Google’s current launch note says voice replication through AI Studio is unavailable in Illinois, Texas, the EEA, the UK, Switzerland, and India.

3. Structured speech metadata changes prompting

The 3.8 TTS schema treats the text payload as the transcript to speak. Sustained delivery instructions such as speaker identity or style should move into structured speech_metadata instead of being embedded in the text. Point-in-time vocal events such as <laugh>, <sigh>, and short pauses can remain inline.

This matters for migration because prompts like “Say cheerfully: Hello” may be spoken literally if left in the transcript. Existing 3.1 TTS integrations should move delivery instructions into metadata before switching models.

4. Unary output now defaults to WAV

Google’s migration guide says Gemini 3.8 TTS returns standard WAV audio by default for unary requests. Older Gemini TTS integrations often returned headerless raw PCM. If your application previously wrapped raw bytes in a WAV header, remove that extra wrapper or explicitly request a raw format such as audio/l16, mu-law, or A-law.

Flash vs Flash-Lite: which should you choose?

Choose Gemini 3.8 Flash TTS when voice quality is the product

Flash is the stronger fit for audiobooks, premium narration, games, scripted dialogue, difficult regional accents, pronunciation-sensitive work, and character-heavy audio. Google positions it as the flagship creative tier, with deeper acting nuance and broader language coverage than Flash-Lite.

Choose Gemini 3.8 Flash-Lite TTS when throughput and cost matter more

Flash-Lite is built for bulk dubbing, read-aloud features, large content libraries, notification systems, and cascaded voice agents where speech synthesis is one stage of a larger pipeline. It uses the same 3.8 TTS schema, so applications can switch between the two models with a model-name change after testing quality and cost.

If you need a model that can listen, reason, call tools, and speak in a live conversation, these TTS models are not the same product as Google’s real-time voice stack. See our Gemini 3.8 Live API guide for the audio-to-audio models built for interactive agents.

Availability: API, AI Studio, Notebook, Vids, and Enterprise

  • Developers: both Gemini 3.8 TTS models are rolling out through the Gemini API and Google AI Studio.
  • Gemini Notebook: Google says Gemini 3.8 Flash TTS is rolling out for users there.
  • Google Vids: Gemini 3.8 Flash-Lite TTS is rolling out for users.
  • Gemini Enterprise: API access is listed as coming soon for both models.

For subscription-level questions around Google’s consumer AI plans, see our Gemini pricing guide.

Safety and provenance

Voice replication creates obvious impersonation risk, so Google is pairing it with several controls. The company says voice replication requires a verbal consent recording from the voice owner that matches the reference speaker. Generated Gemini Audio clips receive SynthID watermarking, and Google says voice-replication workflows also use C2PA credentials.

Those controls reduce risk but do not remove the need for application-level safeguards. Products that generate or publish synthetic voices should still consider user consent records, abuse monitoring, disclosure, access controls, and high-risk impersonation policies.

Migration checklist from Gemini 3.1 Flash TTS Preview

  1. Choose gemini-3.8-flash-lite-tts for the closest cost-efficient replacement, or gemini-3.8-flash-tts for higher creative fidelity.
  2. Move speaker and sustained style instructions into speech_metadata.
  3. Keep only point-in-time vocal events such as laughs or pauses inline in the transcript.
  4. Specify a speaker on every turn in multi-speaker requests.
  5. Remove any manual WAV-header wrapping if you accept the new default WAV output.
  6. Re-test output format, pronunciation, voice identity, latency, and cost before production rollout.
  7. If you use voice replication, verify regional availability and consent requirements.

Practical implications

The most consequential part of this release is that Google is lowering the cost of high-quality speech generation while simultaneously making custom voice creation more programmable. That makes a wider set of workloads practical: localized video, long-form narration, personalized read-aloud, multilingual support, and voice-agent backends can all use one Gemini speech stack without depending entirely on a fixed preset library.

The tradeoff is complexity around identity and governance. Custom voice design is easy to scale; voice replication is sensitive by definition. For teams building autonomous or semi-autonomous voice workflows, combine these audio controls with the permission and review patterns in our AI agents guide rather than treating TTS as an isolated media feature.

Bottom line

Gemini 3.8 Flash TTS is the better choice when expressive quality, character consistency, accent control, and long-form narration matter most. Gemini 3.8 Flash-Lite TTS is the practical default for high-volume production and cost-sensitive speech pipelines.

The current 2026 API rates are especially aggressive compared with Gemini 3.1 Flash TTS Preview, but Google has already published higher rates for January 2027. If you are budgeting a production rollout, model both price periods and benchmark your own scripts, languages, and voices before committing volume.

Primary sources

Source check: September 24, 2026.

AI-XBlog Weekly Brief

Keep up with AI that actually works

Join the AI-XBlog Weekly Brief for major AI updates, practical workflows, useful tools, and editor’s picks. No daily noise.

Double opt-in. Unsubscribe anytime. See our Privacy Policy.

Reader discussion

Join the discussion

Have you tried this tool or workflow? Share your experience, corrections, or questions. Useful reader feedback may help us improve this article.

All comments are reviewed before publication. Your email address will not be published. Promotional links and low-value spam are removed.

Add a comment

Comments are moderated to keep the discussion useful and trustworthy.

About the author

AI-XBlog Editorial Team researches and maintains practical coverage of AI tools, automation, agents and applied artificial intelligence. We prioritize primary sources, clear evidence and useful real-world guidance.

Editorial Policy · Review Methodology · Corrections Policy