Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS are Google’s new production text-to-speech models for expressive narration, custom voice design, dubbing, and voice-agent pipelines. Google announced both on September 23, 2026 and made them available to developers through the Gemini API and Google AI Studio.
The important change is bigger than a quality refresh. Google is splitting TTS into two clear tiers: Flash for maximum voice fidelity and creative control, and Flash-Lite for high-throughput, lower-cost generation. The release also adds persistent custom voices, 30-second voice replication with consent verification, structured per-turn direction, and materially lower audio-output pricing than the older Gemini 3.1 Flash TTS Preview.
This guide focuses on the details that matter for developers and buyers: model IDs, API pricing, approximate per-minute output cost, supported languages, voice replication limits, migration changes, availability, and when to choose Flash versus Flash-Lite. It is based on Google’s official launch post and developer documentation; it is not a hands-on listening benchmark.
Gemini 3.8 TTS at a glance
| Model | Model ID | Best fit | Languages |
|---|---|---|---|
| Gemini 3.8 Flash TTS | gemini-3.8-flash-tts |
Studio narration, audiobooks, complex dialogue, difficult pronunciation, creative voice design | 130 |
| Gemini 3.8 Flash-Lite TTS | gemini-3.8-flash-lite-tts |
High-volume generation, dubbing, read-aloud features, voice-agent cascades | 101 |
Both models accept text and output audio. Google’s model documentation lists an 8,192-token input limit and a 16,384-token output limit. Both support caching, Batch API, Flex inference, and Priority inference. They do not support function calling, code execution, Search grounding, structured outputs, or the Live API because they are dedicated speech-generation models rather than general-purpose conversational models.
Gemini 3.8 TTS pricing
Google’s current Gemini Developer API pricing has two date windows: lower rates through December 31, 2026, followed by scheduled higher rates beginning January 1, 2027.
| Model | Text input | Audio output | Approx. output cost/min* |
|---|---|---|---|
| Gemini 3.8 Flash TTS | $0.50 / 1M tokens | $9.00 / 1M audio tokens | ~$0.0135 |
| Gemini 3.8 Flash-Lite TTS | $0.50 / 1M tokens | $6.00 / 1M audio tokens | ~$0.0090 |
*Google states that audio output uses 25 audio tokens per second. That is roughly 1,500 audio tokens per minute, so the figures above estimate output cost only and exclude text input, caching, storage, or other workload-specific charges.
Starting January 1, 2027, Google’s published standard rates rise to $1.00 / 1M text tokens for both models, $18 / 1M audio tokens for Flash TTS, and $12 / 1M audio tokens for Flash-Lite TTS.
Batch and Flex inference are cheaper. Through December 31, 2026, both use $0.25 / 1M text input; audio output is $4.50 / 1M for Flash and $3.00 / 1M for Flash-Lite. Google also lists a free tier for standard and priority access, while Batch and Flex do not have a free tier.
How much cheaper is this than Gemini 3.1 Flash TTS Preview?
The older gemini-3.1-flash-tts-preview is currently listed at $1.00 / 1M text input and $20 / 1M audio output for standard paid usage. At the current 2026 rates, Gemini 3.8 Flash TTS cuts the published audio-output rate from $20 to $9, while Flash-Lite cuts it to $6. Google explicitly positions Flash-Lite as the recommended replacement for the 3.1 preview model.
Do not treat that as a permanent 2027 discount, however. Google has already published higher rates that begin January 1, 2027, so production budgets should model both periods.
What changed beyond price
1. Voice design moves from presets to custom personas
Gemini 3.8 Flash TTS can create new voice personas from natural-language descriptions. Google’s launch materials say developers can control role, accent, and vocal characteristics, while the developer docs expose a Voice design workflow that can save a created voice and return a reusable voice_... ID.
Google also says Flash TTS can access more than 2,000 production-ready voices, while custom voices can be saved for consistent reuse across longer projects.
2. Voice replication uses a 30-second reference sample
Google says the new TTS system can recreate a consistent vocal profile from a 30-second audio sample when the user has the right to use that voice. The process includes consent verification, and generated audio is protected with SynthID watermarking; Google also references C2PA credentials for voice-replication workflows.
This feature is not available everywhere. Google’s current launch note says voice replication through AI Studio is unavailable in Illinois, Texas, the EEA, the UK, Switzerland, and India.
3. Structured speech metadata changes prompting
The 3.8 TTS schema treats the text payload as the transcript to speak. Sustained delivery instructions such as speaker identity or style should move into structured speech_metadata instead of being embedded in the text. Point-in-time vocal events such as <laugh>, <sigh>, and short pauses can remain inline.
This matters for migration because prompts like “Say cheerfully: Hello” may be spoken literally if left in the transcript. Existing 3.1 TTS integrations should move delivery instructions into metadata before switching models.
4. Unary output now defaults to WAV
Google’s migration guide says Gemini 3.8 TTS returns standard WAV audio by default for unary requests. Older Gemini TTS integrations often returned headerless raw PCM. If your application previously wrapped raw bytes in a WAV header, remove that extra wrapper or explicitly request a raw format such as audio/l16, mu-law, or A-law.
Flash vs Flash-Lite: which should you choose?
Choose Gemini 3.8 Flash TTS when voice quality is the product
Flash is the stronger fit for audiobooks, premium narration, games, scripted dialogue, difficult regional accents, pronunciation-sensitive work, and character-heavy audio. Google positions it as the flagship creative tier, with deeper acting nuance and broader language coverage than Flash-Lite.
Choose Gemini 3.8 Flash-Lite TTS when throughput and cost matter more
Flash-Lite is built for bulk dubbing, read-aloud features, large content libraries, notification systems, and cascaded voice agents where speech synthesis is one stage of a larger pipeline. It uses the same 3.8 TTS schema, so applications can switch between the two models with a model-name change after testing quality and cost.
If you need a model that can listen, reason, call tools, and speak in a live conversation, these TTS models are not the same product as Google’s real-time voice stack. See our Gemini 3.8 Live API guide for the audio-to-audio models built for interactive agents.
Availability: API, AI Studio, Notebook, Vids, and Enterprise
- Developers: both Gemini 3.8 TTS models are rolling out through the Gemini API and Google AI Studio.
- Gemini Notebook: Google says Gemini 3.8 Flash TTS is rolling out for users there.
- Google Vids: Gemini 3.8 Flash-Lite TTS is rolling out for users.
- Gemini Enterprise: API access is listed as coming soon for both models.
For subscription-level questions around Google’s consumer AI plans, see our Gemini pricing guide.
Safety and provenance
Voice replication creates obvious impersonation risk, so Google is pairing it with several controls. The company says voice replication requires a verbal consent recording from the voice owner that matches the reference speaker. Generated Gemini Audio clips receive SynthID watermarking, and Google says voice-replication workflows also use C2PA credentials.
Those controls reduce risk but do not remove the need for application-level safeguards. Products that generate or publish synthetic voices should still consider user consent records, abuse monitoring, disclosure, access controls, and high-risk impersonation policies.
Migration checklist from Gemini 3.1 Flash TTS Preview
- Choose
gemini-3.8-flash-lite-ttsfor the closest cost-efficient replacement, orgemini-3.8-flash-ttsfor higher creative fidelity. - Move speaker and sustained style instructions into
speech_metadata. - Keep only point-in-time vocal events such as laughs or pauses inline in the transcript.
- Specify a speaker on every turn in multi-speaker requests.
- Remove any manual WAV-header wrapping if you accept the new default WAV output.
- Re-test output format, pronunciation, voice identity, latency, and cost before production rollout.
- If you use voice replication, verify regional availability and consent requirements.
Practical implications
The most consequential part of this release is that Google is lowering the cost of high-quality speech generation while simultaneously making custom voice creation more programmable. That makes a wider set of workloads practical: localized video, long-form narration, personalized read-aloud, multilingual support, and voice-agent backends can all use one Gemini speech stack without depending entirely on a fixed preset library.
The tradeoff is complexity around identity and governance. Custom voice design is easy to scale; voice replication is sensitive by definition. For teams building autonomous or semi-autonomous voice workflows, combine these audio controls with the permission and review patterns in our AI agents guide rather than treating TTS as an isolated media feature.
Bottom line
Gemini 3.8 Flash TTS is the better choice when expressive quality, character consistency, accent control, and long-form narration matter most. Gemini 3.8 Flash-Lite TTS is the practical default for high-volume production and cost-sensitive speech pipelines.
The current 2026 API rates are especially aggressive compared with Gemini 3.1 Flash TTS Preview, but Google has already published higher rates for January 2027. If you are budgeting a production rollout, model both price periods and benchmark your own scripts, languages, and voices before committing volume.
Primary sources
Source check: September 24, 2026.
- Google: Gemini 3.8 text-to-speech says hello
- Google AI for Developers: Gemini 3.8 Flash TTS model documentation
- Google AI for Developers: Gemini 3.8 Flash-Lite TTS model documentation
- Google AI for Developers: Gemini Developer API pricing
- Google AI for Developers: Voice design
AI-XBlog Weekly Brief
Keep up with AI that actually works
Join the AI-XBlog Weekly Brief for major AI updates, practical workflows, useful tools, and editor’s picks. No daily noise.
Reader discussion
Join the discussion
Have you tried this tool or workflow? Share your experience, corrections, or questions. Useful reader feedback may help us improve this article.
All comments are reviewed before publication. Your email address will not be published. Promotional links and low-value spam are removed.
