- DeepMind released Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS on September 23, 2026.
- DeepMind calls them its most expressive audio generation models yet.
- They support custom character voices and directable scene dialogue.
- The launch escalates the voice-AI race with dedicated speech specialists.
What Happened
Google DeepMind released Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS on September 23, 2026, describing them as “our most expressive audio generation models yet,” with the ability to generate custom character voices and direct scene dialogue, per the announcement.
Why It Matters
Text-to-speech has moved from reading text aloud to performing it — multiple characters, emotional direction, scene-level control. That capability set is the foundation for AI audiobooks, game dialogue, dubbing, and voice agents. Shipping it in Flash and Flash-Lite tiers signals Google wants expressive speech at commodity prices, squeezing dedicated voice specialists such as ElevenLabs from the platform side.
Technical Details
The two tiers trade quality against cost and latency: Flash for expressiveness, Flash-Lite for cheaper high-volume use. “Directing” scene dialogue implies prompt-level control over delivery — pacing, emotion, character switching — rather than a fixed voice list. Voice cloning and safety constraints around custom character voices are the sensitive details to examine as access widens.
Who’s Affected
Developers building voice agents and audio content get a first-party Google option at two price points. Voice-AI specialists face platform competition on their core feature, expressiveness. Creators gain cheaper multi-character audio production.
What’s Next
Independent quality comparisons against ElevenLabs and OpenAI‘s audio models will establish where the new tiers actually land. Pricing and rate limits will determine whether Flash-Lite becomes the default for high-volume speech.