Nano Banana Pro prompted by THE DECODER
Key Points
- Google has released two new text-to-speech models, Gemini 3.8 Flash TTS and Flash-Lite TTS, which support more than 100 languages. Flash TTS can create new voices from text descriptions, and a voice cloning feature builds voice profiles from 30-second audio samples.
- Flash TTS is aimed at creative uses such as podcasts, audiobooks, and game characters, while Flash-Lite TTS is designed for low-cost speech generation at scale for dubbing, audio content, and voice agents.
- Both models support stage directions for each line, two-voice dialogue, and nonverbal sounds like laughter and sighs. They're rolling out through the Gemini API and Google AI Studio, and Gemini Enterprise API access will follow.
Google is introducing Gemini 3.8 Flash TTS and Flash-Lite TTS, two new models for speech generation. Flash TTS can create new voices from text descriptions, and both models support more than 100 languages and let users add stage directions to individual lines of dialogue.
Gemini 3.8 Flash TTS is designed for creative projects such as game characters, audiobooks, and podcasts, while Gemini 3.8 Flash-Lite TTS focuses on low-cost speech generation at scale for dubbing, audio content, and voice agents, according to Google.
Flash TTS lets users create voices from text descriptions
With Gemini 3.8 Flash TTS, users can design voices from scratch. According to Google, a text prompt can define a voice's role, accent, and vocal traits across a wide range of languages and dialects. For users who don't want to start from zero, Google offers a library of more than 2,000 preset voices, including regional variants such as Mexican Spanish, Quebec French, and Scottish English.
A voice cloning feature can build a voice profile from a 30-second audio sample. To use it, the person whose voice is being cloned has to record a spoken statement of consent, and the voice in that recording must match the sample. Every clip the Gemini audio models generate carries an inaudible SynthID watermark to help detect AI-generated speech, according to Google.
Google has also announced "Voice Remixing," a feature that will let users adjust the timbre, pitch, tempo, and accent of library voices, but it isn't available yet.
Script directions give users control over dialogue and delivery
Both models let users write directions for each line or have the model interpret script cues on its own. Google says the models can generate hours of audio with minimal "speaker drift," meaning the voice barely changes over time.
A two-voice mode generates dialogue from a single script while keeping the voices distinct, according to the company. Users can also script laughter, sighs, and sounds like "mhm" to place reactions and pauses exactly where they want them.
In two of my own tests, I used a preset voice with a style prompt asking it to imitate an annoyed Berliner speaking English with a thick German accent. The style controls produced a convincing accent and intonation, but both tests had a high-pitched whine in the background at some points. In one of them, the voice also changed at the end of the clip.
Style prompt: A native German man from Berlin speaking English as a foreign language, with a thick, unmistakable German accent. He is clearly not a native English speaker: he pronounces English words the German way, applying German rhythm and intonation to English sentences. "Th" becomes "z" or "d" ("ze," "sink," "dat"), "w" becomes "v" ("vat," "vell"), and final consonants become harder ("goot," "bat" for "bad"). The "r" is guttural and throaty, never the English "r." Vowels are flat and short, lacking the softness found in American or British English.
The voice is nasal and slightly strained, in the mid-range, with a raspy quiver and a Kermit-like wobble, but grittier and less puppet-like. It sounds like a tired Berliner at 2 a.m.
Delivery: quick, clipped, choppy. He hesitates briefly when searching for an English word and sometimes inserts the German word flatly without translating it. Occasional exasperated upward pitch breaks. Dry, sardonic, unimpressed, and slightly annoyed that he has to explain anything at all—and even more annoyed that he has to do it in English.

Google starts rolling out both models, with enterprise API access to follow
Google is rolling out both models through the Gemini API and Google AI Studio. Flash TTS is also available in Gemini Notebook, while Flash-Lite TTS is available in Google Vids. Google says access through the Gemini Enterprise API will follow soon for both models.
Developer platforms including Agora, LiveKit, Pipecat, and Vercel already support integration through the Gemini API. Google hasn't listed regional endpoints for the new models yet, though earlier TTS models offered EU data processing.
According to Google, data from the free tier is used to improve its products, while data from the paid tier isn't. Paid pricing is listed in US dollars per million tokens, with text tokens billed for input and audio tokens for output.
| Billing | Flash TTS (through the end of 2026) | Flash TTS (starting in 2027) | Flash-Lite TTS (through the end of 2026) | Flash-Lite TTS (starting in 2027) |
|---|---|---|---|---|
| Text Input | $0.50 | $1.00 | $0.50 | $1.00 |
| Audio Output | $9.00 | $18.00 | $6.00 | $12.00 |
Google says one second of generated audio equals 25 audio tokens, which puts an hour at 90,000 tokens. That means an hour of audio output costs $0.81 with Flash TTS and $0.54 with Flash-Lite TTS through the end of 2026. On January 1, 2027, those costs rise to $1.62 and $1.08, with text input billed separately.