Gemini 3.8 Flash TTS 和 Gemini 3.8 Flash-Lite TTS 是我们迄今最具表现力的音频生成模型。可在 Google AI Studio、Gemini API、Gemini Enterprise、Gemini Notebook 和 Google Vids 中生成自定义角色语音并指导场景对话。
Leland Rechis
产品组经理
Alan Cowen
研究科学总监,代表 Gemini Audio 团队
今天,我们为 Gemini 家族推出两款全新的文本转语音模型,将语音生成从静态预设转变为动态创意工作室。这些模型让创作者、开发者和企业能够打造更丰富、更具表现力的音频体验,同时为 Gemini Notebook 和 Google Vids 等产品带来更出色的用户体验。
- Gemini 3.8 Flash TTS:专为深度创意指导和角色设计而打造。使用自然语言提示词从零开始创造全新语音,让角色在游戏、沉浸式有声书、播客和互动媒体中栩栩如生。逐行指导每一段表演,对表演提示、节奏、方言转换和应和语进行精细控制。
- Gemini 3.8 Flash-Lite TTS: 专为高吞吐量、高性价比的规模化场景打造。针对大批量配音、音频内容创作以及富有表现力的语音智能体进行了优化,可对语调、节奏和表达细节进行精细控制。
这些模型补充了我们快速壮大的 Gemini Audio 家族,此前已推出 3.5 Live Translate、3.5 Transcribe、3.8 Live 和 3.8 Live Extended Thinking。
创建并自定义你自己的声音
从 30 种原创声音扩展到无限的声音库。无论你需要一个完全原创的角色声音,还是一个保持一致的品牌代言人声音,我们的 3.8 Flash TTS 模型都能驱动一整套声音工作室。这让你能够为每一个时刻创建并使用富有表现力、听起来自然的声音,同时赋能开发者和企业轻松构建自定义音频体验。
- 生成式声音设计: 借助 Gemini 3.8 Flash TTS,通过自然语言提示词,在 100 多种语言和方言中自定义角色、口音和声音特征,从零开始打造定制声音——无论你是要让一条戏剧性的喷火龙活灵活现,还是打造一位带有独特地域韵律的魅力旁白者。
听听 Gemini 3.8 Flash TTS 如何生成一位来自墨尔本、充满活力的 DJ 声音。
听听 Gemini 3.8 Flash TTS 如何生成一个极其单薄、单调的机器人声音。
听听 Gemini 3.8 Flash TTS 如何让一条日本龙栩栩如生。
- 海量语音库:可访问 2,000 多种可直接用于生产的语音,语言覆盖广泛——包括墨西哥西班牙语、魁北克法语和苏格兰英语等地区变体。
- 语音复刻:只需一段 30 秒的音频样本,即可用你自己的声音或你拥有使用权的他人声音,重建一致的语音特征,并以内置的同意验证、SynthID 水印和 C2PA 凭证作为支撑,从而同时保护开发者及其配音人才。
- 保存与扩展:保存并管理你设计的自定义语音,以确保在持续进行的项目中保持一致的性能并最大限度减少漂移。
- 语音混音:即将推出,从我们的语音库中挑选一个声音,并微调音色、音高、语速和口音。使用提示词来调整特征(例如“加入微妙的美国南部口音”或“让表达更柔和”)。
逐行指导表演
选定语音后,两个 TTS 模型都能让你精确控制每一行的表达方式。
- 逐行指导表演:你可以自己撰写舞台指示,也可以让 Gemini 借助自然的脚本提示来引导表达——从冷静的客服智能体到低声细语的悬疑场景,皆可实现。
聆听 Gemini 3.8 Flash TTS 如何为交互式语音智能体实现自然、极具表现力的对话。
观看并聆听 Gemini 3.8 Flash TTS 如何利用精细的脚本控制,打造深度沉浸、引人入胜的音频体验。
- 长音频生成:在长达数小时的连续音频中保持高音质、自然的节奏和角色音色,说话人漂移极小——非常适合播客和有声书。
- 原生双说话人场景编排:从单一脚本无缝指导多轮对话——无论是播客还是戏剧化叙事——同时让两个声音清晰分离,并具备自然的对话轮替。
- 脚本化的人声爆发与附和应答:使用非语言提示(如 <laughs>、<sigh>、<gasp>)和积极倾听的插入语(如 |mhm| 或 |yeah|)添加逼真的对话质感,以实现精准的喜剧节奏和反应节拍。
了解 Gemini 3.8 Flash TTS 如何从零开始将自然语言提示词转化为定制的语音人设。
观看 Gemini 3.8 Flash TTS 如何让创作者设计自定义场景,将动画对话生动呈现。
看看 Gemini 3.8 Flash TTS 如何将脚本转化为完整演绎的对话场景,让创作者能够指导语音表达,并实现自然的轮流对话。
获取专为全球规模打造的表现力丰富的高质量语音生成能力
Gemini 3.8 Flash TTS 提供领先的语音定制能力,在 Hume AI 的语音设计基准测试中夺得总分第一(71.4),并在口音建模方面同样领先(60.8)。
Gemini 3.8 Flash TTS 和 Gemini 3.8 Flash-Lite TTS 能够实现真正富有表现力的演绎,同时不牺牲可靠性,并分别在 Hume AI 的总体质量指数中夺得第一和第二名。与 Gemini 3.1 Flash TTS 相比,该模型在长篇内容和双人剧本控制等广泛用例上展现出重大改进。
在 Voice Arena 的盲测人类偏好评估中,Gemini 3.8 Flash 和 Flash-Lite TTS 在日语、巴西葡萄牙语、越南语、现代标准阿拉伯语(MSA)、墨西哥西班牙语和印地语等关键全球语言中,于竞争对手之间占据领先位置。这些模型支持超过 100 种语言,赋能创作者、开发者和企业,在全球范围内构建高质量的多语言语音体验。
以信任、同意和透明为基础进行构建
我们在构建语音创建与复刻能力时,配备了严格的防护措施,以帮助保护配音人才、尊重身份,并确保内容透明度。对于语音复刻,我们的系统采用同意验证机制:用户必须提供一段来自声音所有者的口头同意录音,且该录音需与参考说话人匹配,之后才能创建语音。
更广泛地说,我们的 Gemini Audio 模型生成的每一段音频片段都带有 SynthID 水印。这种不可感知的水印被直接织入音频输出之中,确保 AI 生成的语音始终可被检测,从而帮助防止虚假信息。如需进一步了解我们在安全与责任方面的做法,请查阅 模型卡。
试用我们全新的 Google AI Studio 音频演练场
从今天起,开发者可以体验这些全新的语音生成能力,在Google AI Studio中。它被打造为一个语音设计工作区,你可以从零开始通过提示词创造全新的声音身份,或复刻你自己的声音1,然后将它们直接带入双说话人剧本编辑器,逐行指导语音的表达方式。
在 Google AI Studio 中试用语音复刻。
轻松部署高性能语音交互界面
通过使用 Gemini API,Agora、LiveKit、Pipecat、Vercel 等开发者平台让开发者能够轻松构建和部署高性能的语音生成体验。
我们正与 Figma、HeyGen、Linguana、Wondercraft、99.co 和 Ollang 等公司合作,它们正在集成我们最新的 TTS 模型,以帮助加速全球配音、以细腻的地区口音本地化媒体内容,并大规模驱动对话式语音智能体。
开始使用我们最新的 Gemini Audio 模型:
Gemini 3.8 Flash TTS 从今天起开始推送:
- 面向开发者:在 Gemini API 和 Google AI Studio 中
- 面向企业:即将通过 Gemini Enterprise 中的 API 提供
- 面向所有人:在 Gemini Notebook 中。
Gemini 3.8 Flash-Lite TTS 从今天起开始推送:
- 面向开发者:在 Gemini API 和 Google AI Studio 中
- 面向企业:即将通过 Gemini Enterprise 中的 API 提供
- 面向所有人:在 Google Vids 中
Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS are our most expressive audio generation models yet. Generate custom character voices and direct scene dialogue across Google AI Studio, Gemini API, Gemini Enterprise, Gemini Notebook, and Google Vids.
Leland Rechis
Group Product Manager
Alan Cowen
Director, Research Science, on Behalf of the Gemini Audio Team
Today, we’re introducing two new text-to-speech models to the Gemini family, transforming voice generation from static presets into a dynamic creative studio. These models enable creators, developers, and enterprises to create richer, more expressive audio experiences, while enabling improved user experiences in products like Gemini Notebook and Google Vids.
- Gemini 3.8 Flash TTS: Built for deep creative direction and character design. Create entirely new voices from scratch using natural language prompts to bring characters to life across gaming, immersive audiobooks, podcasts, and interactive media. Direct every performance line by line with granular control over acting cues, pacing, dialect shifts, and backchanneling.
- Gemini 3.8 Flash-Lite TTS: Built for high-volume, cost-efficient scale. Optimized for high-volume dubbing, audio content creation, and expressive voice agents with fine-grained control over tone, pacing, and expressive nuance.
These models complement our fast-growing Gemini Audio family, following 3.5 Live Translate, 3.5 Transcribe, 3.8 Live, and 3.8 Live Extended Thinking.
Create and customize your own voices
Scale up from 30 original voices to an infinite library. Whether you need an entirely original character voice or a consistent brand ambassador, our 3.8 Flash TTS model powers a full vocal studio. This enables you to create and use expressive, natural-sounding voices for every moment, while empowering developers and enterprises to easily build custom audio experiences.
- Generative voice design: With Gemini 3.8 Flash TTS, create bespoke voices from scratch by customizing role, accent and voice characteristics across more than 100 languages and dialects using natural language prompting — whether you're bringing a dramatic, fire-breathing dragon to life or crafting a charismatic narrator with a distinct regional cadence.
Hear how Gemini 3.8 Flash TTS generates a high-energy DJ voice from Melbourne.
Hear how Gemini 3.8 Flash TTS generates a super-tinny, monotone robot voice.
Hear how Gemini 3.8 Flash TTS brings a Japanese dragon to life.
- Expansive voice library: Access 2,000+ production-ready voices with broad language coverage — including regional varieties like Mexican Spanish, Quebec French, and Scots English.
- Voice replication: Recreate consistent vocal profiles from just a 30-second audio sample of your voice or a voice you have the rights to use, backed by built-in consent verification, SynthID watermarking, and C2PA credentials to protect both developers and their vocal talent.
- Save and scale: Save and manage the custom voices you designed to ensure consistent performance and minimal drift across ongoing projects.
- Voice remixing: Coming soon, pick a voice from our voice library and fine-tune timbre, pitch, pace, and accent. Use prompts to dial in characteristics (e.g. “add subtle Southern US accent” or “soften the delivery”).
Direct the performance, line by line
Once you've selected your voices, both TTS models give you precise control over how each line is delivered.
- Direct performance line by line: Write your own stage directions or let Gemini steer delivery with natural script cues — from a calm customer service agent to a whispered suspense scene.
Hear how Gemini 3.8 Flash TTS enables natural, highly expressive conversations for interactive voice agents.
Watch and hear how Gemini 3.8 Flash TTS uses granular script control to build a deeply engaging, immersive audio experience.
- Long-form generation: Maintain high voice quality, natural pacing, and character timbre across hours of continuous audio with minimal speaker drift — ideal for podcasts and audiobooks.
- Native two-speaker scene staging: Direct multi-turn conversations seamlessly from a single script —whether for a podcast or dramatic storytelling—while keeping both voices distinctly separated with natural conversational turn-taking.
- Scripted vocal bursts & backchanneling: Add realistic conversational texture using non verbal cues (like <laughs>, <sigh>, <gasp> and active-listening interjections (like |mhm| or|yeah|) for precise comedic timing and reaction beats.
See how Gemini 3.8 Flash TTS turns natural language prompts into bespoke vocal personas from scratch.
Watch how Gemini 3.8 Flash TTS enables creators to design custom scenes to bring animated dialogue to life.
See how Gemini 3.8 Flash TTS turns scripts into fully performed dialogue scenes, letting creators direct vocal delivery, and natural turn-taking.
Get expressive high-quality speech generation built for global scale
Gemini 3.8 Flash TTS delivers leading voice customization capabilities, securing the #1 overall spot on Hume AI’s Voice Design Benchmark (71.4) and also leading in accent modeling (60.8).
Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS enable truly expressive performances without sacrificing reliability, also securing the #1 and #2 spots respectively on Hume AI’s Overall Quality Index. The model shows major improvements on a wide range of use cases such as long-form content and dual-speaker screenplay control compared to Gemini 3.1 Flash TTS.
In blind human preference evaluations on Voice Arena, Gemini 3.8 Flash and Flash-Lite TTS secure top positions amongst competitors in key global languages, including Japanese, Brazilian Portuguese, Vietnamese, Modern Standard Arabic (MSA), Mexican Spanish and Hindi. With support for over 100 languages, these models empower creators, developers, and enterprises to build high-quality, multilingual voice experiences worldwide.
Build with trust, consent, and transparency
We built our voice creation and replication capabilities with strict safeguards to help protect voice talent, respect identity, and ensure content transparency. For voice replication our system leverages consent verification: users must provide a verbal consent recording from the voice owner that matches the reference speaker before a voice can be created.
More broadly, every audio clip generated by our Gemini Audio models is watermarked with SynthID. This imperceptible watermark is woven directly into the audio output, ensuring AI-generated speech remains detectable to help prevent misinformation. For more details on our approach to safety and responsibility, review the model card.
Try our new Google AI Studio audio playground
Starting today, developers can experience these new speech generation capabilities in Google AI Studio. Built like a voice design workspace, you can prompt entirely new vocal identities from scratch or replicate your own voice 1 , then bring them directly into a dual-speaker screenplay editor to direct line-by-line delivery.
Try voice replication in Google AI Studio.
Deploy high-performance voice interfaces with ease
By using the Gemini API, developer platforms such as Agora, LiveKit, Pipecat, Vercel enable developers to build and deploy high-performance speech generation experiences with ease.
We’re partnering with companies like Figma, HeyGen, Linguana, Wondercraft, 99.co, and Ollang, who are integrating our latest TTS models to help accelerate global dubbing, localize media with nuanced regional accents, and power conversational voice agents at scale.
Start using our latest Gemini Audio models:
Gemini 3.8 Flash TTS is rolling out starting today:
- For developers: In the Gemini API and Google AI Studio
- For enterprises: Coming soon via API in Gemini Enterprise
- For everyone: In Gemini Notebook.
Gemini 3.8 Flash-Lite TTS is rolling out starting today:
- For developers: In the Gemini API and Google AI Studio
- For enterprises: Coming soon via API in Gemini Enterprise
- For everyone: In Google Vids