Google 已推出Gemini 3.8 Live 和 Gemini 3.8 Live Extended Thinking,这是其迄今为止最先进的实时对话模型。两者都是为实时语音智能体打造的原生语音到语音模型。它们扩展了 Google 的 Gemini Audio 系列,该系列于上个月通过以下模型得到扩展Gemini 3.5 Transcribe。此次发布针对一个特定的空白:能够进行推理并执行工具、同时不打断对话流程的语音智能体。
它可以部署吗?可以,适用于基于 API 的生产环境使用。这两款模型目前均已在 Gemini Live API 和 Google AI Studio 中上线。它们是托管模型,而非开放权重,因此没有自托管选项。
Google 发布了什么
此次发布包含两款角色定位不同的模型。Gemini 3.8 Live 专为规模化与成本效率打造。它将对话智能与流畅对话及视觉接地能力相结合。Gemini 3.8 Live Extended Thinking 专为高复杂度任务打造。它在说话的同时增强了智能水平与多步推理能力。Google 将两者定位为对级联语音流水线(串联 ASR、LLM 和 TTS)的精简替代方案。
基准测试结果
Gemini 3.8 Live Extended Thinking 以 82.6 分在 Artificial Analysis 的语音到语音质量指数中拿下总榜第一。它在智能体任务完成度上领先,在 τ-Voice 上达到 68.6%,在 Sierra 的 τ-Voice-banking 基准上达到 35.1%。它还在 Big Bench Audio 这一音频模型推理基准上取得 97.7% 的成绩。Gemini 3.8 Live 在人类偏好评估 Speech Agent Arena 中位列第二。在 ServiceNow 的 EVA-Bench 上,Google 报告称这些模型推动了复杂工作流的帕累托前沿。它们在 Gemini Enterprise Agent Platform 上通过 Live API 测量,在任务准确性与对话质量之间取得平衡。
面向开发者的能力
Live API 在新模型中暴露了 5 项核心能力:
- 异步函数调用:模型在后台执行 API 和工具调用。音频响应持续向用户流式输出,同时任务在后台完成。
- 视觉上下文:模型近乎实时地处理实时视觉输入,因此智能体能够理解用户所说和所看到的内容。
- 字母数字精度:它能准确解析确认码、理赔编号和技术数据,这是语音系统中常见的失败点。
- 多语言支持: 它能在对话过程中自动检测并在 97 种受支持语言之间切换,同时保持口音一致性。
- 增量内容更新: 它将实时音频与结构化数据融合,从而返回具备上下文感知能力的响应。
扩展思考(Extended Thinking)为后台多步推理增加了可配置的思考。它能一边推理一边说话,使用诸如“让我查一下”这样的早期口头提示来确认收到提示词。随后,在长时间运行的任务执行过程中,它会逐步叙述进展。Google 的演示展示了该模型将草图加上语音反馈转换为可运行的 React 组件,并协调多步预订流程。
定价与生态
两款模型的定价均为音频输入 $0.005/min、音频输出 $0.018/min。Google 表示,这一估算基于 $3/1M 输入 tokens 和 $12/1M 输出 tokens。开发者还可以通过负责实时媒体流基础设施的 Live API 集成合作伙伴进行构建。这些合作伙伴包括 Agora、Fishjam、LangChain、LiveKit、Pipecat、Vercel 和 Vision Agents。Google 还在与 Salesforce、Genspark 和 Lumeris 合作,这些公司对模型的延迟、流畅性和工具调用能力给予了肯定。示例应用可在 GitHub 上获取。
核心要点
- Gemini 3.8 Live Extended Thinking 以 82.6 分在 Artificial Analysis 的语音到语音质量指数(Speech to Speech Quality Index)上排名第 1。
- 它在 τ-Voice 上得分 68.6%,在 Sierra 的 τ-Voice-banking 上得分 35.1%,在 Big Bench Audio 上得分 97.7%。
- Gemini 3.8 Live 在继续对话的同时,在后台运行工具和 API 调用。
- 通过 Live API,音频输入定价为 $0.005/min,音频输出定价为 $0.018/min。
- 所有生成的音频都带有 Google DeepMind 不可感知的 SynthID 水印。
Google has introduced Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, its most advanced live dialogue models to date. Both are native speech to speech models built for real time voice agents. They extend the Gemini Audio family that Google expanded last month with Gemini 3.5 Transcribe. The release targets a specific gap: voice agents that can reason and execute tools without breaking conversational flow.
Is it deployable? Yes, for API based production use. Both models are live today in the Gemini Live API and Google AI Studio. They are hosted models, not open weights, so there is no self hosted option.
What Google Released
The launch covers 2 models with distinct roles. Gemini 3.8 Live is built for scale and cost efficiency. It combines conversational intelligence with fluid dialogue and visual grounding. Gemini 3.8 Live Extended Thinking is built for high complexity tasks. It adds increased intelligence and multi step reasoning while it speaks. Google positions both as a streamlined alternative to cascaded speech pipelines that chain ASR, an LLM, and TTS.
Benchmark Results
Gemini 3.8 Live Extended Thinking takes the #1 overall spot on Artificial Analysis’ Speech to Speech Quality Index with a score of 82.6. It leads agentic task completion with 68.6% on τ-Voice and 35.1% on Sierra’s τ-Voice-banking benchmark. It also scores 97.7% on Big Bench Audio, a reasoning benchmark for audio models. Gemini 3.8 Live secured second place in the Speech Agent Arena, a human preference evaluation. On ServiceNow’s EVA-Bench, Google reports that the models push the Pareto Frontier for complex workflows. They balance task accuracy with conversational quality, measured on the Live API on Gemini Enterprise Agent Platform.
Capabilities for Developers
The Live API exposes 5 core capabilities in the new models:
- Asynchronous function calling: The model executes API and tool calls in the background. Audio responses keep streaming to the user while tasks finish.
- Visual context: The model processes live visual inputs in near real time, so agents can understand what users say and see.
- Alphanumeric precision: It accurately parses confirmation codes, claim numbers, and technical data, a common failure point in voice systems.
- Multilingual support: It automatically detects and transitions between 97 supported languages mid conversation, with accent consistency.
- Incremental content updates: It merges real time audio with structured data to return context aware responses.
Extended Thinking adds configurable thinking for multi step reasoning in the background. It reasons and speaks simultaneously, using early verbal cues such as “Let me check that” to acknowledge prompts. It then narrates progress step by step while long running tasks execute. Google’s demos show the model converting sketches plus voice feedback into working React components and coordinating multi step bookings.
Pricing and Ecosystem
Both models are priced at $0.005/min for audio input and $0.018/min for audio output. Google states this estimate is based on $3/1M input tokens and $12/1M output tokens. Developers can also build through Live API integration partners that handle real time media streaming infrastructure. These include Agora, Fishjam, LangChain, LiveKit, Pipecat, Vercel, and Vision Agents. Google is also partnering with Salesforce, Genspark, and Lumeris, which cite the models’ latency, fluidity, and tool calling. Example apps are available on GitHub.
Key Takeaways
- Gemini 3.8 Live Extended Thinking ranks #1 on Artificial Analysis’ Speech to Speech Quality Index with 82.6.
- It scores 68.6% on τ-Voice, 35.1% on Sierra’s τ-Voice-banking, and 97.7% on Big Bench Audio.
- Gemini 3.8 Live runs tools and API calls in the background while continuing the conversation.
- Pricing is $0.005/min for audio input and $0.018/min for audio output via the Live API.
- All generated audio carries Google DeepMind’s imperceptible SynthID watermark.