我们最新的 Gemini 模型提供在大规模构建 AI 智能体时所需的效率、延迟和可靠性。
Tulsee Doshi
产品管理高级总监,代表 Gemini 团队
构建生产级 AI 智能体的开发者和客户需要更高的 token 效率、更低的延迟以及更可靠的性能。我们的 Flash 系列模型正是为满足效率与质量的最佳平衡点而打造,从而支持智能体工作流的规模化扩展。在 Gemini 3.5 Flash 的基础上,我们推出全新的 Gemini 模型:
- 3.6 Flash:我们的主力模型,在编程、知识工作和多模态性能方面表现更佳。根据 Artificial Analysis Index,相比 3.5 Flash,它将输出 token 用量减少了 17%,而在 Datacurve 的 DeepSWE 等部分基准测试中,我们观察到降幅最高可达 65%,且每输出 token 的成本更低。
- 3.5 Flash-Lite:我们最快、最具成本效益的 3.5 级模型,根据 Artificial Analysis Index,每秒可输出 350 个 token,在智能体工作流中也显著优于前几代 Flash-Lite。
- CodeMender 中的 3.5 Flash Cyber:成功的网络安全应用需要将模型与智能体基础设施精心编排。我们推出了一款全新的、高效的、专注于网络安全的专用模型,并将其与我们的 CodeMender 代码安全智能体相结合,在前沿水平上提供极具竞争力的性能。
在今日发布之外,Gemini 3.5 Pro 目前正在与合作伙伴进行测试,我们计划在准备就绪后尽快将其广泛开放。与此同时,我们的团队已经在专注于构建下一代模型。我们已经启动了迄今为止最具雄心的预训练运行,用于 Gemini 4,并对其进展感到兴奋。
3.6 Flash:比 3.5 Flash 更高效、质量更好
Gemini 3.6 Flash 直接建立在来自 3.5 Flash 的开发者和客户反馈之上。3.6 Flash 不仅在编码和知识工作方面实现了提升,同时还在显著改善 token 效率。例如,在 Artificial Analysis Index 上,我们看到 3.6 Flash 的输出 token 消耗比 3.5 Flash 少 17%。它完成多步骤工作流所需的推理步骤和工具调用也更少。
这种增强的效率还结合了比 3.5 Flash 更低的价格。在 $1.50/1M 输入 token 和 $7.50/1M 输出 token 的价格下,3.6 Flash 降低了每个智能体任务的总体成本,使智能体的构建和运行更具成本效益。
在一个 OSWorld 验证任务(API)中,3.6 Flash 相比 3.5 Flash 展现出更好的 token 效率并减少了冗长程度
即便效率更高,3.6 Flash 在各类用例中的表现相比 3.5 Flash 仍有提升:
- 3.6 Flash 精度更高,不必要的代码编辑更少,执行循环也更少,这一点在 DeepSWE 上有所体现(49% 对 37%);在 ML Research 上也有显著提升,这一点在 MLE Bench 上有所体现(63.9% 对 49.7%)。
- 其计算机使用能力有所提升,这一点在 OSWorld-Verified 上有所体现(83.0% 对 78.4%)。计算机使用现已成为通过 Gemini API 和 Gemini Enterprise 提供的内置客户端工具。
- 它在知识工作方面优于 3.5 Flash,GDPval-AA v2 等基准测试显示了这一点(1421 对 1349)。Hebbia 和 Harvey 等客户发现,它在文档解析、图表与数据分析以及报告起草等多模态任务上尤为出色。
3.6 Flash 在 AIS 上使用 Managed Agents,能够比 3.5 Flash 更高效、更准确地解析和分析财务数据与文字记录(AIS)
3.6 Flash 在 AGY 上利用多智能体编排执行代码迁移,相比 3.5 Flash 具有更低的延迟和更高的质量(AGY)
3.6 Flash 借助 canvas 帮助开发用于 3D 工作流的照片级纹理提取器(Gemini App)
3.6 Flash 利用 AGY 和 tldraw 离线编辑器,凭借其强大的视觉理解能力构建交互式主题工作室(AGY)。
客户反馈,3.6 Flash 在成本和质量上均实现了进步,在复杂工作流和知识型任务中平衡了 token 效率、准确性和速度:
安全构建
3.6 Flash 在化学、生物、放射性及核(CBRN)以及网络攻击滥用领域,随附了增强的 Frontier Safety 防护措施。这些防护措施使模型对越狱的抵抗力大幅提升。与此同时,该模型经过训练,能够最大限度地减少对有益用途的拒绝。
如需了解更多信息,请参阅 3.6 Flash 模型卡。
3.5 Flash-Lite:为规模化智能体工作流而构建
除 Flash 之外,我们还发布了 Gemini 3.5 Flash-Lite,专为低延迟任务以及高吞吐量对开发者工作流至关重要的任务而设计,例如智能体搜索和文档处理。
3.5 Flash-Lite 是 3.5 系列中速度最快的模型。根据 Artificial Analysis 的测量,其运行速度为 350 输出 tokens/s。定价为 $0.3/1M 输入 tokens 和 $2.5/1M 输出 tokens,且质量显著优于 3.1 Flash-Lite,3.5 Flash-Lite 为运行高吞吐量生产流量的开发者和客户提供了强劲的性价比。
3.5 Flash-Lite 以比 3.5 Flash 更低的延迟执行高容量任务。
3.5 Flash-Lite 为智能体系统带来了高效的扩展能力。在各个思考层级上,该模型都显著优于 3.1 Flash-Lite。根据工作负载的不同,开发者可以将模型配置为以低延迟、低成本执行为优先,利用最低和低思考层级处理高并发任务,或者启用更高的思考层级来处理多步骤的子智能体工作负载。该模型现在还内置了计算机使用工具,以可靠地支持跨平台的这些智能体任务。
在编码和智能体任务上实现了显著提升,Terminal-Bench 2.1 上的表现为(54% 对 31%),长上下文方面体现在 GDM-MRCR v2(72.2% 对 60.1%),真实世界任务执行方面体现在 GDPval-AA v2(1140 对 642)。
事实上,在许多智能体和编码评测中,3.5 Flash-Lite 甚至超越了 3 Flash,包括 SWE-Bench Pro(54.2% 对 49.6%)和 OSWorld-Verified(74.0% 对 65.1%),使其成为 2.5 和 3 Flash 上工作负载更快且能力更强的选择。
3.5 Flash-Lite 从海量电商数据集中提取产品特征并进行综合整理。
3.5 Flash-Lite 与 3.6 Flash 作为主智能体协同工作,即时生成 25 个独特的、可直接探索的网页设计概念。
3.5 Flash-Lite 凭借其多模态理解能力,可以大规模地完成收据翻译和摘要。
3.5 Flash-Lite 通过即时生成并迭代多个方案来构建一款游戏。
3.5 Flash-Lite 的早期客户正在强调它在速度、智能和成本效率方面的独特组合,可用于扩展智能体工作流和数据处理任务:
如需了解有关该模型的更多信息,请参阅 3.5 Flash-Lite 模型卡。
CodeMender 中的 3.5 Flash Cyber:高效发现并修复漏洞
AI 模型发现安全漏洞的速度已经超过了现有系统修复漏洞的速度。应对这一日益增长的威胁,需要一种能力强大且高效的软件安全防护方法。
Flash 的性能和效率使其成为大规模检测、验证和修补代码安全问题的理想基础。Gemini 3.5 Flash Cyber 构建于 3.5 Flash 之上,并经过微调,用于发现和修复网络安全漏洞,其每 token 价格低于更大的模型。
在 CodeMender 中,多个 3.5 Flash Cyber 智能体协同工作以生成一份合并报告,3.5 Flash Cyber 在热门基准 CyberGym 上达到了前沿水平的竞争力表现。
鉴于这项技术的两用性质,我们对 3.5 Flash Cyber 的部署采取了审慎的做法。该模型将很快通过 CodeMender 以限量访问试点计划的形式,独家提供给各国政府和可信合作伙伴。这将使一线防御者能够抢先发现并修复关键漏洞,赶在其被利用之前,同时降低被更广泛滥用的风险。
3.6 Flash 与 3.5 Flash-Lite:即刻开始使用
3.6 Flash 与 3.5 Flash-Lite 从今天起即可使用:
- 面向开发者,可通过 Google AI Studio 和 Android Studio 中的 Gemini API 使用。3.6 Flash 也可在 Google Antigravity 中使用。请参阅 开发者指南开始使用。
- 面向企业,可在 Gemini Enterprise Agent Platform 中使用。3.6 Flash 也可在 Gemini Enterprise 应用中使用。
- 面向所有人,可通过 Gemini 应用使用。3.5 Flash-Lite 也正在 Google 搜索中逐步推出。
当你开始使用 3.6 Flash 和 3.5 Flash-Lite 进行构建时,我们欢迎你的反馈,以改进未来的 Gemini 模型,并期待很快发布 3.5 Pro。
Our newest Gemini models deliver the efficiency, latency, and reliability to build AI agents at scale.
Tulsee Doshi
Senior Director, Product Management, on behalf of the Gemini team
Developers and customers building production AI agents need higher token efficiency, lower latency, and more reliable performance. Our Flash series of models is built to meet the sweet spot of efficiency and quality to enable scaling agentic workflows. Building on Gemini 3.5 Flash, we’re introducing new Gemini models:
- 3.6 Flash: Our workhorse model that delivers better coding, knowledge work, and multimodal performance. According to the Artificial Analysis Index, it reduces output token usage by 17% compared to 3.5 Flash, and in some benchmarks like DeepSWE by Datacurve, we observe up to 65%, all at a lower cost per output token.
- 3.5 Flash-Lite: Our fastest, most cost-effective 3.5-class model, delivering 350 output tokens per second according to the Artificial Analysis Index, also significantly outperforming prior Flash-Lite generations in agentic workflows.
- 3.5 Flash Cyber in CodeMender: Successful cybersecurity applications require careful orchestration of a model alongside an agent infrastructure. We’re introducing a combination of a new, highly efficient, specialized cyber-focused model paired with our CodeMender code security agent that delivers competitive performance at the frontier.
Beyond today’s releases, Gemini 3.5 Pro is currently testing with partners and we plan to make it broadly available as soon as it’s ready. In parallel, our team is already focusing on building the next generation of models. We have started our most ambitious pre-training run yet, for Gemini 4, and are excited by the progress.
3.6 Flash: More efficient and better quality than 3.5 Flash
Gemini 3.6 Flash builds directly on developer and customer feedback from 3.5 Flash. 3.6 Flash not only delivers a step up in coding and knowledge work, but it does this while meaningfully improving token efficiency. For example, on the Artificial Analysis Index, we see 3.6 Flash consuming 17% fewer output tokens than 3.5 Flash. It also takes fewer reasoning steps and tool calls to accomplish multi-step workflows.
This enhanced efficiency is also combined with a lower price than 3.5 Flash. At $1.50/1M input tokens and $7.50/1M output tokens, 3.6 Flash reduces the overall cost per agentic task, making agents more cost-effective to build and run.
3.6 Flash shows better token efficiency and reduced verbosity than 3.5 Flash in an OSWorld verified task (API)
Even while being more efficient, 3.6 Flash sees performance gains compared to 3.5 Flash across use cases:
- 3.6 Flash delivers higher precision with fewer unwanted code edits and reduced execution loops, as seen in DeepSWE (49% vs. 37%), and shows significant improvement in ML Research, as seen in MLE Bench (63.9% vs. 49.7%).
- It has improved computer use capabilities as seen in OSWorld-Verified (83.0% vs. 78.4%). Computer use is now a built-in client side tool via the Gemini API and Gemini Enterprise.
- It outperforms 3.5 Flash in knowledge work, as shown by benchmarks like GDPval-AA v2 (1421 vs. 1349). Customers like Hebbia and Harvey have found it particularly capable at multimodal tasks like document parsing, chart and data analysis, and report drafting.
3.6 Flash, using Managed Agents on AIS, can help parse through and analyze financial data and transcripts more efficiently and accurately than 3.5 Flash (AIS)
3.6 Flash executes code migrations, using multi-agent orchestration on AGY, with lower latency and higher quality than 3.5 Flash (AGY)
3.6 Flash helps develop a photographic texture extractor for 3D workflows, using canvas (Gemini App)
3.6 Flash uses AGY and the tldraw offline editor to build interactive theme studios with its strong visual understanding skills (AGY).
Customers report 3.6 Flash is a step forward in both cost and quality, balancing token efficiency, accuracy, and speed across complex workflows and knowledge-based tasks:
Built with safety
3.6 Flash is shipping with enhanced Frontier Safety safeguards in the domains of Chemical, Biological, Radiological, and Nuclear (CBRN) and cyber offense misuses. These safeguards make the model substantially more resistant to jailbreaks. At the same time, the model has been trained to minimize refusals for beneficial uses.
For more information, see the 3.6 Flash model card.
3.5 Flash-Lite: Built to scale agentic workflows
Beyond Flash, we’re also releasing Gemini 3.5 Flash-Lite, designed for both low-latency tasks and tasks where high throughput is critical for developers workflows, like agentic search and document processing.
3.5 Flash-Lite is the fastest model in the 3.5 series. As measured by Artificial Analysis, it runs at 350 output tokens/s. Priced at $0.3/1M input tokens and $2.5/1M output tokens and with significantly better quality than 3.1 Flash-Lite, 3.5 Flash-Lite offers a strong price-to-performance ratio for developers and customers running high throughput production traffic.
3.5 Flash-Lite executes high volume tasks at a lower latency than 3.5 Flash.
3.5 Flash-Lite enables efficient scaling for agentic systems. Across thinking levels, the model significantly outperforms 3.1 Flash-Lite. Depending on the workload, developers can configure the model to prioritize low-latency, low-cost execution for high-volume tasks with the minimal and low thinking levels, or engage higher thinking levels to process multi-step subagent workloads. The model now also has computer use as a built-in tool to reliably support these agentic tasks across surfaces.
It’s a significant step up in coding and agentic tasks as seen in Terminal-Bench 2.1 (54% vs 31%), long context as seen in GDM-MRCR v2 (72.2% vs. 60.1%), and real-world task execution as seen in GDPval-AA v2 (1140 vs. 642).
In fact, on many agentic and coding evals, 3.5 Flash-Lite even outperforms 3 Flash, including on SWE-Bench Pro (54.2% vs. 49.6%) and OSWorld-Verified (74.0% vs. 65.1%), making it a faster & more capable option for workloads on both 2.5 and 3 Flash.
3.5 Flash-Lite extracts product features from a massive e-commerce dataset and synthesizes it.
Working alongside 3.6 Flash as the master agent, 3.5 Flash-Lite instantly generates 25 unique, ready-to-explore web design concepts.
3.5 Flash-Lite can scale receipt translation and summarization with its multimodal understanding.
3.5 Flash-Lite builds a game by instantly generating and iterating through multiple options.
Early customers of 3.5 Flash-Lite are highlighting its unique combination of speed, intelligence, and cost efficiency for scaling agentic workflows and data processing tasks:
For more information about the model, see the 3.5 Flash-Lite model card.
3.5 Flash Cyber in CodeMender: finding and fixing vulnerabilities efficiently
AI models have become capable of finding security vulnerabilities faster than current systems can fix them. Tackling this growing threat requires an approach to securing software that is highly capable and efficient.
Flash’s performance and efficiency makes it an ideal foundation to detect, validate, and patch code security issues at scale. Gemini 3.5 Flash Cyber is built on top of 3.5 Flash, and fine-tuned for finding and fixing cybersecurity vulnerabilities at a lower price per token than larger models.
Within CodeMender, which uses multiple 3.5 Flash Cyber agents working together to produce a single combined report, 3.5 Flash Cyber reaches competitive performance at the frontier on the popular benchmark CyberGym.
Given the dual-use nature of this technology, we have taken an intentional approach to deploying 3.5 Flash Cyber. The model will be exclusively available to governments and trusted partners via CodeMender soon as part of a limited-access pilot program. This will give frontline defenders a head start in finding and fixing critical vulnerabilities before they can be exploited, while mitigating against broader misuse.
3.6 Flash and 3.5 Flash-Lite: Get started today
3.6 Flash and 3.5 Flash-Lite are available starting today:
- For developers in the Gemini API via Google AI Studio and Android Studio. 3.6 Flash is also available in Google Antigravity. Get started with the Developer Guide.
- For enterprises in Gemini Enterprise Agent Platform. 3.6 Flash is also available in the Gemini Enterprise app.
- For everyone via the Gemini app. 3.5 Flash-Lite is also rolling out in Google Search.
As you start building with 3.6 Flash and 3.5 Flash-Lite, we welcome your feedback to improve future Gemini models and look forward to releasing 3.5 Pro soon.