今天,我们的技术和产品支撑着 300 多种语言的日常交流,使用人数超过 70 亿——占全球人口的 86%。达成这一里程碑意义重大,但它也凸显了对我们的使命至关重要的工作。几十年来,技术始终只为少数几种主导语言提供最佳支持,致使数千种仍在使用的语言和方言在数字世界中代表性不足,甚至完全缺席。
当我们在 2006 年推出Google Translate时,目标很简单:打破语言之间的壁垒。AI 的进步帮助我们把这个愿景带给更多人,让 Translate 从支持少数几种语言扩展到今天的 250 多种。但仅仅翻译文本还不够。技术需要理解人们在现实世界中究竟如何交流。因此,我们把研发重点放在构建能够尊重文化细微差别和人类语言丰富性的系统上,让每个人都能按自己的方式参与其中并被理解。
以下是这项工作在实践中的样子。
从文本走向真正的理解
过去,语音识别系统遵循一套僵化的多步骤流程:把音频转写成文本,处理该文本,然后再将其合成为音频。这套流水线虽然能用,却剥离了人类交流中最丰富的部分:语气、节奏、情感和上下文。
人们说话时并不会用完美工整、合乎语法的句子。我们会笑、会抢话、会犹豫,还会在一句话中间把多种语言交织在一起,就像我们说西班牙英语(Spanglish)或印度英语(Hinglish)时那样。
为了捕捉这一点,我们超越了文本转写,转向原生音频智能——训练像 Gemini 这样的模型直接按音频本来的样子进行处理,同时理解声音和意图。这些工作包括:
- Fluid real-time dialogue tools:
- 如今,Gemini 3.5 Live Translate 支持覆盖 70 种语言、2,000+ 语言对的实时口语翻译,并在过程中自然地捕捉语码转换和情感线索。
- Gemini 3.5 Transcribe 是我们迄今最精准的语音转文本模型,能把原始音频转化为经过润色、格式规范的文本,即便在嘈杂环境中或面对复杂术语也不例外。它还为 Android Gboard 上的 Rambler 等功能提供支持,该功能可以去除填充词、修正语法和标点,让你通过语音指令编辑或改写,并在语言之间无缝切换。
- 1000 种语言计划:AI 正在帮助我们以前所未有的规模打破语言障碍。但要让更多人用自己偏好的语言使用 AI,就意味着不能只停留在目前 AI 表现最好的那些语言上:我们的目标是支持全球使用人数最多的 1000 种语言。为实现这一目标,我们的通用语音模型——基于 1200 万小时音频训练——采用了跨语言迁移学习,这类技术能让模型将从数据丰富语言中学到的知识迁移过来,从而提升训练数据远为稀少的语言的语音理解能力。这使得模型能够把从数据丰富语言中学到的模式应用到资源匮乏的语言上。
- 严谨的基础研究:这项工作建立在 25 年开放研究以及超过400 篇同行评审语音论文的基础之上,这些研究帮助推动了前沿进展并促进了语音模型的发展。
将社区置于语言数据的核心
由于网络内容不成比例地偏向少数几种主导语言,教会 AI 理解代表性不足的语言,要求我们重新思考收集数据的方式。解决方案是本地化的基层合作。这种本地化方法推动了我们三项最具雄心的开放数据合作:
- WAXAL(沃洛夫语中意为“说话”,发音为“Wah-hal”):WAXAL 与 Makerere University、Digital Umuganda 等合作伙伴共同打造,是一个大规模开放语音数据集,覆盖 27 种撒哈拉以南非洲语言,这些语言由超过 26 个国家的逾 1 亿人使用,捕捉了传统数据集常常缺失的声调变化和对话节奏。
- Project Vaani: 与印度科学研究所(IISc)和 Bhashini 合作,Project Vaani 通过以地区而非语言为锚点的方法来绘制印度的语言多样性,迄今已收集了来自超过 155,000 名说话者、涵盖 109 种语言的逾 30,000 小时语音。
- Amplify Initiative: 我们与四大洲的 1,600 多名本地专家和 20 所大学合作,包括巴西的 UFMG、印度的 IIT Kharagpur 和乌干达的 Makerere University,贡献了 15,000 个捕捉本地细微差异的多模态数据点。
我们还在通过新工具 Language Explorer 推进我们优先支持开源语言创新的工作。这是一个交互式工具,可对 LinguaMeta 进行可视化——后者是全球最大的开源语言数据存储库。它因设计创新而获得 Fast Company 认可,持续绘制超过 7,000 种口语、书面语和手语语言。
这些创新与合作,只有在惠及那些能够将新数据和洞察转化为社区切实改变的人群时,才能发挥最大影响。Google.org 支持的相关工作,包括 数字语言包容中心和 AI Singapore 的 Project Aquarium,正在帮助将多语言工具带给全球的农民、医护人员、教师以及其他重要的社区成员。
克服现实世界的种种限制
对于超过 30 亿人 1 来说,可靠的互联网接入仍然遥不可及。技术只有能在人们生活的地方——包括连接有限或时断时续的地区——正常运转,才算真正可及。
为帮助解决这一问题,我们开发了 TranslateGemma,这是一个基于 Gemini 构建、在 55 种语言上训练的轻量级开放翻译模型系列。由于 TranslateGemma 能够在设备端高效运行,高质量翻译不再需要连接云端或互联网。
不过,运行强大的 AI 模型需要性能足够的硬件,而低资源地区仍有数亿人使用功能手机,被排除在外。为弥合这一鸿沟,我们正在支持 Viamo 等组织打造“Ask Viamo Anything”(AVA),这是一款语音 AI 助手,将 Gemini 的能力带到标准功能手机上。Viamo 已在其现有的交互式语音应答用户中于卢旺达试点 AVA,该服务已利用 Gemini 回答了超过 200 万个问题。
为无障碍而设计
语言不仅仅关乎地区方言或词汇,还关乎人们沟通的许多其他方式。传统语音工具常常无法服务于使用非标准语音的人群,迫使他们去适应技术,而不是让技术来适应他们。
我们正致力于改变这一现状,从底层开始为无障碍而设计,例如手语转文本(SL2T)。SL2T 在 50 多种手语上训练,为 Pixel 11 上 Gboard 和 Live Transcribe 的手语转文本听写提供支持,首批支持美国手语(ASL)转英语。这是朝着让我们的产品对全球 7000 万依赖手语沟通的人更无障碍迈出的重要第一步。
正确处理本地发音
细节很重要,在导航这类日常工具中尤其如此。当导航应用读错一个小镇或街道的名称时,不仅会造成困惑,还可能忽视该地的文化遗产。
我们认为,文化语境应当成为语言技术构建方式的一部分,这意味着要直接与当地社区合作。例如,在新西兰,我们与毛利语专家合作,改进了 Google Maps 中的地名发音,帮助让体验更加准确、真正具有本地特色。将文化上真实的发音直接融入我们的文本转语音模型,确保技术能够反映其所服务社区的语言和传承。
为每个人打造语言工具
在 20 年的 AI 语言研究与开发中,我们学到的一个重要经验是:技术绝不应缩小人类表达的频谱——而应拓展它。
如今,我们的语言技术已经嵌入我们的核心生态系统,连接着九大平台上超过 50 亿人,包括 Search、Android、Chrome、YouTube 和 Google Play。但规模只是故事的一部分。更大的目标是深度与丰富性——构建能够理解语境、尊重文化,并赞颂人们多种多样沟通方式的系统。
随着我们不断扩大利用 AI 来应对社会一些最大的机遇,我们将继续与当地社区紧密合作,构建帮助更多人按自己的方式沟通、参与并被理解的技术。
Today, our technologies and products power everyday interactions in more than 300 languages, spoken by more than 7 billion people — representing 86% of the global population. Reaching this milestone is meaningful, but it also underscores work that is critical to our mission. For decades, technology has worked best for a handful of dominant languages, leaving thousands of living languages and dialects poorly represented or absent altogether from the digital world.
When we launched Google Translate in 2006, our goal was simple: to break down the barriers between languages. Advances in AI have helped us bring that vision to more people, expanding Translate from a handful of languages to more than 250 today. But translating text isn’t enough. Technology needs to understand how people actually communicate in the real world. So we focus our research and development on building systems that honor cultural nuance and the richness of human language, enabling everyone to participate and be understood on their own terms.
Here’s what that work looks like in practice.
Going from text to true understanding
Historically, speech recognition systems followed a rigid, multi-step process: transcribing audio into text, processing that text, and then synthesizing it back into audio. While functional, this pipeline strips away the richest parts of human communication: tone, pacing, emotion, and context.
People don't speak in perfectly neat, grammatical sentences. We laugh, overlap, hesitate, and weave multiple languages together mid-sentence, like when we speak Spanglish or Hinglish.
To capture this, we moved beyond text transcripts to native audio intelligence — training models like Gemini to process audio directly as is, while also grasping both sound and intent. These efforts include:
- Fluid real-time dialogue tools:
- Today, Gemini 3.5 Live Translate powers real-time spoken translation across 70 languages and 2,000+ language pairs, naturally capturing code-switching and emotional cues along the way.
- Gemini 3.5 Transcribe is our most precise speech-to-text model yet, turning raw audio into polished, formatted text, even in noisy environments or with complex jargon. It also powers features like Rambler on Android Gboard, which removes filler words, fixes grammar and punctuation, and lets you edit or rewrite with voice commands and switch seamlessly between languages.
- The 1,000 Languages Initiative: AI is helping us break down language barriers at a scale that was previously unimaginable. But reaching more people in their preferred language means going beyond the languages where AI performs best today: Our goal is to support the world’s 1,000 most-spoken languages. To help make that possible, our Universal Speech Model — trained on 12 million hours of audio — used cross-lingual transfer learning, techniques that enable models to transfer what they learn from data-rich languages, to improve speech understanding in languages with far less training data. This allows models to apply patterns learned from data-rich languages to under-resourced ones.
- Rigorous foundational research: This work builds on 25 years of open research and more than 400 peer-reviewed speech papers, which have helped push the frontier and advance speech models.
Putting communities at the heart of language data
Because the web disproportionately represents a few dominant languages, teaching AI to understand underrepresented languages required us to rethink how we gather data. The solution is local grassroots partnerships. This localized approach has driven three of our most ambitious open-data partnerships:
- WAXAL (Wolof for “speaking,” pronounced "Wah-hal"): Built with partners including Makerere University and Digital Umuganda, WAXAL is a large-scale, open speech dataset covering 27 Sub-Saharan African languages spoken by more than 100 million people across more than 26 countries, capturing tonal variation and conversational rhythms often missing from traditional datasets.
- Project Vaani: In partnership with the Indian Institute of Science (IISc) and Bhashini, Project Vaani is mapping India’s linguistic diversity through a region-anchored rather than language-anchored approach, enabling it to collect to date more than 30,000 hours of speech across 109 languages from more than 155,000 speakers.
- Amplify Initiative: We teamed up with more than 1,600 local experts and 20 universities across four continents, including Brazil’s UFMG, India’s IIT Kharagpur, and Uganda’s Makerere University, to contribute 15,000 multimodal data points capturing local nuance.
We’re also building on our work prioritizing open-source language innovation through our new tool Language Explorer. It’s an interactive tool that visualizes LinguaMeta, the world’s largest open-source language data repository. Recognized by Fast Company for design innovation, it continuously maps more than 7,000 spoken, written, and signed languages.
The impact of these innovations and partnerships is greatest when they reach the people who can turn new data and insights into meaningful change in their communities. Google.org-supported efforts, including the Centre for Digital Language Inclusion and AI Singapore’s Project Aquarium, are helping bring multilingual tools to farmers, healthcare workers, teachers, and other essential community members around the world.
Overcoming real-world constraints
For more than 3 billion people 1 , reliable internet access is still out of reach. Technology is only truly accessible if it works where people live, including areas with limited or intermittent connectivity.
To help address this, we developed TranslateGemma, a family of lightweight open translation models built from Gemini and trained across 55 languages. Because TranslateGemma runs efficiently on-device, high quality translation no longer requires a connection to the cloud or the internet.
Still, running powerful AI models requires capable hardware, which excludes the hundreds of millions of people still using feature phones in low-resource regions. To bridge this divide, we’re supporting organizations like Viamo to power “Ask Viamo Anything” (AVA), a voice AI assistant that brings the power of Gemini to standard feature phones. Viamo has successfully piloted AVA in Rwanda with its existing interactive voice response users, and the service has already used Gemini to answer more than 2 million questions.
Designing for accessibility
Language isn’t just about regional dialects or vocabulary. It’s also about the many other ways people communicate. Conventional speech tools frequently fail people with non-standard speech, making them adapt to the technology rather than the other way around.
We’re working to change that by designing for accessibility from the ground up, for example, with Sign Language-to-Text (SL2T). Trained across 50+ sign languages, SL2T powers sign-to-text dictation in Gboard and Live Transcribe on Pixel 11, starting with American Sign Language (ASL) to English. This is an important first step toward making our products more accessible to the 70 million people worldwide who rely on sign language to communicate.
Getting local pronunciation right
Details matter, certainly in everyday tools like navigation. When a navigation app mispronounces a town or street name, it doesn’t just cause confusion, it can overlook the cultural heritage of the place.
We believe cultural context should be part of how language technology is built, and that means working directly with local communities. For example, in New Zealand, we worked with Māori language experts to improve the place name pronunciation in Google Maps, helping make the experience more accurate and genuinely local. Incorporating culturally authentic pronunciations directly into our text-to-speech models ensures technology reflects the language and heritage of the communities it serves.
Building language tools for everyone
One essential lesson we’ve learned over 20 years of AI language research and development is that technology should never narrow the spectrum of human expression — it should expand it.
Today, our language technologies are already embedded across our core ecosystem, connecting more than five billion people across nine platforms, including Search, Android, Chrome, YouTube, and Google Play. But scale is only part of the story. The bigger goal is depth and richness — building systems that grasp context, respect culture, and celebrate the many ways people communicate.
As we expand our use of AI to address some of society’s biggest opportunities, we’ll continue to work closely with local communities to build technology that helps more people communicate, participate, and be understood on their own terms.