
Image: METAL
Summary
- Google says its technology supports more than 300 languages and 7 billion people, 86% of the world's population.
- Gemini 3.5 Live Translate interprets in real time across 70 languages and more than 2,000 language pairs, processing audio without first converting it to text.
- WAXAL, Project Vaani and the Amplify Initiative are built so that local communities record the data themselves and partners keep ownership of it.
Google published a post on its company blog on September 15 that gathers 20 years of language AI research into a single overview. The author is James Manyika, Senior Vice President for Research, Labs, Technology and Society, and the first sentence opens with numbers. Google's technology and products now support everyday interactions in more than 300 languages, spoken by more than 7 billion people, or 86% of the world's population. The same day, Yossi Matias of Google Research condensed the main points into a four-post thread on X, and the official Google Research account quoted that thread to share it again.
The starting point is Google Translate in 2006. Manyika writes that the goal at the time was to break down the barriers between languages, and that advances in AI have expanded Translate from a handful of languages to more than 250 today. The weight of the post, though, sits in the next sentence. He writes, "But translating text isn't enough," adding that technology needs to understand how people actually communicate in the real world. Behind that comes his diagnosis: for decades, technology has worked best for a handful of dominant languages, leaving thousands of living languages and dialects poorly represented or absent altogether from the digital world.
The first example is a change in how speech is processed. According to the post, speech recognition used to follow a rigid, multi-step pipeline: transcribing audio into text, processing that text, then synthesizing it back into audio. It worked, but it stripped away the richest parts of communication along the way, including tone, pacing, emotion and context. People do not speak in neat, grammatical sentences. They laugh, overlap, hesitate and weave several languages together mid-sentence, as in Spanglish or Hinglish. Google says it therefore trained models like Gemini to process audio directly, without passing through text.
Two products sit on top of that. Gemini 3.5 Live Translate handles real-time spoken translation across 70 languages and more than 2,000 language pairs, capturing code-switching within a sentence and emotional cues as they come. Gemini 3.5 Transcribe is the most precise speech-to-text model Google has built, turning raw audio into polished text even in noisy environments or with complex jargon. The same model powers the Rambler feature in Android Gboard, which removes filler words, fixes grammar and punctuation, and lets users edit or rewrite by voice command or switch between languages.
The target number of languages is also stated. Google has set a goal of supporting the world's 1,000 most-spoken languages, and the Universal Speech Model underneath that effort was trained on 12 million hours of audio. The method is cross-lingual transfer learning, which carries patterns learned from data-rich languages over to languages with far less data. Manyika writes that this work builds on 25 years of open research and more than 400 peer-reviewed speech papers.
Seen through a sociologist's eyes, the most important part of the post is not the models but the way the data is gathered. Because the web disproportionately represents a few dominant languages, Google acknowledges that teaching AI to understand underrepresented languages required rethinking data collection itself. The answer was local grassroots partnerships, and the result is three open-data collaborations. The first is WAXAL. The name means "speaking" in Wolof, and the dataset was built with partners including Makerere University and Digital Umuganda. It covers 27 Sub-Saharan African languages spoken by more than 100 million people across more than 26 countries.
The WAXAL announcement that METAL reviewed on the Google Research blog, dated March 6, gives more specific figures. The effort began in 2021 and ran for several years, and the initial release contains approximately 1,846 hours of transcribed natural speech for speech recognition and more than 565 hours of high-quality recordings for speech synthesis, released under a CC-BY-4.0 license. Instead of reading scripts, participants described images across more than 50 topics in their native language, a method Google Research says captured tonal variation and code-switching. The collection was led entirely by African academic and community organizations, and the announcement states the principle that partners retain ownership of the data they collected.
The second is Project Vaani in India. Working with the Indian Institute of Science and Bhashini, it maps India's linguistic diversity by region rather than by language, and to date has collected more than 30,000 hours of speech across 109 languages from more than 155,000 speakers. The third is the Amplify Initiative, in which more than 1,600 local experts and 20 universities across four continents, including Brazil's UFMG, India's IIT Kharagpur and Uganda's Makerere University, contributed 15,000 multimodal data points capturing local nuance. All three are structured so that the communities who speak a language record it themselves rather than having it collected from outside.
A tool that shows where this data exists and how much of it there is was introduced alongside. Language Explorer is an interactive tool that visualizes LinguaMeta, the world's largest open-source language data repository, mapping more than 7,000 spoken, written and signed languages. The post says the tool has been recognized externally for design innovation. The Centre for Digital Language Inclusion and AI Singapore's Project Aquarium, both supported by Google.org, carry that data and insight to farmers, healthcare workers, teachers and other community members as multilingual tools.
The post also lays out the conditions technology does not reach. Manyika writes that reliable internet access is still out of reach for more than 3 billion people. The answer to that is TranslateGemma, a family of lightweight open translation models built from Gemini, trained across 55 languages and running on-device, so translation no longer requires the cloud or an internet connection. According to the January 15 announcement METAL reviewed, TranslateGemma is built on Gemma 3 and comes in three sizes, 4B, 12B and 27B, and the 12B model outperformed the Gemma 3 27B baseline, a model twice its size, as measured by MetricX on the WMT24++ benchmark. METAL has reported on Google's release of an offline voice translator that runs on-device.
Some places do not even have smartphones. For the hundreds of millions of people using feature phones, Google supports Viamo's voice AI assistant AVA, which brings Gemini's capabilities to standard feature phones. According to the post, Viamo piloted AVA in Rwanda with its existing interactive voice response users, and the service has already used Gemini to answer more than 2 million questions. On the accessibility side, SL2T, trained across more than 50 sign languages, powers sign-to-text dictation in Gboard and Live Transcribe on Pixel 11, starting with American Sign Language to English. Google calls it a first step toward the 70 million people worldwide who rely on sign language. METAL reported on the SL2T announcement in August.
Something as small as place-name pronunciation falls under the same principle. The post says that when a navigation app mispronounces a town or street name, it does more than cause confusion; it overlooks the cultural heritage of the place, and in New Zealand Google worked with Māori language experts to improve place-name pronunciation in Google Maps. Manyika sets down one lesson from 20 years of language AI research this way: "technology should never narrow the spectrum of human expression — it should expand it." In his X thread, Matias made the same point in different words: "translation and text processing alone are insufficient; systems must reflect how people communicate in the real world across diverse cultural nuances."
Scale comes last. Google's language technologies connect more than 5 billion people across nine platforms, including Search, Android, Chrome, YouTube and Google Play. Yet the post itself says that scale is only part of the story. Who gathers data in which languages, and who owns that data, is a technical question and at the same time a question of which languages come to exist in the digital world. The structure of WAXAL, where partners own the data and communities do the recording, and the design of Vaani, anchored in regions rather than languages, are the parts of this post that will last longest. After 20 years, Google has rewritten the task of language AI from translation to representation.





Comments