
이미지: @GoogleDeepMind (X) 영상 갈무리
Summary
- Google DeepMind posted a screenshot on X showing the "Custom vocabulary" feature for its speech recognition model, Gemini 3.5 Transcribe
- The screenshot shows an input field where users can manually add unfamiliar terms like "NeuroWave"
- The model has been rolling out to the macOS Gemini app since August 26, and this appears to layer a term-registration feature on top of that existing rollout
Google DeepMind posted a new screen from its speech recognition model, Gemini 3.5 Transcribe, on X. The caption is a single short line — "Gemini 3.5 Transcribe is our latest speech-to-text model for precise and intelligent transcriptions" — but the accompanying screenshot says more. The screen is titled "Custom vocabulary," and below an "Add terms..." input field sit registered term tags like "Gemini 3.5 Transcribe" and "NeuroWave." It looks like a feature that lets users manually enter proper nouns the model tends to miss — company names, product names — to improve recognition accuracy. That said, a line at the bottom of the screen reads "Sequences shortened and screen images simulated. Compatibility and availability varies," which is Google itself acknowledging that this screenshot is a reconstructed example rather than a literal capture of the live app, and that where the feature actually works may vary by device and region.
Actual product demo videos
The six clips below aren't reposts — they're sourced directly from Google's official announcement post and from official channels run by Google's product and developer teams. Four come straight from Google's announcement post, and two more were posted on X by Google AI Developers and the product lead for Google AI Studio.
Real-time multilingual transcription
An official Google demo showing Gemini 3.5 Transcribe keeping up without interruption as a speaker switches languages mid-sentence.
On-screen context and function calling in the macOS app
An official Google demo of the macOS Gemini app understanding on-screen context and handling tasks by voice, like drafting emails, analyzing files, and generating images.
Voice input in the macOS Gemini app
An official Google macOS demo showing how holding the Fn key and speaking in any app strips out filler words and mid-sentence corrections, then types the cleaned-up sentence right at the cursor.
Message input with Gboard Rambler
An official Google demo of Gboard Rambler on Android removing filler words and letting users revise sentences by voice.
Code-context awareness in Antigravity
An original clip from Google AI Developers showing Google Antigravity using the current code and file names as context to accurately transcribe voice input.
A dictation app built on Gemini 3.5 Transcribe
An official demo of a dictation app that Ammaar Reshi, Lead Product + Design at Google AI Studio, built himself using the model.
As METAL LAB covered in a previous article, Gemini 3.5 Transcribe is the speech-to-text model that launched on August 26 alongside Gemini 3.5 Live and Live Experimental. At launch, its standout features included automatically stripping filler words like "um" and "uh," auto-detecting more than 85 languages, distinguishing up to three speakers in pre-recorded audio, and attaching word-level timestamps. A long-standing weakness of speech recognition models is that while they handle common words well, they frequently misread brand names, coined terms, and industry jargon — words that rarely show up in training data. Simply expanding the number of auto-detected languages doesn't solve this, which is why letting users fill in their own glossary of terms has been a fix adopted across a number of speech recognition services.
This post alone doesn't confirm exactly when the newly revealed custom vocabulary screen will actually ship, or on which platforms. But given that the August 26 rollout reached the English-language macOS Gemini app and the Android Rambler dictation feature in select countries in stages, it's plausible this term-registration feature will follow a similarly gradual rollout. For anyone regularly transcribing meetings or interviews packed with specific proper nouns or technical jargon, this is a feature worth watching once it actually goes live.

Editor's take
What matters most here isn't recognition accuracy on its own — it's that Google showed a concrete path toward reducing real-world input errors by combining term registration with on-screen context. Looking at the top of this article alongside the demo videos below, you can see how the same model plugs into different input environments: dictation, messaging, and coding work alike.




Comments