
이미지: The Verge AI
Summary
- Google has added three new models to its Gemini audio lineup: 3.5 Transcribe, 3.5 Live, and 3.5 Live Experimental
- 3.5 Transcribe automatically removes filler words and auto-detects more than 85 languages along with specialized terminology
- Rollout began August 26 for the Gemini app on macOS and for Android's Rambler, but the Gemini 3.5 Pro model teased back in June is still nowhere to be seen
- 신규 모델
- Gemini 3.5 Transcribe, Gemini 3.5 Live, Gemini 3.5 Live Experimental
- 지원 언어
- 85개 이상 자동 감지
- 화자 구분
- 사전 녹음 오디오 기준 최대 3명 구분 + 단어 단위 타임스탬프
- 배포 시작
- 2026년 8월 26일, 영어 기준 맥OS 제미나이 앱 전체
- 안드로이드 적용
- 일부 국가·언어의 램블러(Rambler) 받아쓰기 기능
- 개발자 접근
- 제미나이 API, AI 스튜디오, Antigravity 퍼블릭 프리뷰
- 비교 대상
- 이전 트랜스크립션 모델 Chirp 3 대비 다국어 성능·오류율 개선(구글 발표)
- 미출시 항목
- 6월 예고된 제미나이 3.5 프로, 아직 공개 안 됨
Google has refreshed its Gemini audio model lineup, the ones that handle speech-to-text conversion. The new releases are Gemini 3.5 Transcribe, 3.5 Live, and 3.5 Live Experimental — and of the three, 3.5 Transcribe is an entirely new model. It automatically filters out verbal filler like "um" and "uh," and it can detect and transcribe more than 85 languages on its own.
The announcement lands while Gemini 3.5 Pro — the flagship model Google teased back in June — is still nowhere to be found. The flagship has gone quiet while the audio lineup gets updated first. Google mentioned a roadmap for 3.5 Pro in its blog post, but this audio-model announcement didn't include a new timeline for the Pro model.
What's new
Google describes 3.5 Transcribe as "a major step up" from Chirp 3, its previous transcription model. The company says it improved multilingual handling and lowered the word error rate. If users pre-load a custom vocabulary list, the model will automatically catch specialized terms and unusual spellings, cutting down on manual cleanup afterward. For pre-recorded audio, it can distinguish up to three speakers and attaches timestamps to every word.
3.5 Live and 3.5 Live Experimental build on the existing speech-recognition tech that powers Gemini's voice chat mode. 3.5 Live handles interruptions and mid-sentence language switching better, and its real-time screen recognition has also been strengthened. 3.5 Live Experimental goes a step further, narrating what it's doing step by step in real time while it works through complex reasoning tasks.
| Model | Key features | Primary use |
|---|---|---|
| Gemini 3.5 Transcribe | Filler-word removal, custom vocabulary, speaker separation | Transcribing recorded audio |
| Gemini 3.5 Live | Handles interruptions and language switching, real-time screen processing | Voice chat mode |
| Gemini 3.5 Live Experimental | Real-time narration of reasoning process | Complex task handling |
How to try it
The update started rolling out on August 26 for English-language users across the Gemini app on macOS. If you're on macOS, there's nothing to set up — the new model kicks in automatically whenever you use voice input or dictation inside the Gemini app.
On Android, it's rolling out to the Rambler dictation feature, though only in select countries and languages for now. Developers can get early access to 3.5 Transcribe through public preview via the Gemini API, AI Studio, or Antigravity. Google says Chrome support is coming soon as well.
As a practical example: upload a meeting recording to 3.5 Transcribe and it'll break the transcript out by speaker and strip filler words, leaving you with a cleaner set of meeting notes. In fields packed with jargon — medicine, law, and the like — pre-registering a vocabulary list should cut down on the manual corrections you'd otherwise have to make every time.
Where's Gemini 3.5 Pro
Google's Gemini lineup has clearly picked up its release pace lately. As covered in Gemini 3.7 Flash spotted, and it looks like the release pace is accelerating, a Gemini 3.6 Flash test banner and references to a Gemini 3.7 Flash model name both surfaced around August 12-13. Yet the flagship 3.5 Pro that was teased back in June is still missing, and this audio-model announcement didn't come with a new release date for it either.
Editor's take
This isn't the first time Google has held back a flagship model while filling in the gaps with smaller releases first. Finishing a large language model takes time — safety validation, allocating massive amounts of compute — and rolling out specialized models for voice or image in the meantime helps keep up the impression that Gemini is still moving forward. It's the same pattern OpenAI followed when it refreshed its voice and image tools while still working on the next GPT version.
Having actually put transcription models to work, the filler-word removal and speaker separation here are the kind of features that visibly cut down meeting-note cleanup time. Older-generation models tended to rack up errors whenever background noise crept in or speech got cut off mid-sentence, and 3.5 Transcribe looks like an update aimed squarely at that weak spot. That said, features like custom vocabulary registration and speaker separation are still in API preview, so it'll take a few more weeks of stabilization before they're ready to drop into a real production workflow.
For teams in Korea, the practical first step is checking when Korean-language support arrives, since the rollout right now starts with English on macOS. When Chrome support and Korean-language expansion land will really determine when adoption becomes practical here. As for Gemini 3.5 Pro itself, I'd guess an actual release within the next month is likely — stabilizing the audio models first before shipping the flagship lines up with Google's recent pattern.




Comments