Gemini-3.5-Transcribe
- AI
- Developer Tools
- Mobile
- Open Source
Google posted Gemini 3.5 Transcribe as a new speech-to-text model aimed at both API use and Android dictation. The pitch is not just lower word error rate, but cleaner formatting and a more polished dictation experience. People reading it immediately separated those two jobs. For straight transcription, the big question was whether it beats the current mix of Whisper, ElevenLabs, and newer local models on noisy audio, timestamps, multilingual speech, and jargon. For dictation, the concern was different. People want exact wording, not a model that tidies up what they said and quietly changes meaning.
If you buy or build around speech input, test for your actual failure modes instead of headline benchmark scores. Mixed-language audio, domain terms, silence handling, punctuation, and whether the model paraphrases your words will decide whether this is usable in production.
-
blog.google
- Discuss on HN