Quick answer
Google introduced Gemini 3.5 Transcribe on August 26, 2026 as a speech-to-text model for live streaming and pre-recorded audio. It can clean up disfluencies, apply custom vocabulary, detect more than 85 languages, and add speaker labels and word-level timestamps to recorded audio. The model is in public preview through Google developer and enterprise products, while selected consumer features are already using it. Availability in Chat AI has not been verified.
Quick answer: Gemini 3.5 Transcribe is a specialized speech-to-text model
Gemini 3.5 Transcribe converts spoken audio into formatted text rather than acting as a general-purpose chat model. Google offers separate paths for continuous live transcription and complete pre-recorded files. The live path is intended for voice interfaces and captions, while the recorded path adds features such as speaker attribution and word-level timestamps.
Sources: Google, Google Cloud
Smart transcription produces cleaner text than a literal transcript
Google says the model can remove filler words, resolve spoken self-corrections, format text, and adapt to a supplied custom vocabulary. Those features can improve notes and dictation, but the polished result may differ from the exact sequence of words that was spoken. For interviews, legal records, research, or other tasks where wording matters, keep the source audio and compare important passages against it.
Sources: Google, Google DeepMind
The model supports multilingual and multi-speaker audio
Google documents automatic transcription across more than 85 languages and support for language changes during a live stream. For pre-recorded audio, the model can identify up to three speakers and return word-level timestamps; Google labels support beyond three speakers as experimental. Accuracy can still vary with accents, overlapping speech, background noise, microphones, and domain-specific terms.
Sources: Google, Google DeepMind, Google Cloud
Independent word-error benchmarks need workload-specific testing
Google cites Artificial Analysis results of 4.0% average word error rate for streaming and 2.6% for non-streaming transcription. Artificial Analysis describes its score as a weighted average across roughly eight hours of English audio drawn from voice-agent, parliamentary, and earnings-call datasets. A lower word error rate is useful evidence, but it does not predict performance for every language, accent, acoustic environment, or specialized vocabulary.
Sources: Google, Artificial Analysis
Developer access is in public preview
Google says developers can use the model through the Gemini API in Google AI Studio, with live streaming and recorded-audio interfaces, and enterprises can access it through Gemini Enterprise Agent Platform. The model card lists a 96K-token input window and a 32K-token output window. Preview status means interfaces, limits, regional access, and behavior may change, so production teams should verify the current documentation before committing to a workflow.
Sources: Google, Google DeepMind, Google Cloud
Selected Google products already use the transcription model
Google says Gemini 3.5 Transcribe is available in English in the Gemini app on macOS and powers Rambler on Android in selected countries and languages. It is also used with screen context in Google Antigravity, and Google says voice typing with the model is coming to Chrome. These product rollouts are separate from general API access and can differ by platform, country, language, and account.
Sources: Google
Gemini 3.5 Transcribe is not yet verified in Chat AI
Gemini 3.5 Transcribe does not appear in Chat AI's public model directory at publication time, so this article does not claim that the model is available in the app. Check the current directory and in-app model selector before assuming access; model availability can vary by plan, platform, region, provider access, and app version.
Sources: Chat AI
Frequently asked questions
What readers usually ask
What is Gemini 3.5 Transcribe?
Gemini 3.5 Transcribe is Google's specialized speech-to-text model for live streaming and pre-recorded audio. It can produce formatted transcripts, recognize custom vocabulary, and support multilingual audio.
Does Gemini 3.5 Transcribe work in real time?
Yes. Google provides a live streaming interface for continuous transcription and a separate interface for complete recorded audio.
How many languages does Gemini 3.5 Transcribe support?
Google says it automatically detects and transcribes more than 85 languages, including language changes during live audio. Actual accuracy varies by language, accent, audio quality, and terminology.
Can Gemini 3.5 Transcribe identify different speakers?
For pre-recorded audio, Google documents speaker attribution and word-level timestamps for up to three speakers. Support for more than three speakers is described as experimental.
Is Gemini 3.5 Transcribe available in Chat AI?
Availability in Chat AI was not verified at publication time. Check the current Chat AI model directory or in-app model selector before assuming access.
Evidence
Sources
- Gemini 3.5 Transcribe announcementGoogle · Primary source
- Gemini 3.5 Audio model cardGoogle DeepMind · Primary source
- Gemini 3.5 Transcribe documentationGoogle Cloud · Primary source
- Speech-to-text benchmarking methodologyArtificial Analysis · Secondary source
- Chat AI model directoryChat AI · Primary source