Google shipped a speech-to-text model on 26 August 2026. Gemini 3.5 Transcribe. Not a new chat model. Not Gemini 3.5 Live. Transcribe.
I care about one job: you talk like a person, it types like you meant it. Fillers out. Self-corrections kept. Language detected without you picking a menu. That is Google’s claim. Here is where it actually lives, and how I would try it on a messy recording.
Where Gemini 3.5 Transcribe is today
Google’s post splits it three ways. I am sticking to that list, not a rumor board.
- Phone, if you are in a supported country: Rambler on Gboard for Android. You ramble. It formats. You can also edit by voice: fix a spelling, change the tone.
- Mac, English only for now: the Gemini macOS app. Talk, get clean text. Google also says you can issue voice commands that hand work to other Gemini models (summarize a local file, move text, generate an image at the cursor). That last part is their product pitch. I have not run it.
- Developers: public preview in the Gemini API, through Google AI Studio, plus Gemini Enterprise Agent Platform. Two endpoints:
gemini-3.5-transcribe-liveon the Live API for streaming, andgemini-3.5-transcribeon the Interactions API for a file you already recorded.
Chrome “talk to type” in any web field is coming later. It is not shipping with this drop. Do not tell someone to open Chrome and look for it today.
The Verge has the same 26 August date and the same product. After they published, Google told them 3.5 Live was not launching. I am not writing Live into this piece.
How to turn it on
Android, Rambler. You need a current Gboard in a country Google has opened. Open any text field. Look for Rambler in the Gboard extras, not the old one-shot mic if both are there. Hold and talk. If you do not see it, you are not in the rollout. That is a geo gate, not a broken install.
Mac. Install or update the Gemini app. Language has to be English for this. Open a compose box, hit the voice control, talk. If you want it to act on a file or the screen, that is a separate permission Google mentions. I would leave screen context off until you trust the first clean transcript.
A recording you already have. Developers: AI Studio, pick gemini-3.5-transcribe, upload the file. Everyone else: the consumer apps Google listed are live-ish. The post does not give a “drop an MP3 in the Gemini web app” button for Transcribe. I am not inventing one.
What it does to a messy recording
Google’s own example is the one I would test first. You say, “let’s meet Tuesday, no, Wednesday.” A dumb dictation keeps both days. 3.5 Transcribe is built to keep Wednesday and drop the false start. Same pass strips “ums” and “ahs” and formats the sentence. That is the product. Not a word-error trophy.

It auto-detects more than 85 languages, including a switch mid-stream, per Google. You can feed a custom vocabulary so a last name or a SKU does not get “helpfully” fixed. On a file, it can label up to three speakers with word-level timestamps. More than three is experimental.
If I were checking it on a real clip, I would use a two-minute voice note with: one self-correction, a bunch of fillers, a Spanish-English hop (I do that without thinking), and a proper noun it has never seen. Then I would look at three things only. Did the correction stick. Did the fillers die. Did it keep the name. Everything else is marketing.
The scoreboard is a vendor scoreboard
Google cites Artificial Analysis for the numbers, not an independent paper I can re-run. Treat them as vendor-reported:
- Average word error rate of 4.0% streaming and 2.6% non-streaming.
- Time to final transcription 70% faster than Chirp 3, their previous model.
- On FLEURS, across a set of top languages, 5.50% WER streaming and 5.04% non-streaming, better than Chirp 3 on that same bench.
Those are Google’s citations of Artificial Analysis. I am not converting them into “it’s the most accurate STT.” I am saying: if the cleanup works on your voice note, you will feel it before you feel 2.6%.
Try Rambler if you have it. Try the Mac app if you speak English there. Developers can hit the preview in AI Studio. Leave Chrome alone until Google actually ships talk-to-type. And if a headline nearby says 3.5 Live launched yesterday, that is the correction Google already sent The Verge.