August 27, 2026

News, minutes after it breaks

Latest

Home  / Artificial intelligence

Google’s New AI Transcription Edits Out Your “Ums” and “Ahs”

Image: The Verge

Google is giving speech-to-text technology a major upgrade with Gemini 3.5 Transcribe, a new AI model designed to turn natural, messy speech into cleaner and more polished text.

Unlike traditional transcription tools that attempt to capture every word exactly as spoken, Gemini 3.5 Transcribe can understand context, remove filler words such as “um” and “uh,” recognize self-corrections and automatically format the final transcript. Google describes it as its most precise speech-to-text model yet.

No More “Ums” and “Ahs”

One of the most noticeable improvements is the model’s ability to clean up disfluencies.

If someone says, “Let’s meet Tuesday—no, Wednesday,” the system can understand the correction and produce the intended version rather than simply recording every discarded word. It can also remove unnecessary pauses, repeated words and filler sounds while preserving the speaker’s meaning.

That could make AI transcription significantly more useful for people who dictate emails, notes, messages and documents using their voice.

Gemini Understands How People Actually Speak

Google says Gemini 3.5 Transcribe is designed to handle real-world speech rather than perfectly scripted audio.

The model can deal with background noise, accents, specialized terminology and natural speaking patterns. It can also automatically detect and transcribe more than 85 languages, making it useful for multilingual users and international applications.

Developers can also provide custom vocabulary. This allows the system to better recognize unusual names, technical terminology, product codes and other specialized words.

It Can Format Your Speech Automatically

Gemini 3.5 Transcribe doesn’t simply produce a block of unformatted text.

The model can automatically add useful formatting, punctuation and structure based on the context of the spoken content. That means users can dictate information in a conversational style and receive something closer to finished written text.

Google is also adding voice-based editing capabilities, allowing users to correct details, change wording and refine their text without stopping the transcription process.

Already Powering Google’s Voice Features

The technology is already being integrated into Google’s products.

Gemini 3.5 Transcribe powers Rambler, Google’s intelligent dictation feature on Android, and is also available in the Gemini app for macOS. Google says the technology will also come to Chrome, expanding its reach to browser-based voice typing.

Developers can access the model through Google’s AI development tools, allowing them to build applications around the new transcription capabilities.

Better Accuracy for Real-World Audio

Google says Gemini 3.5 Transcribe significantly improves transcription accuracy compared with its previous Chirp 3 speech model.

According to figures reported by Google, the model achieves an average word error rate of approximately 2.6% for non-streaming use cases and 4.0% for streaming transcription, with strong performance in noisy environments and when processing alphanumeric information such as order numbers and postal codes.

For businesses, this could make the technology particularly useful for meetings, customer-service calls, interviews and automated documentation.

Multiple Speakers and Timestamps

The new model also supports speaker identification for recorded conversations.

Gemini 3.5 Transcribe can attribute speech to multiple speakers and provide word-level timestamps, making it easier to identify who said what and locate specific moments within an audio recording. Google says attribution for up to three speakers is supported, while larger speaker counts remain experimental.

What This Means for Voice AI

The launch reflects a broader shift in how Google approaches speech recognition.

Instead of treating transcription as simply converting audio into text, Gemini 3.5 Transcribe treats spoken language as information that can be understood, cleaned up and acted upon.

That could eventually make voice interfaces feel much more natural. Users won’t necessarily need to speak in perfectly structured sentences or pause to correct every mistake. The AI can increasingly understand what they meant and turn that speech into useful output.

Google’s new model therefore represents more than an improvement to dictation. It is another step toward AI systems that can understand everyday human speech with greater accuracy and context—and turn those conversations directly into useful work.

More in Artificial intelligence

Leave a Reply

Your email address will not be published. Required fields are marked *