Technology
The Verge

Google’s new AI transcription edits out your ‘ums’ and ‘ahs’

Source Entity

Jess Weatherbed

August 27, 2026
Google’s new AI transcription edits out your ‘ums’ and ‘ahs’

Google has launched Gemini 3.5 Transcribe, a new AI model designed for high-precision, noise-resistant speech-to-text conversion. The tool automatically removes disfluencies like 'ums' and 'ahs' while supporting over 85 languages and specialized jargon.

The Evolution of Google’s Voice Intelligence

Google has officially expanded its AI portfolio with the introduction of Gemini 3.5 Transcribe, a specialized model aimed at refining the speech-to-text experience. While users continue to anticipate the broader release of the Gemini 3.5 Pro model, this new addition to the 3.5 branch focuses specifically on the challenges of natural language processing in real-world environments. By prioritizing the cleanup of human speech patterns, Google is positioning its AI to become an indispensable tool for productivity and accessibility.

Overcoming the Barriers of Human Speech

The core innovation behind Gemini 3.5 Transcribe lies in its ability to handle the inconsistencies of human communication. Conventional transcription models have historically struggled with background noise, interruptions, and the natural disfluencies—such as 'ums,' 'ahs,' and self-corrections—that define spontaneous speech. Gemini 3.5 Transcribe addresses these issues by converting raw audio directly into polished, formatted text, effectively acting as an intelligent editor that processes input in real-time.

Technical Advancements and Performance Metrics

Compared to its predecessor, Chirp 3, the new model demonstrates significant technical improvements. Google reports that Gemini 3.5 Transcribe is approximately 70 percent faster, drastically reducing the latency between voice input and final text output. Furthermore, the live-speech error rate has been refined to 5.5 percent, offering a measurable improvement over Chirp 3’s 7.32 percent error rate. These metrics suggest a focused effort to make voice interaction feel more fluid and less prone to the mechanical delays that often hinder user adoption.

Broad Ecosystem Integration

The utility of Gemini 3.5 Transcribe is not limited to a single application. It is already integrated into the Gboard “Rambler” feature on the Pixel 11 and is accessible via the Gemini app on macOS. By making this model available to developers through the Gemini API in Google AI Studio and the Gemini Enterprise Agent Platform, Google is facilitating the integration of these advanced transcription capabilities into a wider array of third-party products and services, potentially standardizing high-quality voice input across the Android and Chrome ecosystems.

Global Accessibility and Jargon Recognition

Beyond simple transcription, the model supports more than 85 languages and is specifically optimized to detect and interpret specialized jargon. This is a critical development for professional sectors—such as legal, medical, or technical fields—where accurate terminology is essential. By automating the recognition of complex vocabulary, Google is reducing the need for manual post-transcription editing, thereby increasing the overall efficiency for professionals who rely on voice-to-text workflows.

Future Trends in AI-Driven Communication

The launch of Gemini 3.5 Transcribe signals a shift toward 'invisible' AI, where the software works in the background to improve communication without requiring constant user intervention. As this technology becomes more deeply embedded in browsers, mobile operating systems, and developer tools, we can expect voice-controlled AI to become the standard for human-computer interaction. The focus on 'polished' output suggests that Google envisions a future where the spoken word is treated with the same precision as written documentation, bridging the gap between casual conversation and professional output.

Verification Required?

Read the full report from the primary source

Go to The Verge