Tech
Google Launches Gemini 3.5 Transcribe as Its Most Precise Speech-to-Text Model
Google today unveiled Gemini 3.5 Transcribe as its “most precise speech-to-text model yet,” with the technology already powering multiple first-party products.
Gemini Live Gains Productivity Features With Spark, Gmail and More
Unlike traditional speech recognition systems that can have difficulty with background sounds, specialized terminology and cleaning up disfluent speech, Gemini 3.5 Transcribe transforms raw audio into accurate, refined and properly formatted text.
The model is “designed to capture your natural speaking style to better understand your intent and recognize custom vocabulary.” As demonstrated in Rambler, Gemini 3.5 Transcribe can recognize self-corrections (“let’s meet Tuesday—no, Wednesday”), eliminate “ums,” “ahs,” and other filler words from the final transcription, automatically format text and support natural voice-based editing. Google also highlights:
- More precise transcription: As measured by Artificial Analysis, it achieves an average Word Error Rate (WER) of 4.0% for streaming and 2.6% for non-streaming use-cases. The model performs strongly in noisy, real-world environments and can accurately capture alphanumeric information such as postal codes and order IDs.
- Custom vocabulary: The system can identify specialized terminology and unusual spellings by smoothly adapting transcriptions to the custom vocabulary supplied by users.
- Global language support: It automatically detects and transcribes more than 85 languages while handling regional accents and varied dialects.
- Multi-speaker identification: It can accurately assign speech to speakers in pre-recorded audio and provide timestamps for up to three speakers, while support for 3+ speakers remains experimental.
From a performance perspective, Gemini 3.5 Transcribe delivers what Google describes as a “major advancement” in its capabilities, with better word error rates and substantially improved latency compared with its Chirp 3 transcription model from 2025.
Another objective is enabling users to “execute tasks with your voice.” Through function calling, Gemini 3.5 Transcribe can “delegate complex tasks (such as image generation and file analysis) to other Gemini models.” The capability is demonstrated through Speak to Window in the Gemini app for macOS.
In addition to the Gemini macOS app and Gboard Rambler on Android, Gemini 3.5 Transcribe is available through the microphone in Google Antigravity’s prompt box, where it “pairs screen context and chat history, with your permission, to ensure pinpoint transcription accuracy across file names, agent thoughts, and active documents.”
The technology is also coming to the Chrome browser, allowing users to “talk to type in any web field — making it effortless to dictate replies, draft posts, or prompt Gemini in Chrome more naturally and easily with your voice.”
Gemini 3.5 Transcribe is available:
- For developers: In public preview through the Gemini API via Google AI Studio and Google Antigravity.
- For enterprises: In public preview through Gemini Enterprise Agent Platform and coming soon to Gemini Enterprise for Customer Experience.
-
Entertainment3 weeks agoUniversal Becomes First Studio to Pass $4 Billion at 2026 Global Box Office
-
Business2 weeks ago
Black Banx H1 2026 Results Reveal the Power of Operating Leverage
-
Sports4 days agoConcacaf Nations League 2026/27: Groups and Schedule
-
Festivals & Events1 week agoGamescom 2026 Opening Night Live: Schedule, Date, Streaming Details, Expected Games and How to Watch
-
Business3 weeks agoWhat to Know before Using Insurance for Scratches, Dents and Small Repairs
-
Tech1 week agoYouTube View Count Changes: What Creators and Marketers Need
-
Science3 weeks ago2026 Total Solar Eclipse: Where to Watch and How to View It Safely
-
Entertainment3 weeks agoMeet IKON: The Independent Promoter Powering Ye’s U.S. Stadium Run

