Connect with us

Tech

Google Launches Gemini 3.5 Transcribe as Its Most Precise Speech-to-Text Model

Published

on

Google Launches Gemini 3.5 Transcribe as Its Most Precise Speech to Text Model

Google today unveiled Gemini 3.5 Transcribe as its “most precise speech-to-text model yet,” with the technology already powering multiple first-party products.

Gemini Live Gains Productivity Features With Spark, Gmail and More

Unlike traditional speech recognition systems that can have difficulty with background sounds, specialized terminology and cleaning up disfluent speech, Gemini 3.5 Transcribe transforms raw audio into accurate, refined and properly formatted text.

The model is “designed to capture your natural speaking style to better understand your intent and recognize custom vocabulary.” As demonstrated in Rambler, Gemini 3.5 Transcribe can recognize self-corrections (“let’s meet Tuesday—no, Wednesday”), eliminate “ums,” “ahs,” and other filler words from the final transcription, automatically format text and support natural voice-based editing. Google also highlights:

  • More precise transcription: As measured by Artificial Analysis, it achieves an average Word Error Rate (WER) of 4.0% for streaming and 2.6% for non-streaming use-cases. The model performs strongly in noisy, real-world environments and can accurately capture alphanumeric information such as postal codes and order IDs.
  • Custom vocabulary: The system can identify specialized terminology and unusual spellings by smoothly adapting transcriptions to the custom vocabulary supplied by users.
  • Global language support: It automatically detects and transcribes more than 85 languages while handling regional accents and varied dialects.
  • Multi-speaker identification: It can accurately assign speech to speakers in pre-recorded audio and provide timestamps for up to three speakers, while support for 3+ speakers remains experimental.

From a performance perspective, Gemini 3.5 Transcribe delivers what Google describes as a “major advancement” in its capabilities, with better word error rates and substantially improved latency compared with its Chirp 3 transcription model from 2025.

Another objective is enabling users to “execute tasks with your voice.” Through function calling, Gemini 3.5 Transcribe can “delegate complex tasks (such as image generation and file analysis) to other Gemini models.” The capability is demonstrated through Speak to Window in the Gemini app for macOS.

In addition to the Gemini macOS app and Gboard Rambler on Android, Gemini 3.5 Transcribe is available through the microphone in Google Antigravity’s prompt box, where it “pairs screen context and chat history, with your permission, to ensure pinpoint transcription accuracy across file names, agent thoughts, and active documents.”

The technology is also coming to the Chrome browser, allowing users to “talk to type in any web field — making it effortless to dictate replies, draft posts, or prompt Gemini in Chrome more naturally and easily with your voice.”

Gemini 3.5 Transcribe is available:

  • For developers: In public preview through the Gemini API via Google AI Studio and Google Antigravity.
  • For enterprises: In public preview through Gemini Enterprise Agent Platform and coming soon to Gemini Enterprise for Customer Experience.

Advertisement
follow us on google news banner black

Facebook

Recent Posts

Trending

error: Content is protected !!