|

Google AI Releases Gemini 3.5 Transcribe: A Speech-to-Text Model Reporting 2.6% Average WER Across 85+ Languages

Google has launched Gemini 3.5 Transcribe, a speech-to-text mannequin for real-time voice interfaces and recorded audio. It ships as two endpoints, not one. gemini-3.5-transcribe handles pre-recorded information by the Interactions API. gemini-3.5-transcribe-live handles bidirectional streaming by the Live API. Google stories common phrase error charges of 4.0% streaming and a pair of.6% non-streaming, as measured by Artificial Analysis. Time to remaining transcription improves 70% over Chirp 3, the earlier mannequin. Automatic detection covers greater than 85 languages, together with mid-sentence code-switching. The break up between the 2 endpoints is the half value planning round. They don’t share the identical characteristic set, limits, or worth.

Is it deployable?

Yes, however API-only. There aren’t any open weights and no self-hosted path. This is a managed-service determination, not an infrastructure one.

  • Company stage: Any. Solo builders and startups can begin on the Gemini API free tier by way of Google AI Studio. Mid-market groups transfer to the paid tier for increased charge limits. The paid tier additionally ensures content material shouldn’t be used to enhance Google’s merchandise. Regulated enterprises route by the Gemini Enterprise Agent Platform, which provides provisioned throughput, compliance controls, and quantity reductions. Both developer and enterprise tracks are in public preview, so deal with manufacturing commitments accordingly.
  • Industries: Contact facilities and CX platforms, scientific documentation, media captioning and localization, authorized and insurance coverage consumption, assembly tooling, and voice-driven developer instruments.
  • Applications: Real-time voice brokers, stay captioning, post-call analytics pipelines, assembly transcription with speaker attribution, dictation, and voice-controlled interfaces.