LLM & Text GenerationDev5 min reading time

Google AI Releases Gemini 3.5 Transcribe: A Speech-to-Text Model Reporting 2.6% Average WER Across 85+ Languages

MarkTechPost
Read full post
Google has launched Gemini 3.5 Transcribe, a speech-to-text model supporting over 85 languages with a 2.6% average word error rate for recorded audio and 4.0% for streaming. It offers two APIs: one for pre-recorded files and another for live streaming, each with distinct features and limits. The model is accessible via API only, with no open weights or self-hosting options, targeting industries like contact centers, clinical documentation, and media captioning.

More on this story


More in LLM & Text Generation

DeepSeek V4.1 Flash now available on AI Gateway

Covered by 2 sources

Cohere Debuts Open-Weight 218B Mixture-of-Experts Machine Translation Model

Unite.AI