Gemini 3.5 Transcribe now available on AI Gateway
Google's Gemini 3.5 Transcribe model is now available on AI Gateway, supporting both streaming and batch audio transcription with multi-language detection and custom vocabulary. The AI SDK V7 adds streamTranscribeReadable for real-time streaming transcription via WebSocket.
from Google is now available on AI Gateway. It takes audio and returns text, in two variants:Gemini 3.5 Transcribe
The model detects the language on its own, covers 85+, and follows a speaker who switches language partway through. You can also supply custom vocabulary so it recognizes names, jargon, and spellings.
Streaming transcription is new in AI SDK V7:
opens the socket and takes a of raw audio chunks, so you can pass a microphone straight through. Tell it the format you are sending with :streamTranscribeReadableStreaminputAudioFormat
For audio you already have on disk, sends it in one request and returns the text:transcribe
You can also try the model without writing any code. Open and send audio to read the transcript in the browser.Gemini 3.5 Transcribe Live
AI Gateway provides a unified API for calling models, tracking usage and cost, failover, and performance optimizations for higher-than-provider uptime. It includes built-in , , , and more.custom reportingbudgets for API keysrouting rules
AI Gateway reflects provider pricing with no markup and does not charge a platform fee on inference, including on (BYOK) requests.Bring Your Own Key
You can view available on AI Gateway, or start from the .all transcription modelsspeech quickstart
transcribes a complete recording in a single request.
google/gemini-3.5-transcribetranscribes audio over a WebSocket, returning a transcript that updates while the recording is still going.
google/gemini-3.5-transcribe-live
Live transcription
Complete recordings
Source: original entry ↗