More on the topic...
Generating detailed summary...
Failed to generate summary. Please try again.
Google just rolled out Gemini 3.5 Live Translate, a speech-to-speech model that catches more than 70 languages on the fly and spits out natural-sounding audio. Instead of waiting for you to finish each sentence, it translates in real time, tracking your intonation, pacing and pitch. The delay stays just a few seconds behind, so conversations flow without awkward silences.
Developers can try it now through the Gemini Live API in public preview or via Google AI Studio. Enterprises get private preview in Google Meet this month, and everyone on Android and iOS will see it inside the Google Translate app. No manual language setup needed—you speak, it detects, it translates. Background noise won’t trip it up, making it handy for calls, lessons, conferences or live broadcasts.
Platforms such as Agora, Fishjam, LiveKit, Pipecat and Vision Agents have already hooked into the Gemini Live API, handling real-time streaming so developers can focus on UX. Grab is testing this for drivers and riders, spanning 10 million voice calls each month. CJ ENM and other early users praise the accuracy, low latency and smooth audio.
In Google Meet, the model expands speech translation from five languages to 70+, covering over 2,000 language pairs instead of just English-based ones. On mobile, a new “listening mode” for Android pipes translations through your earpiece—no headphones needed. All output carries an invisible SynthID watermark, so AI-generated audio stays traceable and helps curb misinformation.
Questions about this article
No questions yet.