A Google DeepMind developer advocate presents Gemini's audio capabilities stack at a conference. The talk covers three main areas: audio understanding (transcription with speaker identification, emotion detection, multilingual support via Gemini 3 Flash), speech generation (directing ~30 base voices with accent and style prompts via a 'director's note' approach), and the newly launched Gemini 3.1 Flash Live real-time multimodal model (speech-to-speech via WebSocket, supporting text/audio/video input). A live demo shows the Live model responding with an Irish accent, switching languages, and using a tool to generate a German techno song about the UK startup scene via the Lyra 3 music generation model. Code examples in Python and JavaScript are referenced, along with Google AI Studio as a free experimentation platform.

19m watch time
132 Impressions