---
title: "From Transcription to Live Music: Gemini's Audio Stack — Thor Schaeff, Google DeepMind"
url: https://daily.dev/posts/from-transcription-to-live-music-gemini-s-audio-stack-thor-schaeff-google-deepmind-8nfdaujv7
source_url: https://www.youtube.com/watch?v=Bc6Ojl2XS1w
type: video:youtube
source: "AI Engineer"
published: 2026-06-09T16:43:26.551Z
updated: 2026-06-09T16:43:45.537Z
tags: ["deep-learning", "google-gemini", "audio-processing"]
reading_time: 19
upvotes: 0
comments: 0
language: en
---

> ## Documentation Index
> Fetch the complete documentation index at: https://daily.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# From Transcription to Live Music: Gemini's Audio Stack — Thor Schaeff, Google DeepMind

**[AI Engineer](https://daily.dev/sources/aidotengineer)** · 19 min read · 0 upvotes · 0 comments

## Summary

A Google DeepMind developer advocate presents Gemini's audio capabilities stack at a conference. The talk covers three main areas: audio understanding (transcription with speaker identification, emotion detection, multilingual support via Gemini 3 Flash), speech generation (directing ~30 base voices with accent and style prompts via a 'director's note' approach), and the newly launched Gemini 3.1 Flash Live real-time multimodal model (speech-to-speech via WebSocket, supporting text/audio/video input). A live demo shows the Live model responding with an Irish accent, switching languages, and using a tool to generate a German techno song about the UK startup scene via the Lyra 3 music generation model. Code examples in Python and JavaScript are referenced, along with Google AI Studio as a free experimentation platform.

## Full article

daily.dev links to this article rather than hosting it. Read it at the original source: <https://www.youtube.com/watch?v=Bc6Ojl2XS1w>

---

Tags: [#deep-learning](https://daily.dev/tags/deep-learning), [#google-gemini](https://daily.dev/tags/google-gemini), [#audio-processing](https://daily.dev/tags/audio-processing)

[View this post on daily.dev](https://daily.dev/posts/from-transcription-to-live-music-gemini-s-audio-stack-thor-schaeff-google-deepmind-8nfdaujv7)
