<!-- mobian-agent-page publisher="dailydev" canonical="https://daily.dev/posts/intelligent-transcription-with-gemini-3-5-transcribe-poly4ngnn" -->

---
title: Intelligent transcription with Gemini 3.5 Transcribe
description: Google introduces Gemini 3.5 Transcribe, a new speech-to-text model available via the Live API (streaming) and Interactions API (pre-recorded audio). It offers...
canonical: https://daily.dev/posts/intelligent-transcription-with-gemini-3-5-transcribe-poly4ngnn
twitter:card: summary_large_image
twitter:site: @dailydotdev
og:type: website
og:site_name: daily.dev
og:title: Intelligent transcription with Gemini 3.5 Transcribe | daily.dev
og:description: Google introduces Gemini 3.5 Transcribe, a new speech-to-text model available via the Live API (streaming) and Interactions API (pre-recorded audio). It offers...
og:url: https://daily.dev/posts/intelligent-transcription-with-gemini-3-5-transcribe-poly4ngnn
og:image: https://api.daily.dev/og/posts/POLY4NgNn.png
og:image:alt: Intelligent transcription with Gemini 3.5 Transcribe
og:image:width: 1200
og:image:height: 630
og:locale: en
---

[DeepMind](https://daily.dev/sources/dm)

[Read post](https://api.daily.dev/r/POLY4NgNn)

# [Intelligent transcription with Gemini 3.5 Transcribe](https://api.daily.dev/r/POLY4NgNn "Go to post")

Google introduces Gemini 3.5 Transcribe, a new speech-to-text model available via the Live API (streaming) and Interactions API (pre-recorded audio). It offers self-correction handling, filler-word removal, custom vocabulary support, speaker attribution, and over 85 languages. Compared to the previous Chirp 3 model, it achieves lower word error rates (4.0% streaming, 2.6% non-streaming per Artificial Analysis) and 70% faster time-to-final-transcription. It's rolling out in Google AI Studio, Gemini Enterprise Agent Platform, Gboard's Rambler feature, Google Antigravity, the Gemini app on macOS, and soon Chrome, with third-party platforms like LiveKit, Pipecat, and Vercel already integrating it.

[#google](/tags/google "Check all #google posts")[#google-gemini](/tags/google-gemini "Check all #google-gemini posts")[#speech-recognition](/tags/speech-recognition "Check all #speech-recognition posts")[#voice-ai](/tags/voice-ai "Check all #voice-ai posts")

Yesterday•5m read time•From [blog.google](https://api.daily.dev/r/POLY4NgNn "blog.google")

[![Post cover image](https://media.daily.dev/image/upload/f_auto,q_auto/v1/posts/57ca82980b7b1524508eaba413ca9a16?_a=AQAEuop)](https://api.daily.dev/r/POLY4NgNn "Go to post")

Questions this post answers

What is Gemini 3.5 Transcribe and how does it compare to Chirp 3?

Gemini 3.5 Transcribe is Google's newest speech-to-text model, replacing the previous Chirp 3 model with improved accuracy and speed. As measured by Artificial Analysis, it achieves a 4.0% word error rate for streaming and 2.6% for non-streaming transcription, plus a 70% improvement in time to final transcription over Chirp 3\. Developers picking a speech-to-text model can track releases like this one on daily.dev.

How can I access Gemini 3.5 Transcribe as a developer?

It is available in public preview through the Gemini API in Google AI Studio and Google Antigravity, using the model ids gemini-3.5-transcribe-live for real-time streaming via the Live API and gemini-3.5-transcribe for pre-recorded audio via the Interactions API. Enterprises can access it via the Gemini Enterprise Agent Platform, with Gemini Enterprise for Customer Experience coming soon. Track new model previews and API rollouts like this one on daily.dev before wiring them into production.

How many speakers can Gemini 3.5 Transcribe identify in pre-recorded audio?

It accurately attributes speech with timestamps for up to three speakers in pre-recorded audio, while support for more than three speakers is labeled experimental. It also supports over 85 languages with automatic detection and handles regional accents and dialects. Compare voice transcription capabilities like this before choosing a model on daily.dev.

Comment

Bookmark

Copy

![Placeholder image for anonymous user](https://media.daily.dev/image/upload/s--qsFuKGv_--/t_logo,f_auto/public/noProfile)Share your thoughtsPost

[![DeepMind's image](https://media.daily.dev/image/upload/t_logo,f_auto/v1/logos/dm)](https://daily.dev/sources/dm)

[DeepMind](https://daily.dev/sources/dm "https://daily.dev/sources/dm")

DM provides a diverse range of content spanning technology, business, and culture, offering articles... Read more

60 Followers

•

127 Upvotes

#### Would you recommend this post?

Copy link

WhatsApp

Facebook

X

New Squad

Copy linkShare with your friends

```json
{"@context":"https://schema.org","@graph":[{"@type":"Organization","@id":"https://daily.dev/#organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180},"sameAs":["https://twitter.com/dailydotdev","https://github.com/dailydotdev","https://www.linkedin.com/company/daily-dev-ltd"]},{"@type":"WebSite","@id":"https://daily.dev/#website","url":"https://daily.dev","name":"daily.dev","publisher":{"@id":"https://daily.dev/#organization"},"potentialAction":{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https://daily.dev/search?q={search_term_string}"},"query-input":"required name=search_term_string"}}]}
{"@context":"https://schema.org","@type":"TechArticle","headline":"Intelligent transcription with Gemini 3.5 Transcribe","url":"https://daily.dev/posts/intelligent-transcription-with-gemini-3-5-transcribe-poly4ngnn","mainEntityOfPage":{"@type":"WebPage","@id":"https://daily.dev/posts/intelligent-transcription-with-gemini-3-5-transcribe-poly4ngnn"},"datePublished":"2026-08-26T17:21:02.683Z","dateModified":"2026-08-26T21:57:26.227Z","description":"Google introduces Gemini 3.5 Transcribe, a new speech-to-text model available via the Live API (streaming) and Interactions API (pre-recorded audio). It offers...","image":"https://media.daily.dev/image/upload/f_auto,q_auto/v1/posts/57ca82980b7b1524508eaba413ca9a16?_a=AQAEuop","thumbnailUrl":"https://media.daily.dev/image/upload/f_auto,q_auto/v1/posts/57ca82980b7b1524508eaba413ca9a16?_a=AQAEuop","isAccessibleForFree":true,"articleSection":"DeepMind","inLanguage":"en","publisher":{"@type":"Organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180}},"author":{"@type":"Organization","name":"DeepMind","logo":"https://media.daily.dev/image/upload/t_logo,f_auto/v1/logos/dm","url":"https://daily.dev/sources/dm"},"commentCount":0,"discussionUrl":"https://daily.dev/posts/intelligent-transcription-with-gemini-3-5-transcribe-poly4ngnn","interactionStatistic":[{"@type":"InteractionCounter","interactionType":{"@type":"LikeAction"},"userInteractionCount":0},{"@type":"InteractionCounter","interactionType":{"@type":"CommentAction"},"userInteractionCount":0}],"keywords":"google,google-gemini,speech-recognition,voice-ai","timeRequired":"PT5M"}
{"@context":"https://schema.org","@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https://daily.dev"},{"@type":"ListItem","position":2,"name":"DeepMind","item":"https://daily.dev/sources/dm"},{"@type":"ListItem","position":3,"name":"Intelligent transcription with Gemini 3.5 Transcribe"}]}
{"@context":"https://schema.org","@type":"FAQPage","@id":"https://daily.dev/posts/intelligent-transcription-with-gemini-3-5-transcribe-poly4ngnn#faq","mainEntity":[{"@type":"Question","name":"What is Gemini 3.5 Transcribe and how does it compare to Chirp 3?","acceptedAnswer":{"@type":"Answer","text":"Gemini 3.5 Transcribe is Google's newest speech-to-text model, replacing the previous Chirp 3 model with improved accuracy and speed. As measured by Artificial Analysis, it achieves a 4.0% word error rate for streaming and 2.6% for non-streaming transcription, plus a 70% improvement in time to final transcription over Chirp 3. Developers picking a speech-to-text model can track releases like this one on daily.dev."}},{"@type":"Question","name":"How can I access Gemini 3.5 Transcribe as a developer?","acceptedAnswer":{"@type":"Answer","text":"It is available in public preview through the Gemini API in Google AI Studio and Google Antigravity, using the model ids gemini-3.5-transcribe-live for real-time streaming via the Live API and gemini-3.5-transcribe for pre-recorded audio via the Interactions API. Enterprises can access it via the Gemini Enterprise Agent Platform, with Gemini Enterprise for Customer Experience coming soon. Track new model previews and API rollouts like this one on daily.dev before wiring them into production."}},{"@type":"Question","name":"How many speakers can Gemini 3.5 Transcribe identify in pre-recorded audio?","acceptedAnswer":{"@type":"Answer","text":"It accurately attributes speech with timestamps for up to three speakers in pre-recorded audio, while support for more than three speakers is labeled experimental. It also supports over 85 languages with automatic detection and handles regional accents and dialects. Compare voice transcription capabilities like this before choosing a model on daily.dev."}}]}
```

