<!-- mobian-agent-page publisher="dailydev" canonical="https://daily.dev/posts/gemini-3-8-flash-tts-with-voice-cloning-oanjbesje" -->

---
title: Gemini 3.8 Flash TTS  with Voice Cloning | daily.dev
description: A walkthrough of Google&#x27;s two new Gemini 3.8 Flash TTS models (Flash TTS and Flash Light TTS), covering voice design via natural-language descriptions, a...
canonical: https://daily.dev/posts/gemini-3-8-flash-tts-with-voice-cloning-oanjbesje
twitter:card: summary_large_image
twitter:site: @dailydotdev
og:type: website
og:site_name: daily.dev
og:title: Gemini 3.8 Flash TTS  with Voice Cloning | daily.dev
og:description: A walkthrough of Google&#x27;s two new Gemini 3.8 Flash TTS models (Flash TTS and Flash Light TTS), covering voice design via natural-language descriptions, a...
og:url: https://daily.dev/posts/gemini-3-8-flash-tts-with-voice-cloning-oanjbesje
og:image: https://api.daily.dev/og/posts/OanJbesjE.png
og:image:alt: Gemini 3.8 Flash TTS  with Voice Cloning
og:image:width: 1200
og:image:height: 630
og:locale: en
---

> ## Documentation Index
> Fetch the complete documentation index at: https://daily.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# Gemini 3.8 Flash TTS  with Voice Cloning

**[Sam Witteveen](https://daily.dev/sources/samwitteveenai)** · 17 min read · 0 upvotes · 0 comments

## Summary

A walkthrough of Google's two new Gemini 3.8 Flash TTS models (Flash TTS and Flash Light TTS), covering voice design via natural-language descriptions, a 2,000+ voice library, newly launched voice cloning with consent-clip safeguards and SynthID/C2PA watermarking, stage-direction and inline emotion tags, multi-speaker dialogue generation, and hands-on demos in AI Studio and Colab. The presenter also scrutinizes Google's benchmark claims, noting a conflict of interest with Hume AI's numbers, and compares independent Artificial Analysis rankings against competitors like Qwen Audio 3, Cartesia, and ElevenLabs, along with pricing comparisons between the two new models.

## Full article

daily.dev links to this article rather than hosting it. Read it at the original source: <https://www.youtube.com/watch?v=GDUBXR-ql78>

## Questions this post answers

### What is the difference between Gemini 3.8 Flash TTS and Gemini 3.8 Flash Light TTS?

Flash TTS is designed for creative, nuance-sensitive work like character voices, audiobooks, podcasts, and games, while Flash Light TTS targets high-volume use cases like dubbing, bulk audio generation, and voice agents where cost matters more than perfect nuance. Flash Light tends to sound slightly more lazy or drawn out, and pricing between the two is not hugely different, with Flash TTS costing about 50% more.

_Developers weighing TTS models for cost versus quality can track comparisons like this on daily.dev._

### How does voice cloning work in Gemini 3.8 TTS and what safeguards does Google use?

A user provides a 30-second reference audio sample plus a separate spoken consent clip that must match the reference speaker before a clone is created. Every generated clip is watermarked with SynthID and carries C2PA credentials, and the feature is unavailable in many countries and some U.S. states due to legal restrictions around voice cloning.

_Anyone building voice features that touch consent and safety should follow rollouts like this via daily.dev._

### How does Gemini 3.8 Flash TTS compare to competitors like Qwen Audio 3 TTS Plus on independent benchmarks?

On Artificial Analysis's preference ELO ranking, an independent test, Gemini 3.8 Flash TTS ranks second while Flash Light TTS ranks sixth, with Flash TTS placing just ahead of Qwen Audio 3 TTS Plus, which remains closed at the time of testing. Gemini models lead on pronunciation benchmarks, useful for reading numbers and IDs aloud, and on multilingual support, though Cartesia and ElevenLabs are making strong moves in multilingual TTS.

_Teams picking a TTS provider can weigh independent benchmark shifts like these through daily.dev._

## Similar posts on daily.dev

- [Gemini 3.1 Flash TTS: the next generation of expressive AI speech](https://daily.dev/posts/gemini-3-1-flash-tts-the-next-generation-of-expressive-ai-speech-1av40deos) · DeepMind · 9 upvotes · 1 comments
- [Gemini 3.1 Flash TTS on Google Cloud](https://daily.dev/posts/gemini-3-1-flash-tts-on-google-cloud-pl4nonfnc) · Google Cloud · 0 upvotes · 0 comments

---

Tags: [#ai](https://daily.dev/tags/ai), [#google-gemini](https://daily.dev/tags/google-gemini), [#text-to-speech](https://daily.dev/tags/text-to-speech), [#google-deepmind](https://daily.dev/tags/google-deepmind), [#voice-cloning](https://daily.dev/tags/voice-cloning)

[View this post on daily.dev](https://daily.dev/posts/gemini-3-8-flash-tts-with-voice-cloning-oanjbesje)

```json
{"@context":"https://schema.org","@graph":[{"@type":"Organization","@id":"https://daily.dev/#organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180},"sameAs":["https://twitter.com/dailydotdev","https://github.com/dailydotdev","https://www.linkedin.com/company/daily-dev-ltd"]},{"@type":"WebSite","@id":"https://daily.dev/#website","url":"https://daily.dev","name":"daily.dev","publisher":{"@id":"https://daily.dev/#organization"},"potentialAction":{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https://daily.dev/search?q={search_term_string}"},"query-input":"required name=search_term_string"}}]}
{"@context":"https://schema.org","@type":"TechArticle","headline":"Gemini 3.8 Flash TTS  with Voice Cloning","url":"https://daily.dev/posts/gemini-3-8-flash-tts-with-voice-cloning-oanjbesje","mainEntityOfPage":{"@type":"WebPage","@id":"https://daily.dev/posts/gemini-3-8-flash-tts-with-voice-cloning-oanjbesje"},"datePublished":"2026-09-24T17:27:55.988Z","dateModified":"2026-09-25T20:07:43.121Z","description":"A walkthrough of Google's two new Gemini 3.8 Flash TTS models (Flash TTS and Flash Light TTS), covering voice design via natural-language descriptions, a...","image":"https://i.ytimg.com/vi/GDUBXR-ql78/sddefault.jpg","thumbnailUrl":"https://i.ytimg.com/vi/GDUBXR-ql78/sddefault.jpg","isAccessibleForFree":true,"articleSection":"Sam Witteveen","inLanguage":"en","publisher":{"@type":"Organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180}},"author":{"@type":"Organization","name":"Sam Witteveen","logo":"https://media.daily.dev/image/upload/s--gJm-KsgL--/f_auto/v1711727006/logos/samwitteveenai","url":"https://daily.dev/sources/samwitteveenai"},"commentCount":0,"discussionUrl":"https://daily.dev/posts/gemini-3-8-flash-tts-with-voice-cloning-oanjbesje","interactionStatistic":[{"@type":"InteractionCounter","interactionType":{"@type":"LikeAction"},"userInteractionCount":0},{"@type":"InteractionCounter","interactionType":{"@type":"CommentAction"},"userInteractionCount":0}],"keywords":"ai,google-gemini,text-to-speech,google-deepmind,voice-cloning","timeRequired":"PT17M","video":{"@type":"VideoObject","name":"Gemini 3.8 Flash TTS  with Voice Cloning","description":"A walkthrough of Google's two new Gemini 3.8 Flash TTS models (Flash TTS and Flash Light TTS), covering voice design via natural-language descriptions, a...","thumbnailUrl":"https://i.ytimg.com/vi/GDUBXR-ql78/sddefault.jpg","uploadDate":"2026-09-24T17:27:55.988Z","duration":"PT17M","url":"https://api.daily.dev/r/OanJbesjE","embedUrl":"https://www.youtube.com/embed/GDUBXR-ql78"}}
{"@context":"https://schema.org","@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https://daily.dev"},{"@type":"ListItem","position":2,"name":"Sam Witteveen","item":"https://daily.dev/sources/samwitteveenai"},{"@type":"ListItem","position":3,"name":"Gemini 3.8 Flash TTS  with Voice Cloning"}]}
{"@context":"https://schema.org","@type":"FAQPage","@id":"https://daily.dev/posts/gemini-3-8-flash-tts-with-voice-cloning-oanjbesje#faq","mainEntity":[{"@type":"Question","name":"What is the difference between Gemini 3.8 Flash TTS and Gemini 3.8 Flash Light TTS?","acceptedAnswer":{"@type":"Answer","text":"Flash TTS is designed for creative, nuance-sensitive work like character voices, audiobooks, podcasts, and games, while Flash Light TTS targets high-volume use cases like dubbing, bulk audio generation, and voice agents where cost matters more than perfect nuance. Flash Light tends to sound slightly more lazy or drawn out, and pricing between the two is not hugely different, with Flash TTS costing about 50% more. Developers weighing TTS models for cost versus quality can track comparisons like this on daily.dev."}},{"@type":"Question","name":"How does voice cloning work in Gemini 3.8 TTS and what safeguards does Google use?","acceptedAnswer":{"@type":"Answer","text":"A user provides a 30-second reference audio sample plus a separate spoken consent clip that must match the reference speaker before a clone is created. Every generated clip is watermarked with SynthID and carries C2PA credentials, and the feature is unavailable in many countries and some U.S. states due to legal restrictions around voice cloning. Anyone building voice features that touch consent and safety should follow rollouts like this via daily.dev."}},{"@type":"Question","name":"How does Gemini 3.8 Flash TTS compare to competitors like Qwen Audio 3 TTS Plus on independent benchmarks?","acceptedAnswer":{"@type":"Answer","text":"On Artificial Analysis's preference ELO ranking, an independent test, Gemini 3.8 Flash TTS ranks second while Flash Light TTS ranks sixth, with Flash TTS placing just ahead of Qwen Audio 3 TTS Plus, which remains closed at the time of testing. Gemini models lead on pronunciation benchmarks, useful for reading numbers and IDs aloud, and on multilingual support, though Cartesia and ElevenLabs are making strong moves in multilingual TTS. Teams picking a TTS provider can weigh independent benchmark shifts like these through daily.dev."}}]}
```

