<!-- mobian-agent-page publisher="dailydev" canonical="https://daily.dev/posts/gemini-3-8-live-thinks-in-the-background-while-you-talk-and-the-benchmarks-are-hard-to-ignore-vqety7y4w" -->

---
title: Gemini 3.8 Live thinks in the background while you talk,...
description: Google has released Gemini 3.8 Live and 3.8 Live Extended Thinking, models that can reason and run tool calls in the background without pausing a live...
canonical: https://daily.dev/posts/gemini-3-8-live-thinks-in-the-background-while-you-talk-and-the-benchmarks-are-hard-to-ignore-vqety7y4w
twitter:card: summary_large_image
twitter:site: @dailydotdev
og:type: website
og:site_name: daily.dev
og:title: Gemini 3.8 Live thinks in the background while you talk, and the benchmarks are hard to ignore | daily.dev
og:description: Google has released Gemini 3.8 Live and 3.8 Live Extended Thinking, models that can reason and run tool calls in the background without pausing a live...
og:url: https://daily.dev/posts/gemini-3-8-live-thinks-in-the-background-while-you-talk-and-the-benchmarks-are-hard-to-ignore-vqety7y4w
og:image: https://api.daily.dev/og/posts/vqeTY7Y4w.png
og:image:alt: Gemini 3.8 Live thinks in the background while you talk, and the benchmarks are hard to ignore
og:image:width: 1200
og:image:height: 630
og:locale: en
---

> ## Documentation Index
> Fetch the complete documentation index at: https://daily.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# Gemini 3.8 Live thinks in the background while you talk, and the benchmarks are hard to ignore

**[Trends](https://daily.dev/sources/trends)** · 3 min read · 4 upvotes · 0 comments

## Summary

Google has released Gemini 3.8 Live and 3.8 Live Extended Thinking, models that can reason and run tool calls in the background without pausing a live conversation. The architecture moves away from cascaded ASR-LLM-TTS pipelines toward a native end-to-end multimodal model, preserving tone and emotional cues that get lost at pipeline handoffs. Gemini 3.8 scores 82.6 on the Artificial Analysis Quality Index (top spot) and 35.1 on τ-banking agentic task completion. Pricing is $0.005/min input and $0.018/min output. New Live API features include async function calling, proactive audio, and send_client_content for injecting context mid-session. The model is available in Google AI Studio, the Gemini API, Search, and the Gemini app, with partner integrations in LiveKit, Pipecat, LangChain, and Vercel. Google also highlighted a 70+ language real-time translation demo in Google Meet, noting most Gemini users are non-English speakers.

## Content

Google and OpenAI shipped competing real-time voice architectures within five days of each other, and the gap is sharper than the timing suggests.

Gemini 3.8 Live scores 82.6 on the Artificial Analysis Quality Index (first place), 68.6% on τ-Voice versus GPT-Live-1's 67.9%, and costs $0.84/hour — roughly 80% cheaper than OpenAI's comparable tier. That's the headline. The more interesting part is *why* the two systems diverge so much in price and design.

**Two very different bets on architecture**

Google kept everything in one place. Gemini 3.8 Live runs reasoning, speech, and tool execution inside a single stateful session. The Extended Thinking variant can think in the background and run async tool calls while still talking to you — no awkward pause while it waits for a database lookup to finish.

OpenAI went the opposite direction. GPT-Live-1 splits the job: a small, fast voice model handles the conversation while a separate frontier model (GPT-5.5 or similar) does the reasoning and tool calls. The engineering is genuinely clever — WARP, a custom protocol that collapses WebRTC's six-step handshake into one round trip, sessions kept resident in GPU memory to avoid reprocessing context, prefetching for the backend model before delegation is even needed. But that split architecture means developers are on the hook for orchestrating two systems, and they're paying for both: $0.05/min for voice alone, plus separate backend charges.

Gemini's $0.005/min input and $0.018/min output pricing isn't just cheaper — it's a different order of magnitude.

**What the benchmarks actually test**

τ-Voice is worth understanding here. It doesn't measure whether a model sounds natural. It measures whether it can finish a real multi-step task: hold a conversation, follow domain policies, use tools correctly, and reach the right outcome across airline, retail, and telecom scenarios. Gemini's narrow lead there (68.6% vs 67.9%) matters more than it looks, because that's the benchmark closest to what production voice agents actually need to do.

The caveat: both companies used different test suites for their own benchmarks, so direct comparisons have limits. And Anthropic is still sitting this one out — Claude has no comparable real-time speech-to-speech API yet.

For developers building voice agents today, the calculus is pretty clear: Google's integrated architecture is cheaper and slightly ahead on agentic tasks. OpenAI's split model gives you more control but hands you the orchestration problem. Neither is obviously wrong — they're just different bets on where the complexity should live.

## Questions this post answers

### What is new in Gemini 3.8 Live compared to previous voice models?

Gemini 3.8 Live can reason and run tool calls in the background without pausing an ongoing conversation, using async function calling so tools execute while the user keeps talking. It also introduces proactive audio, which only responds when addressed or when something relevant occurs, and send_client_content, which injects context into a session without forcing a conversational turn. It scores 82.6 on the Artificial Analysis Quality Index and 35.1 on τ-banking agentic task completion.

_Teams building voice agents can track new Gemini Live capabilities like these on daily.dev._

### How much does Gemini 3.8 Live cost to use via the API?

Gemini 3.8 Live is priced at $0.005 per minute of input audio and $0.018 per minute of output audio, positioned as competitive for real-time voice applications. It is available through Google AI Studio, the Gemini API, Google Search, and the Gemini app, with partner integrations in LiveKit, Pipecat, LangChain, and Vercel.

_Compare real-time voice API pricing across providers before committing on daily.dev._

### Why are native multimodal voice models better than cascaded ASR-LLM-TTS pipelines?

Cascaded pipelines transcribe audio to text, reason over the text, then synthesize speech, losing paralinguistic information like tone, hesitation, and emotion at each handoff. A natively multimodal end-to-end model like Gemini 3.8 Live processes audio directly, preserving that information throughout the interaction rather than discarding it during transcription.

_Developers evaluating voice AI architectures can follow shifts like this one on daily.dev._

## Community take

How the wider developer community reacted, aggregated from 2 discussions and 12 comments across x (as of 2026-09-22).

**TL;DR:** Reactions center on the low pricing and the async 'thinks while talking' architecture as genuinely useful for voice agents, though several people flag latency gaps, missing regional/app availability, and want clearer speech-specific benchmarks before trusting the headline numbers.

**Sentiment:** 30% positive · 55% mixed · 15% skeptical

**The case for**

- The per-minute pricing is seen as cheap enough to genuinely undercut call-center economics.
- Background thinking plus async tool calls while still speaking is viewed as solving the hard turn-taking problem in voice agents.
- Topping τ-banking agentic task completion is seen as evidence the latency bottleneck has been addressed.

**The pushback**

- A multi-second thinking pause in a live call can feel like the call dropped.
- Concern about who's accountable when the agent confidently takes a wrong action, like booking the wrong flight.
- Doubts that the headline Quality Index reflects speech-to-speech performance specifically, since Artificial Analysis benchmarks audio separately.
- Feature isn't yet available in some regions (EU) or in the consumer Gemini App for paid subscribers.

**By community**

- x (mixed): Replies mix enthusiasm about pricing and async architecture with requests for clearer audio-specific benchmarks, regional availability gaps, and concerns about latency and reliability of async tool calls.

**Hottest debate:** Whether the touted Quality Index and benchmark wins actually reflect real speech-to-speech performance, or just text-based capability.

**Open questions**

- How does 3.8 Live actually score on the dedicated Speech to Speech Index rather than the general Quality Index?
- How should turn-taking be handled when the model wants to interject mid-sentence versus waiting for a pause?
- What is the cost per correctly completed workflow, including retries and recovery, rather than just raw audio cost?
- Does a changed user request mid-tool-execution still get routed and completed correctly?

**Highlights**

> @_philschmid Artificial Analysis describes the Quality Index as a primarily text and English language suite, with speech inputs benchmarked separately from it. Their native audio measure is the Speech to Speech Index, built on Big Bench Audio and tau-Voice. Where does 3.8 Live land there?
> — [jatingargiitk on x](https://x.com/jatingargiitk/status/2099908846273863778)

> @_philschmid An agent that keeps talking while it thinks in the background. So we finally shipped the confident coworker. Jokes aside, $0.005/min input is the first voice pricing that genuinely undercuts a call center. The question is who picks up when it confidently books the wrong flight.
> — [daniel\_priscu on x](https://x.com/daniel_priscu/status/2099927014748610691)

> @_philschmid Background thinking + async tool calls while still talking. That's the hard part of voice agents. τ-banking at #1 tracks, latency was the bottleneck. Curious how I'd wire turn-taking: pre-empt mid-sentence or wait for a pause?
> — [miratechgeek on x](https://x.com/miratechgeek/status/2099934854498378068)

> @_philschmid The partner-plugin ecosystem makes a shared acceptance suite valuable. Keep the business scenario fixed across integrations, record the order of acknowledgements and tool completions, and verify identical final state after an interruption. Adapter differences should be visible
> — [AgomaMitchell on x](https://x.com/AgomaMitchell/status/2099969255764857321)

> @rohanpaul_ai The task-completion distinction is important. For an application evaluation, I would also report cost per correctly completed workflow, including retries and recovery, rather than audio cost alone. Then test whether a changed request during tool execution still reaches the right
> — [AgomaMitchell on x](https://x.com/AgomaMitchell/status/2099968098484756975)

**Source threads**

- [x](https://x.com/_philschmid/status/2099908172899357093) · 0 points · 10 comments
- [x](https://x.com/rohanpaul_ai/status/2099942230379401250) · 0 points · 2 comments

## Similar posts on daily.dev

- [Introducing Gemini 3.8 Live and 3.8 Live Extended Thinking](https://daily.dev/posts/introducing-gemini-3-8-live-and-3-8-live-extended-thinking-cicfsv9uc) · DeepMind · 0 upvotes · 0 comments
- [Gemini 3.0 vs GPT-5.1: a clear comparison for builders](https://daily.dev/posts/gemini-3-0-vs-gpt-5-1-a-clear-comparison-for-builders-j1vqnpv6e) · portkey · 1 upvotes · 0 comments

---

Tags: [#ai-agents](https://daily.dev/tags/ai-agents), [#google-gemini](https://daily.dev/tags/google-gemini), [#multimodal](https://daily.dev/tags/multimodal), [#voice-ai](https://daily.dev/tags/voice-ai), [#google-deepmind](https://daily.dev/tags/google-deepmind)

[View this post on daily.dev](https://daily.dev/posts/gemini-3-8-live-thinks-in-the-background-while-you-talk-and-the-benchmarks-are-hard-to-ignore-vqety7y4w)

```json
{"@context":"https://schema.org","@graph":[{"@type":"Organization","@id":"https://daily.dev/#organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180},"sameAs":["https://twitter.com/dailydotdev","https://github.com/dailydotdev","https://www.linkedin.com/company/daily-dev-ltd"]},{"@type":"WebSite","@id":"https://daily.dev/#website","url":"https://daily.dev","name":"daily.dev","publisher":{"@id":"https://daily.dev/#organization"},"potentialAction":{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https://daily.dev/search?q={search_term_string}"},"query-input":"required name=search_term_string"}}]}
{"@context":"https://schema.org","@type":"TechArticle","headline":"Gemini 3.8 Live thinks in the background while you talk, and the benchmarks are hard to ignore","url":"https://daily.dev/posts/gemini-3-8-live-thinks-in-the-background-while-you-talk-and-the-benchmarks-are-hard-to-ignore-vqety7y4w","mainEntityOfPage":{"@type":"WebPage","@id":"https://daily.dev/posts/gemini-3-8-live-thinks-in-the-background-while-you-talk-and-the-benchmarks-are-hard-to-ignore-vqety7y4w"},"datePublished":"2026-09-15T17:25:07.729Z","dateModified":"2026-09-22T15:35:20.152Z","description":"Google has released Gemini 3.8 Live and 3.8 Live Extended Thinking, models that can reason and run tool calls in the background without pausing a live...","image":"https://i.ytimg.com/vi/3CyW24Pkz4o/sddefault.jpg","thumbnailUrl":"https://i.ytimg.com/vi/3CyW24Pkz4o/sddefault.jpg","isAccessibleForFree":true,"articleSection":"Trends","inLanguage":"en","publisher":{"@type":"Organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180}},"author":{"@type":"Organization","name":"Trends","logo":"https://media.daily.dev/image/upload/s--ZfSp3asX--/f_auto,q_auto/v1780996004/logos/trends?_a=BAMAMiWQ0","url":"https://daily.dev/sources/trends"},"commentCount":0,"discussionUrl":"https://daily.dev/posts/gemini-3-8-live-thinks-in-the-background-while-you-talk-and-the-benchmarks-are-hard-to-ignore-vqety7y4w","interactionStatistic":[{"@type":"InteractionCounter","interactionType":{"@type":"LikeAction"},"userInteractionCount":4},{"@type":"InteractionCounter","interactionType":{"@type":"CommentAction"},"userInteractionCount":0}],"keywords":"ai-agents,google-gemini,multimodal,voice-ai,google-deepmind","timeRequired":"PT3M"}
{"@context":"https://schema.org","@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https://daily.dev"},{"@type":"ListItem","position":2,"name":"Trends","item":"https://daily.dev/sources/trends"},{"@type":"ListItem","position":3,"name":"Gemini 3.8 Live thinks in the background while you talk, and the benchmarks are hard to ignore"}]}
{"@context":"https://schema.org","@type":"FAQPage","@id":"https://daily.dev/posts/gemini-3-8-live-thinks-in-the-background-while-you-talk-and-the-benchmarks-are-hard-to-ignore-vqety7y4w#faq","mainEntity":[{"@type":"Question","name":"What is new in Gemini 3.8 Live compared to previous voice models?","acceptedAnswer":{"@type":"Answer","text":"Gemini 3.8 Live can reason and run tool calls in the background without pausing an ongoing conversation, using async function calling so tools execute while the user keeps talking. It also introduces proactive audio, which only responds when addressed or when something relevant occurs, and send_client_content, which injects context into a session without forcing a conversational turn. It scores 82.6 on the Artificial Analysis Quality Index and 35.1 on τ-banking agentic task completion. Teams building voice agents can track new Gemini Live capabilities like these on daily.dev."}},{"@type":"Question","name":"How much does Gemini 3.8 Live cost to use via the API?","acceptedAnswer":{"@type":"Answer","text":"Gemini 3.8 Live is priced at $0.005 per minute of input audio and $0.018 per minute of output audio, positioned as competitive for real-time voice applications. It is available through Google AI Studio, the Gemini API, Google Search, and the Gemini app, with partner integrations in LiveKit, Pipecat, LangChain, and Vercel. Compare real-time voice API pricing across providers before committing on daily.dev."}},{"@type":"Question","name":"Why are native multimodal voice models better than cascaded ASR-LLM-TTS pipelines?","acceptedAnswer":{"@type":"Answer","text":"Cascaded pipelines transcribe audio to text, reason over the text, then synthesize speech, losing paralinguistic information like tone, hesitation, and emotion at each handoff. A natively multimodal end-to-end model like Gemini 3.8 Live processes audio directly, preserving that information throughout the interaction rather than discarding it during transcription. Developers evaluating voice AI architectures can follow shifts like this one on daily.dev."}}]}
```

