<!-- mobian-agent-page publisher="dailydev" canonical="https://daily.dev/posts/one-second-voice-to-voice-latency-with-modal-pipecat-and-open-models-aejtkcqop" -->

---
title: One-Second Voice-to-Voice Latency with Modal, Pipecat,...
description: Building a conversational voice AI bot with sub-second response latency using Modal&#x27;s serverless platform, the Pipecat framework, and open-source models. The...
canonical: https://daily.dev/posts/one-second-voice-to-voice-latency-with-modal-pipecat-and-open-models-aejtkcqop
twitter:card: summary_large_image
twitter:site: @dailydotdev
og:type: website
og:site_name: daily.dev
og:title: One-Second Voice-to-Voice Latency with Modal, Pipecat, and Open Models | daily.dev
og:description: Building a conversational voice AI bot with sub-second response latency using Modal&#x27;s serverless platform, the Pipecat framework, and open-source models. The...
og:url: https://daily.dev/posts/one-second-voice-to-voice-latency-with-modal-pipecat-and-open-models-aejtkcqop
og:image: https://api.daily.dev/og/posts/AeJTKCqop.png
og:image:alt: One-Second Voice-to-Voice Latency with Modal, Pipecat, and Open Models
og:image:width: 1200
og:image:height: 630
og:locale: en
---

> ## Documentation Index
> Fetch the complete documentation index at: https://daily.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# One-Second Voice-to-Voice Latency with Modal, Pipecat, and Open Models

**[Modal](https://daily.dev/sources/modal_labs)** · 14 min read · 0 upvotes · 0 comments

## Summary

Building a conversational voice AI bot with sub-second response latency using Modal's serverless platform, the Pipecat framework, and open-source models. The implementation achieves ~1 second voice-to-voice latency by combining Parakeet STT, Qwen3 LLM with vLLM, and Kokoro TTS. Key optimizations include using Modal Tunnels to bypass the input plane, WebRTC for client connections, regional pinning to minimize network latency, and independent autoscaling of GPU services. The demo includes a RAG-powered assistant for Modal's documentation with structured outputs and animated avatars.

## Full article

daily.dev links to this article rather than hosting it. Read it at the original source: <https://modal.com/blog/low-latency-voice-bot>

## Similar posts on daily.dev

- [Real-time voice AI low-latency techniques that work](https://daily.dev/posts/real-time-voice-ai-low-latency-techniques-that-work-noffnjl2b) · Netguru · 0 upvotes · 0 comments
- [How I built a sub-500ms latency voice agent from scratch](https://daily.dev/posts/how-i-built-a-sub-500ms-latency-voice-agent-from-scratch-neao8lbjr) · Hacker News · 0 upvotes · 0 comments
- [How I Made My Own Offline Voice Assistant](https://daily.dev/posts/how-i-made-my-own-offline-voice-assistant-lvephkfkw) · Medium · 0 upvotes · 0 comments
- [How Decagon shipped real-time voice AI on Modal](https://daily.dev/posts/how-decagon-shipped-real-time-voice-ai-on-modal-y0nigfyhy) · Modal · 1 upvotes · 0 comments

---

Tags: [#python](https://daily.dev/tags/python), [#real-time-systems](https://daily.dev/tags/real-time-systems), [#voice-ai](https://daily.dev/tags/voice-ai), [#webrtc](https://daily.dev/tags/webrtc)

[View this post on daily.dev](https://daily.dev/posts/one-second-voice-to-voice-latency-with-modal-pipecat-and-open-models-aejtkcqop)

```json
{"@context":"https://schema.org","@graph":[{"@type":"Organization","@id":"https://daily.dev/#organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180},"sameAs":["https://twitter.com/dailydotdev","https://github.com/dailydotdev","https://www.linkedin.com/company/daily-dev-ltd"]},{"@type":"WebSite","@id":"https://daily.dev/#website","url":"https://daily.dev","name":"daily.dev","publisher":{"@id":"https://daily.dev/#organization"},"potentialAction":{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https://daily.dev/search?q={search_term_string}"},"query-input":"required name=search_term_string"}}]}
{"@context":"https://schema.org","@type":"TechArticle","headline":"One-Second Voice-to-Voice Latency with Modal, Pipecat, and Open Models","url":"https://daily.dev/posts/one-second-voice-to-voice-latency-with-modal-pipecat-and-open-models-aejtkcqop","mainEntityOfPage":{"@type":"WebPage","@id":"https://daily.dev/posts/one-second-voice-to-voice-latency-with-modal-pipecat-and-open-models-aejtkcqop"},"datePublished":"2025-11-04T22:13:07.962Z","dateModified":"2026-03-30T02:12:48.400Z","description":"Building a conversational voice AI bot with sub-second response latency using Modal's serverless platform, the Pipecat framework, and open-source models. The...","image":"https://media.daily.dev/image/upload/f_auto,q_auto/v1/posts/541320873264f1ac7dc894ba522a25ce?_a=AQAEulh","thumbnailUrl":"https://media.daily.dev/image/upload/f_auto,q_auto/v1/posts/541320873264f1ac7dc894ba522a25ce?_a=AQAEulh","isAccessibleForFree":true,"articleSection":"Modal","inLanguage":"en","publisher":{"@type":"Organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180}},"author":{"@type":"Organization","name":"Modal","logo":"https://media.daily.dev/image/upload/s--HfKZHzC1--/f_auto/v1750946306/logos/modal_labs","url":"https://daily.dev/sources/modal_labs"},"commentCount":0,"discussionUrl":"https://daily.dev/posts/one-second-voice-to-voice-latency-with-modal-pipecat-and-open-models-aejtkcqop","interactionStatistic":[{"@type":"InteractionCounter","interactionType":{"@type":"LikeAction"},"userInteractionCount":0},{"@type":"InteractionCounter","interactionType":{"@type":"CommentAction"},"userInteractionCount":0}],"keywords":"python,real-time-systems,voice-ai,webrtc","timeRequired":"PT14M"}
{"@context":"https://schema.org","@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https://daily.dev"},{"@type":"ListItem","position":2,"name":"Modal","item":"https://daily.dev/sources/modal_labs"},{"@type":"ListItem","position":3,"name":"One-Second Voice-to-Voice Latency with Modal, Pipecat, and Open Models"}]}
```

