Modal
Read post

One-Second Voice-to-Voice Latency with Modal, Pipecat, and Open Models

Building a conversational voice AI bot with sub-second response latency using Modal's serverless platform, the Pipecat framework, and open-source models. The implementation achieves ~1 second voice-to-voice latency by combining Parakeet STT, Qwen3 LLM with vLLM, and Kokoro TTS. Key optimizations include using Modal Tunnels to bypass the input plane, WebRTC for client connections, regional pinning to minimize network latency, and independent autoscaling of GPU services. The demo includes a RAG-powered assistant for Modal's documentation with structured outputs and animated avatars.

    #python#real-time-systems#voice-ai#webrtc
Nov 04, 2025•14m read time•From modal.com
Post cover image
Table of contents
Conversational Voice AI ApplicationsWhy Modal and Pipecat work so well together for Voice AIVoice-to-Voice LatencyBuilding a Conversational Voice AI for Modal’s DocsTesting PerformanceDeploy Your Own Conversational Voice AI on ModalBonus: Animating Modal’s Mascots Moe and Dal
569 Impressions
Modal's image
Modal

35 Followers

•

346 Upvotes

Would you recommend this post?

Copy link
WhatsApp
Facebook
X
New Squad
  • © 2026 Daily Dev Ltd.
  • Guidelines
  • Explore
  • Tags
  • Sources
  • Squads
  • Leaderboard