<!-- mobian-agent-page publisher="dailydev" canonical="https://daily.dev/posts/how-real-time-ai-avatars-actually-work-every-piece-explained-built-w-synthesia--2pfmmqrxr" -->

---
title: How Real-Time AI Avatars Actually Work (Every Piece...
description: A real-time interactive AI avatar system is demonstrated on a restaurant website where a talking face responds to voice, fills orders, and manipulates the UI...
canonical: https://daily.dev/posts/how-real-time-ai-avatars-actually-work-every-piece-explained-built-w-synthesia--2pfmmqrxr
twitter:card: summary_large_image
twitter:site: @dailydotdev
og:type: website
og:site_name: daily.dev
og:title: How Real-Time AI Avatars Actually Work (Every Piece Explained &amp; Built w/ Synthesia) | daily.dev
og:description: A real-time interactive AI avatar system is demonstrated on a restaurant website where a talking face responds to voice, fills orders, and manipulates the UI...
og:url: https://daily.dev/posts/how-real-time-ai-avatars-actually-work-every-piece-explained-built-w-synthesia--2pfmmqrxr
og:image: https://api.daily.dev/og/posts/2pFmMqrxR.png
og:image:alt: How Real-Time AI Avatars Actually Work (Every Piece Explained &amp; Built w/ Synthesia)
og:image:width: 1200
og:image:height: 630
og:locale: en
---

> ## Documentation Index
> Fetch the complete documentation index at: https://daily.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# How Real-Time AI Avatars Actually Work (Every Piece Explained & Built w/ Synthesia)

**[Tech With Tim](https://daily.dev/sources/TechWithTim)** · 20 min read · 0 upvotes · 0 comments

## Summary

A real-time interactive AI avatar system is demonstrated on a restaurant website where a talking face responds to voice, fills orders, and manipulates the UI live. The pipeline is broken down: WebRTC for connection, voice activity detection, speech-to-text, turn detection, an LLM (GPT/Claude/Gemini) as the brain, tool calling to manipulate the page, text-to-speech streaming, and finally Synthesia's interactive avatar API for the live-rendered, lip-synced face. LiveKit handles the real-time room, audio/video routing, and agent orchestration, while a small Python token server issues secure join credentials. A working code sample (server.py, agent.py, index.html) is walked through, showing how to wire LiveKit and Synthesia's plugin together, including using a Synthesia skill for coding agents like Claude Code or Cursor to scaffold the integration automatically. Synthesia's API is pay-as-you-go at 12 cents per minute with free credits for testing.

## Full article

daily.dev links to this article rather than hosting it. Read it at the original source: <https://www.youtube.com/watch?v=OtpdA8Ye2OY>

## Questions this post answers

### How much does Synthesia's interactive avatar API cost to use?

It is pay-as-you-go pricing at 12 cents per minute of avatar usage, with a batch of free credits provided to test the API before paying anything. Developers access it through an API designed as a plugin for LiveKit, with support for additional frameworks planned, and there is even a setup skill for coding agents like Claude Code and Cursor.

_Developers pricing out real-time avatar features for a product can track tools like Synthesia's API on daily.dev._

### What components do I need to build a real-time talking AI avatar for a website?

You need three pieces you write yourself: a web page, a small token server (for secure credential issuance), and an AI agent, plus two external services: LiveKit for real-time audio/video room routing and Synthesia for rendering the live, lip-synced avatar face. The agent runs speech-to-text, an LLM, tool calling, and text-to-speech, then hands its audio to Synthesia via a LiveKit plugin.

_Anyone assembling a voice-driven agent stack can follow architecture breakdowns like this one on daily.dev._

### Why do real-time AI avatars need to respond so much faster than a text chatbot?

Because a normal human conversation has only about a 200 millisecond gap between turns, so a face that stares silently for even a few seconds feels broken, whereas users tolerate a chatbot taking several seconds to reply in a text window. This latency requirement is why systems stream text-to-speech token by token and overlap pipeline stages rather than waiting for each step to fully finish.

_Teams optimizing latency in voice interfaces can find similar real-time architecture breakdowns on daily.dev._

## Similar posts on daily.dev

- [I created an interactive digital avatar of myself — and you can talk to it](https://daily.dev/posts/i-created-an-interactive-digital-avatar-of-myself-and-you-can-talk-to-it-trznf7yu1) · TechCrunch · 0 upvotes · 0 comments
- [Build a real-time hotel booking voice agent](https://daily.dev/posts/build-a-real-time-hotel-booking-voice-agent-ntorwqf8n) · Daily Dose of Data Science \| Avi Chawla \| Substack · 1 upvotes · 0 comments

---

Tags: [#text-to-speech](https://daily.dev/tags/text-to-speech), [#webrtc](https://daily.dev/tags/webrtc)

[View this post on daily.dev](https://daily.dev/posts/how-real-time-ai-avatars-actually-work-every-piece-explained-built-w-synthesia--2pfmmqrxr)

```json
{"@context":"https://schema.org","@graph":[{"@type":"Organization","@id":"https://daily.dev/#organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180},"sameAs":["https://twitter.com/dailydotdev","https://github.com/dailydotdev","https://www.linkedin.com/company/daily-dev-ltd"]},{"@type":"WebSite","@id":"https://daily.dev/#website","url":"https://daily.dev","name":"daily.dev","publisher":{"@id":"https://daily.dev/#organization"},"potentialAction":{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https://daily.dev/search?q={search_term_string}"},"query-input":"required name=search_term_string"}}]}
{"@context":"https://schema.org","@type":"TechArticle","headline":"How Real-Time AI Avatars Actually Work (Every Piece Explained & Built w/ Synthesia)","url":"https://daily.dev/posts/how-real-time-ai-avatars-actually-work-every-piece-explained-built-w-synthesia--2pfmmqrxr","mainEntityOfPage":{"@type":"WebPage","@id":"https://daily.dev/posts/how-real-time-ai-avatars-actually-work-every-piece-explained-built-w-synthesia--2pfmmqrxr"},"datePublished":"2026-10-02T14:55:42.339Z","dateModified":"2026-10-02T14:56:06.179Z","description":"A real-time interactive AI avatar system is demonstrated on a restaurant website where a talking face responds to voice, fills orders, and manipulates the UI...","image":"https://i.ytimg.com/vi/OtpdA8Ye2OY/sddefault.jpg","thumbnailUrl":"https://i.ytimg.com/vi/OtpdA8Ye2OY/sddefault.jpg","isAccessibleForFree":true,"articleSection":"Tech With Tim","inLanguage":"en","publisher":{"@type":"Organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180}},"author":{"@type":"Organization","name":"Tech With Tim","logo":"https://media.daily.dev/image/upload/s--6vVIDfU0--/f_auto/v1711726146/logos/TechWithTim","url":"https://daily.dev/sources/TechWithTim"},"commentCount":0,"discussionUrl":"https://daily.dev/posts/how-real-time-ai-avatars-actually-work-every-piece-explained-built-w-synthesia--2pfmmqrxr","interactionStatistic":[{"@type":"InteractionCounter","interactionType":{"@type":"LikeAction"},"userInteractionCount":0},{"@type":"InteractionCounter","interactionType":{"@type":"CommentAction"},"userInteractionCount":0}],"keywords":"text-to-speech,webrtc","timeRequired":"PT20M","video":{"@type":"VideoObject","name":"How Real-Time AI Avatars Actually Work (Every Piece Explained & Built w/ Synthesia)","description":"A real-time interactive AI avatar system is demonstrated on a restaurant website where a talking face responds to voice, fills orders, and manipulates the UI...","thumbnailUrl":"https://i.ytimg.com/vi/OtpdA8Ye2OY/sddefault.jpg","uploadDate":"2026-10-02T14:55:42.339Z","duration":"PT20M","url":"https://api.daily.dev/r/2pFmMqrxR","embedUrl":"https://www.youtube.com/embed/OtpdA8Ye2OY"}}
{"@context":"https://schema.org","@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https://daily.dev"},{"@type":"ListItem","position":2,"name":"Tech With Tim","item":"https://daily.dev/sources/TechWithTim"},{"@type":"ListItem","position":3,"name":"How Real-Time AI Avatars Actually Work (Every Piece Explained & Built w/ Synthesia)"}]}
{"@context":"https://schema.org","@type":"FAQPage","@id":"https://daily.dev/posts/how-real-time-ai-avatars-actually-work-every-piece-explained-built-w-synthesia--2pfmmqrxr#faq","mainEntity":[{"@type":"Question","name":"How much does Synthesia's interactive avatar API cost to use?","acceptedAnswer":{"@type":"Answer","text":"It is pay-as-you-go pricing at 12 cents per minute of avatar usage, with a batch of free credits provided to test the API before paying anything. Developers access it through an API designed as a plugin for LiveKit, with support for additional frameworks planned, and there is even a setup skill for coding agents like Claude Code and Cursor. Developers pricing out real-time avatar features for a product can track tools like Synthesia's API on daily.dev."}},{"@type":"Question","name":"What components do I need to build a real-time talking AI avatar for a website?","acceptedAnswer":{"@type":"Answer","text":"You need three pieces you write yourself: a web page, a small token server (for secure credential issuance), and an AI agent, plus two external services: LiveKit for real-time audio/video room routing and Synthesia for rendering the live, lip-synced avatar face. The agent runs speech-to-text, an LLM, tool calling, and text-to-speech, then hands its audio to Synthesia via a LiveKit plugin. Anyone assembling a voice-driven agent stack can follow architecture breakdowns like this one on daily.dev."}},{"@type":"Question","name":"Why do real-time AI avatars need to respond so much faster than a text chatbot?","acceptedAnswer":{"@type":"Answer","text":"Because a normal human conversation has only about a 200 millisecond gap between turns, so a face that stares silently for even a few seconds feels broken, whereas users tolerate a chatbot taking several seconds to reply in a text window. This latency requirement is why systems stream text-to-speech token by token and overlap pipeline stages rather than waiting for each step to fully finish. Teams optimizing latency in voice interfaces can find similar real-time architecture breakdowns on daily.dev."}}]}
```

