<!-- mobian-agent-page publisher="dailydev" canonical="https://daily.dev/posts/teaching-a-child-in-1000-ms-the-architecture-behind-a-real-time-tutor-8t3zegwa2" -->

---
title: Teaching a child in &lt;1000 ms: the architecture behind a...
description: Ello shares the architectural decisions behind building a real-time AI tutor for children ages 4-9, where sub-second response time is non-negotiable. Key...
canonical: https://daily.dev/posts/teaching-a-child-in-1000-ms-the-architecture-behind-a-real-time-tutor-8t3zegwa2
twitter:card: summary_large_image
twitter:site: @dailydotdev
og:type: website
og:site_name: daily.dev
og:title: Teaching a child in &lt;1000 ms: the architecture behind a real-time tutor | daily.dev
og:description: Ello shares the architectural decisions behind building a real-time AI tutor for children ages 4-9, where sub-second response time is non-negotiable. Key...
og:url: https://daily.dev/posts/teaching-a-child-in-1000-ms-the-architecture-behind-a-real-time-tutor-8t3zegwa2
og:image: https://api.daily.dev/og/posts/8T3ZEGWA2.png
og:image:alt: Teaching a child in &lt;1000 ms: the architecture behind a real-time tutor
og:image:width: 1200
og:image:height: 630
og:locale: en
---

> ## Documentation Index
> Fetch the complete documentation index at: https://daily.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# Teaching a child in <1000 ms: the architecture behind a real-time tutor

**[Hacker News](https://daily.dev/sources/hn)** · 10 min read · 1 upvotes · 0 comments

## Summary

Ello shares the architectural decisions behind building a real-time AI tutor for children ages 4-9, where sub-second response time is non-negotiable. Key innovations include: replacing the standard LLM tool loop with a custom streaming harness that parses and executes actions while the model is still generating; an asynchronous planner agent that reasons about pedagogy in the gaps while a converser handles real-time interaction; pre-generating responses to predicted child answers on branched trajectories; and running a safety classifier in parallel with an eager response model so safety checks never add latency. The post explains the tradeoffs of each approach, including cost, observability overhead, and occasional mispredictions.

## Full article

daily.dev links to this article rather than hosting it. Read it at the original source: <https://www.ello.com/blog/teaching-a-child-in-1000-ms>

## Questions this post answers

### Why does the standard LLM tool-call agent loop cause too much latency for a real-time conversational agent?

The standard tool loop waits for the model to finish generating before executing actions, and frontier models take 2-3 seconds to produce a first token and then decode at roughly 30 tokens per second. With actions averaging a few dozen tokens, plus round-trip latency and audio playback, this produces 3-4 seconds of downtime between each spoken sentence or screen change, which is too slow for holding a child's attention.

_Anyone designing responsive AI agents can compare real-time architecture tradeoffs like these on daily.dev._

### How can an AI system run safety checks on user input without adding latency to every response?

Execution can be gated on the safety check while generation runs in parallel with it. As soon as user input finishes, both a safety classifier (taking roughly 500-1000ms) and a small model generating a quick, low-risk acknowledgment response are dispatched simultaneously; the acknowledgment only executes once the classifier confirms the turn is safe, avoiding a sequential delay.

_Teams building safety-gated real-time agents can track patterns like this on daily.dev._

### How can an AI agent predict a user's response before they finish answering, to reduce reply latency?

When a closed-ended question is asked (like a fill-in-the-blank or a math equation), the system hypothesizes the likely answers in advance and pre-generates a response for each one on a separate branch forked from the conversation trajectory. Once the actual answer arrives, it is matched to the corresponding branch and the pre-generated response plays immediately without a fresh model call.

_Developers exploring predictive response generation for low-latency agents can follow techniques like this on daily.dev._

## Similar posts on daily.dev

- [The 200ms latency: A developer’s guide to real-time personalization](https://daily.dev/posts/the-200ms-latency-a-developer-s-guide-to-real-time-personalization-u7s5bwkwz) · InfoWorld · 0 upvotes · 0 comments
- [Is Your “Human-in-the-Loop” Actually Slowing You Down? Here’s What We Learned](https://daily.dev/posts/is-your-human-in-the-loop-actually-slowing-you-down-here-s-what-we-learned-s37uyphem) · Stack Overflow Blog · 0 upvotes · 0 comments
- [Inside Thinking Machines’ Interaction Models](https://daily.dev/posts/inside-thinking-machines-interaction-models-z2ltfxsdw) · ByteByteGo · 2 upvotes · 0 comments
- [4 AI Use Cases Exposing Your EdTech Platform's Data Gap](https://daily.dev/posts/4-ai-use-cases-exposing-your-edtech-platform-s-data-gap-6gqv3ltmq) · SingleStore · 1 upvotes · 0 comments

---

Tags: [#llm](https://daily.dev/tags/llm), [#ai-agents](https://daily.dev/tags/ai-agents), [#edtech](https://daily.dev/tags/edtech), [#real-time-systems](https://daily.dev/tags/real-time-systems)

[View this post on daily.dev](https://daily.dev/posts/teaching-a-child-in-1000-ms-the-architecture-behind-a-real-time-tutor-8t3zegwa2)

```json
{"@context":"https://schema.org","@graph":[{"@type":"Organization","@id":"https://daily.dev/#organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180},"sameAs":["https://twitter.com/dailydotdev","https://github.com/dailydotdev","https://www.linkedin.com/company/daily-dev-ltd"]},{"@type":"WebSite","@id":"https://daily.dev/#website","url":"https://daily.dev","name":"daily.dev","publisher":{"@id":"https://daily.dev/#organization"},"potentialAction":{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https://daily.dev/search?q={search_term_string}"},"query-input":"required name=search_term_string"}}]}
{"@context":"https://schema.org","@type":"TechArticle","headline":"Teaching a child in <1000 ms: the architecture behind a real-time tutor","url":"https://daily.dev/posts/teaching-a-child-in-1000-ms-the-architecture-behind-a-real-time-tutor-8t3zegwa2","mainEntityOfPage":{"@type":"WebPage","@id":"https://daily.dev/posts/teaching-a-child-in-1000-ms-the-architecture-behind-a-real-time-tutor-8t3zegwa2"},"datePublished":"2026-07-10T01:16:36.305Z","dateModified":"2026-09-13T18:26:08.800Z","description":"Ello shares the architectural decisions behind building a real-time AI tutor for children ages 4-9, where sub-second response time is non-negotiable. Key...","image":"https://media.daily.dev/image/upload/f_auto,q_auto/v1/posts/22cd4321263dba5538d0481b85b9f458?_a=AQAEuop","thumbnailUrl":"https://media.daily.dev/image/upload/f_auto,q_auto/v1/posts/22cd4321263dba5538d0481b85b9f458?_a=AQAEuop","isAccessibleForFree":true,"articleSection":"Hacker News","inLanguage":"en","publisher":{"@type":"Organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180}},"author":{"@type":"Organization","name":"Hacker News","logo":"https://media.daily.dev/image/upload/t_logo,f_auto/v1/logos/hn","url":"https://daily.dev/sources/hn"},"commentCount":0,"discussionUrl":"https://daily.dev/posts/teaching-a-child-in-1000-ms-the-architecture-behind-a-real-time-tutor-8t3zegwa2","interactionStatistic":[{"@type":"InteractionCounter","interactionType":{"@type":"LikeAction"},"userInteractionCount":1},{"@type":"InteractionCounter","interactionType":{"@type":"CommentAction"},"userInteractionCount":0}],"keywords":"llm,ai-agents,edtech,real-time-systems","timeRequired":"PT10M"}
{"@context":"https://schema.org","@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https://daily.dev"},{"@type":"ListItem","position":2,"name":"Hacker News","item":"https://daily.dev/sources/hn"},{"@type":"ListItem","position":3,"name":"Teaching a child in <1000 ms: the architecture behind a real-time tutor"}]}
{"@context":"https://schema.org","@type":"FAQPage","@id":"https://daily.dev/posts/teaching-a-child-in-1000-ms-the-architecture-behind-a-real-time-tutor-8t3zegwa2#faq","mainEntity":[{"@type":"Question","name":"Why does the standard LLM tool-call agent loop cause too much latency for a real-time conversational agent?","acceptedAnswer":{"@type":"Answer","text":"The standard tool loop waits for the model to finish generating before executing actions, and frontier models take 2-3 seconds to produce a first token and then decode at roughly 30 tokens per second. With actions averaging a few dozen tokens, plus round-trip latency and audio playback, this produces 3-4 seconds of downtime between each spoken sentence or screen change, which is too slow for holding a child's attention. Anyone designing responsive AI agents can compare real-time architecture tradeoffs like these on daily.dev."}},{"@type":"Question","name":"How can an AI system run safety checks on user input without adding latency to every response?","acceptedAnswer":{"@type":"Answer","text":"Execution can be gated on the safety check while generation runs in parallel with it. As soon as user input finishes, both a safety classifier (taking roughly 500-1000ms) and a small model generating a quick, low-risk acknowledgment response are dispatched simultaneously; the acknowledgment only executes once the classifier confirms the turn is safe, avoiding a sequential delay. Teams building safety-gated real-time agents can track patterns like this on daily.dev."}},{"@type":"Question","name":"How can an AI agent predict a user's response before they finish answering, to reduce reply latency?","acceptedAnswer":{"@type":"Answer","text":"When a closed-ended question is asked (like a fill-in-the-blank or a math equation), the system hypothesizes the likely answers in advance and pre-generates a response for each one on a separate branch forked from the conversation trajectory. Once the actual answer arrives, it is matched to the corresponding branch and the pre-generated response plays immediately without a fresh model call. Developers exploring predictive response generation for low-latency agents can follow techniques like this on daily.dev."}}]}
```

