<!-- mobian-agent-page publisher="dailydev" canonical="https://daily.dev/posts/sspvcvzxg" -->

---
title: Streaming LLM Responses in Next.js: 1.3s to First Token,...
description: A practical walkthrough of implementing token streaming for LLM responses in a Next.js App Router application, backed by measurements against DigitalOcean&#x27;s...
canonical: https://daily.dev/posts/sspvcvzxg
twitter:card: summary_large_image
twitter:site: @dailydotdev
og:type: website
og:site_name: daily.dev
og:title: Streaming LLM Responses in Next.js: 1.3s to First Token, Not 15.7s | daily.dev
og:description: A practical walkthrough of implementing token streaming for LLM responses in a Next.js App Router application, backed by measurements against DigitalOcean&#x27;s...
og:url: https://daily.dev/posts/sspvcvzxg
og:image: https://api.daily.dev/og/posts/sSpVcvZxg.png
og:image:alt: Post cover image
og:image:width: 1200
og:image:height: 630
og:locale: en
---

> ## Documentation Index
> Fetch the complete documentation index at: https://daily.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# Streaming LLM Responses in Next.js: 1.3s to First Token, Not 15.7s

**[DevOps Daily](https://daily.dev/sources/devopsdaily)** · [@bobbyiliev](https://daily.dev/bobbyiliev) · 5 upvotes · 0 comments

## Summary

A practical walkthrough of implementing token streaming for LLM responses in a Next.js App Router application, backed by measurements against DigitalOcean's Inference Engine. Streaming drops time-to-first-token from 15.7s to 1.3s but does not reduce total generation time. The post details a common pitfall (awaiting the full upstream response instead of piping the ReadableStream), the ~120ms overhead a route handler proxy adds, how to correctly parse SSE frames that split across network reads, how to wire cancellation so aborted requests actually stop billing, why the Node runtime should be used instead of Edge for long generations, and two DigitalOcean-specific gotchas: the models list includes models that 403 at request time, and reasoning models like qwen3-32b have long first-token latency regardless of streaming.

## Full article

daily.dev links to this article rather than hosting it. Read it at the original source: <https://devops-daily.com/posts/nextjs-streaming-digitalocean-inference>

## Similar posts on daily.dev

- [Next.js for the Next Billion Users: Optimizing for High-Latency Markets](https://daily.dev/posts/next-js-for-the-next-billion-users-optimizing-for-high-latency-markets-nbxhqfay2) · SitePoint · 1 upvotes · 0 comments
- [Streaming SSR in Next.js: A Deep, Beginner-Friendly Guide for Production Apps](https://daily.dev/posts/streaming-ssr-in-next-js-a-deep-beginner-friendly-guide-for-production-apps-5qirbgtda) · Medium · 3 upvotes · 0 comments
- [How To Reduce LLM Token Costs by 70–90%](https://daily.dev/posts/how-to-reduce-llm-token-costs-by-70-90--lqaosovr2) · C\# Corner · 1 upvotes · 0 comments
- [We Ralph Wiggumed WebStreams to make them 10x faster](https://daily.dev/posts/we-ralph-wiggumed-webstreams-to-make-them-10x-faster-otxpvgii5) · Vercel · 39 upvotes · 1 comments
- [The systems guide to production token optimization](https://daily.dev/posts/the-systems-guide-to-production-token-optimization-y3mhhtlgk) · The New Stack · 0 upvotes · 0 comments

---

Tags: [#llm](https://daily.dev/tags/llm), [#typescript](https://daily.dev/tags/typescript), [#nextjs](https://daily.dev/tags/nextjs), [#digitalocean](https://daily.dev/tags/digitalocean)

[View this post on daily.dev](https://daily.dev/posts/sspvcvzxg)

```json
{"@context":"https://schema.org","@graph":[{"@type":"Organization","@id":"https://daily.dev/#organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180},"sameAs":["https://twitter.com/dailydotdev","https://github.com/dailydotdev","https://www.linkedin.com/company/daily-dev-ltd"]},{"@type":"WebSite","@id":"https://daily.dev/#website","url":"https://daily.dev","name":"daily.dev","publisher":{"@id":"https://daily.dev/#organization"},"potentialAction":{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https://daily.dev/search?q={search_term_string}"},"query-input":"required name=search_term_string"}}]}
{"@context":"https://schema.org","@type":"DiscussionForumPosting","mainEntityOfPage":"https://daily.dev/posts/sspvcvzxg","headline":"Streaming LLM Responses in Next.js: 1.3s to First Token, Not 15.7s","text":"Shared: Streaming LLM Responses in Next.js: 1.3s to First Token, Not 15.7s","url":"https://daily.dev/posts/sspvcvzxg","datePublished":"2026-08-18T04:59:34.965Z","dateModified":"2026-08-18T05:00:47.175Z","author":{"@type":"Person","name":"Bobby Iliev","url":"https://daily.dev/bobbyiliev","image":"https://avatars3.githubusercontent.com/u/21223421?v=4","description":"DevOps, DevEx, cloud & open source\n","worksFor":{"@type":"Organization","name":"Materialize","logo":"https://res.cloudinary.com/daily-now/image/upload/s--4mL2CrlK--/f_auto/v1725263959/companies/materialize"},"interactionStatistic":{"@type":"InteractionCounter","interactionType":{"@type":"EndorseAction"},"userInteractionCount":81300}},"image":"https://media.daily.dev/image/upload/s--2-1xRawN--/f_auto/v1722860399/public/Placeholder%2011","interactionStatistic":[{"@type":"InteractionCounter","interactionType":{"@type":"LikeAction"},"userInteractionCount":5},{"@type":"InteractionCounter","interactionType":{"@type":"CommentAction"},"userInteractionCount":0}],"sharedContent":{"@type":"WebPage","url":"https://api.daily.dev/r/oU68raXJT"},"isPartOf":{"@type":"WebPage","url":"https://daily.dev/squads/devopsdaily","name":"DevOps Daily"}}
{"@context":"https://schema.org","@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https://daily.dev"},{"@type":"ListItem","position":2,"name":"DevOps Daily","item":"https://daily.dev/squads/devopsdaily"},{"@type":"ListItem","position":3,"name":"Streaming LLM Responses in Next.js: 1.3s to First Token, Not 15.7s"}]}
```

