<!-- mobian-agent-page publisher="dailydev" canonical="https://daily.dev/posts/the-end-of-tcp-for-ai-clusters-john-ousterhout-stanford-swtimhc2i" -->

---
title: The End of TCP for AI Clusters — John Ousterhout, Stanford
description: John Ousterhout of Stanford argues AI workloads are shifting from massive, throughput-bound transfers to smaller, latency-sensitive exchanges, especially for...
canonical: https://daily.dev/posts/the-end-of-tcp-for-ai-clusters-john-ousterhout-stanford-swtimhc2i
twitter:card: summary_large_image
twitter:site: @dailydotdev
og:type: website
og:site_name: daily.dev
og:title: The End of TCP for AI Clusters — John Ousterhout, Stanford | daily.dev
og:description: John Ousterhout of Stanford argues AI workloads are shifting from massive, throughput-bound transfers to smaller, latency-sensitive exchanges, especially for...
og:url: https://daily.dev/posts/the-end-of-tcp-for-ai-clusters-john-ousterhout-stanford-swtimhc2i
og:image: https://api.daily.dev/og/posts/SWTiMHC2i.png
og:image:alt: The End of TCP for AI Clusters — John Ousterhout, Stanford
og:image:width: 1200
og:image:height: 630
og:locale: en
---

> ## Documentation Index
> Fetch the complete documentation index at: https://daily.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# The End of TCP for AI Clusters — John Ousterhout, Stanford

**[AI Engineer](https://daily.dev/sources/aidotengineer)** · 18 min read · 0 upvotes · 0 comments

## Summary

John Ousterhout of Stanford argues AI workloads are shifting from massive, throughput-bound transfers to smaller, latency-sensitive exchanges, especially for inference and agentic AI, where small metadata and synchronization messages between GPU nodes can stall compute if tail latency is high. He explains why TCP and RDMA (RoCE) struggle here: sender-side congestion control reacts slowly and unstably, and their byte-stream data model causes head-of-line blocking for short messages. He then introduces Homa, a clean-slate, message-based transport protocol developed at Stanford that uses receiver-driven congestion control, shortest-remaining-processing-time scheduling, and switch priority queues. Benchmarks show Homa cuts P99 tail latency for short messages by roughly 13x versus TCP and still roughly halves latency for the longest messages. Homa is implemented as a Linux kernel module, available on GitHub, and Ousterhout is pursuing kernel upstreaming while working on it full time.

## Full article

daily.dev links to this article rather than hosting it. Read it at the original source: <https://www.youtube.com/watch?v=eZ8WWZzoaR0>

## Questions this post answers

### Why does TCP perform poorly for small message latency in AI inference workloads?

TCP relies on sender-side congestion control, which reacts slowly because it takes several round trips for a sender to learn about congestion and adjust its rate, causing constant oscillation. TCP also serializes messages into an undifferentiated byte stream, so short messages can get stuck in queues behind large ones (head-of-line blocking), producing over a millisecond of tail latency in one benchmark versus under 100 microseconds for Homa.

_Anyone weighing TCP against newer transports for latency-sensitive clusters can track this debate on daily.dev._

### What is the Homa network protocol and how does it reduce tail latency compared to TCP?

Homa is a message-based transport protocol developed at Stanford, built as a clean-slate redesign for data centers, that controls congestion from the receiver rather than the sender and prioritizes shorter messages using shortest-remaining-processing-time scheduling plus switch priority queues. In one benchmark it achieved roughly 13 times lower P99 tail latency than TCP for short messages and nearly halved latency for the largest messages. It ships as a Linux kernel module available on GitHub, with upstreaming into the kernel underway.

_Teams evaluating new transport protocols for AI clusters can follow Homa's progress on daily.dev._

### Why is network latency becoming more important than throughput for AI workloads like inference and agentic applications?

Inference and agentic workloads increasingly rely on frequent small message exchanges, such as checking a distributed KV cache entry or barrier synchronization between compute rounds, rather than the massive gradient transfers typical of training. As computation phases shrink to millisecond scale, GPUs sit idle waiting on synchronization, so tail latency of small messages directly limits overall throughput, unlike in training where large-transfer throughput dominated.

_Developers tuning agentic and inference pipelines can keep up with latency-focused networking shifts via daily.dev._

## Similar posts on daily.dev

- [Latest Linux Patches For Homa Posted: TCP Alternative With 10~100x Lower Tail Latency](https://daily.dev/posts/latest-linux-patches-for-homa-posted-tcp-alternative-with-10-100x-lower-tail-latency-bygqzttkf) · Phoronix · 0 upvotes · 0 comments
- [What is RDMA?](https://daily.dev/posts/what-is-rdma--t1vgee9hu) · Ubuntu · 0 upvotes · 0 comments
- [Building a reliable cloud native foundation for distributed AI training](https://daily.dev/posts/building-a-reliable-cloud-native-foundation-for-distributed-ai-training-pwun78uwc) · CNCF · 1 upvotes · 0 comments
- [MetaRoCE: A New RDMA Transport Built for AI-Scale Ethernet](https://daily.dev/posts/metaroce-a-new-rdma-transport-built-for-ai-scale-ethernet-4rzmhaxu4) · Facebook Engineering
 · 0 upvotes · 0 comments

---

[View this post on daily.dev](https://daily.dev/posts/the-end-of-tcp-for-ai-clusters-john-ousterhout-stanford-swtimhc2i)

```json
{"@context":"https://schema.org","@graph":[{"@type":"Organization","@id":"https://daily.dev/#organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180},"sameAs":["https://twitter.com/dailydotdev","https://github.com/dailydotdev","https://www.linkedin.com/company/daily-dev-ltd"]},{"@type":"WebSite","@id":"https://daily.dev/#website","url":"https://daily.dev","name":"daily.dev","publisher":{"@id":"https://daily.dev/#organization"},"potentialAction":{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https://daily.dev/search?q={search_term_string}"},"query-input":"required name=search_term_string"}}]}
{"@context":"https://schema.org","@type":"TechArticle","headline":"The End of TCP for AI Clusters — John Ousterhout, Stanford","url":"https://daily.dev/posts/the-end-of-tcp-for-ai-clusters-john-ousterhout-stanford-swtimhc2i","mainEntityOfPage":{"@type":"WebPage","@id":"https://daily.dev/posts/the-end-of-tcp-for-ai-clusters-john-ousterhout-stanford-swtimhc2i"},"datePublished":"2026-09-17T13:00:32.607Z","dateModified":"2026-09-17T13:00:58.338Z","description":"John Ousterhout of Stanford argues AI workloads are shifting from massive, throughput-bound transfers to smaller, latency-sensitive exchanges, especially for...","image":"https://i.ytimg.com/vi/eZ8WWZzoaR0/sddefault.jpg","thumbnailUrl":"https://i.ytimg.com/vi/eZ8WWZzoaR0/sddefault.jpg","isAccessibleForFree":true,"articleSection":"AI Engineer","inLanguage":"en","publisher":{"@type":"Organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180}},"author":{"@type":"Organization","name":"AI Engineer","logo":"https://media.daily.dev/image/upload/s--u5PucxNT--/f_auto/v1724338940/logos/aidotengineer","url":"https://daily.dev/sources/aidotengineer"},"commentCount":0,"discussionUrl":"https://daily.dev/posts/the-end-of-tcp-for-ai-clusters-john-ousterhout-stanford-swtimhc2i","interactionStatistic":[{"@type":"InteractionCounter","interactionType":{"@type":"LikeAction"},"userInteractionCount":0},{"@type":"InteractionCounter","interactionType":{"@type":"CommentAction"},"userInteractionCount":0}],"keywords":"","timeRequired":"PT18M","video":{"@type":"VideoObject","name":"The End of TCP for AI Clusters — John Ousterhout, Stanford","description":"John Ousterhout of Stanford argues AI workloads are shifting from massive, throughput-bound transfers to smaller, latency-sensitive exchanges, especially for...","thumbnailUrl":"https://i.ytimg.com/vi/eZ8WWZzoaR0/sddefault.jpg","uploadDate":"2026-09-17T13:00:32.607Z","duration":"PT18M","url":"https://api.daily.dev/r/SWTiMHC2i","embedUrl":"https://www.youtube.com/embed/eZ8WWZzoaR0"}}
{"@context":"https://schema.org","@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https://daily.dev"},{"@type":"ListItem","position":2,"name":"AI Engineer","item":"https://daily.dev/sources/aidotengineer"},{"@type":"ListItem","position":3,"name":"The End of TCP for AI Clusters — John Ousterhout, Stanford"}]}
{"@context":"https://schema.org","@type":"FAQPage","@id":"https://daily.dev/posts/the-end-of-tcp-for-ai-clusters-john-ousterhout-stanford-swtimhc2i#faq","mainEntity":[{"@type":"Question","name":"Why does TCP perform poorly for small message latency in AI inference workloads?","acceptedAnswer":{"@type":"Answer","text":"TCP relies on sender-side congestion control, which reacts slowly because it takes several round trips for a sender to learn about congestion and adjust its rate, causing constant oscillation. TCP also serializes messages into an undifferentiated byte stream, so short messages can get stuck in queues behind large ones (head-of-line blocking), producing over a millisecond of tail latency in one benchmark versus under 100 microseconds for Homa. Anyone weighing TCP against newer transports for latency-sensitive clusters can track this debate on daily.dev."}},{"@type":"Question","name":"What is the Homa network protocol and how does it reduce tail latency compared to TCP?","acceptedAnswer":{"@type":"Answer","text":"Homa is a message-based transport protocol developed at Stanford, built as a clean-slate redesign for data centers, that controls congestion from the receiver rather than the sender and prioritizes shorter messages using shortest-remaining-processing-time scheduling plus switch priority queues. In one benchmark it achieved roughly 13 times lower P99 tail latency than TCP for short messages and nearly halved latency for the largest messages. It ships as a Linux kernel module available on GitHub, with upstreaming into the kernel underway. Teams evaluating new transport protocols for AI clusters can follow Homa's progress on daily.dev."}},{"@type":"Question","name":"Why is network latency becoming more important than throughput for AI workloads like inference and agentic applications?","acceptedAnswer":{"@type":"Answer","text":"Inference and agentic workloads increasingly rely on frequent small message exchanges, such as checking a distributed KV cache entry or barrier synchronization between compute rounds, rather than the massive gradient transfers typical of training. As computation phases shrink to millisecond scale, GPUs sit idle waiting on synchronization, so tail latency of small messages directly limits overall throughput, unlike in training where large-transfer throughput dominated. Developers tuning agentic and inference pipelines can keep up with latency-focused networking shifts via daily.dev."}}]}
```

