<!-- mobian-agent-page publisher="dailydev" canonical="https://daily.dev/posts/routing-llm-inference-in-production-from-engine-signals-to-policy-qianru-lao-lu-zhang-openai-ll852jstc" -->

---
title: Routing LLM Inference in Production: From Engine Signals...
description: OpenAI inference engineers describe how their internal inference load balancer (IRB) evolved from a feedback-loop-driven weighted consistent hashing system to...
canonical: https://daily.dev/posts/routing-llm-inference-in-production-from-engine-signals-to-policy-qianru-lao-lu-zhang-openai-ll852jstc
twitter:card: summary_large_image
twitter:site: @dailydotdev
og:type: website
og:site_name: daily.dev
og:title: Routing LLM Inference in Production: From Engine Signals to Policy — Qianru Lao &amp; Lu Zhang, OpenAI | daily.dev
og:description: OpenAI inference engineers describe how their internal inference load balancer (IRB) evolved from a feedback-loop-driven weighted consistent hashing system to...
og:url: https://daily.dev/posts/routing-llm-inference-in-production-from-engine-signals-to-policy-qianru-lao-lu-zhang-openai-ll852jstc
og:image: https://api.daily.dev/og/posts/LL852JSTC.png
og:image:alt: Routing LLM Inference in Production: From Engine Signals to Policy — Qianru Lao &amp; Lu Zhang, OpenAI
og:image:width: 1200
og:image:height: 630
og:locale: en
---

> ## Documentation Index
> Fetch the complete documentation index at: https://daily.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# Routing LLM Inference in Production: From Engine Signals to Policy — Qianru Lao & Lu Zhang, OpenAI

**[AI Engineer](https://daily.dev/sources/aidotengineer)** · 18 min read · 1 upvotes · 0 comments

## Summary

OpenAI inference engineers describe how their internal inference load balancer (IRB) evolved from a feedback-loop-driven weighted consistent hashing system to a control-plane/data-plane architecture with an explicit optimization policy. The original approach used periodic performance scoring akin to a proportional controller, which adapted automatically but was hard to reason about and prone to oscillations that disrupted KV cache locality. The newer architecture separates a control plane, which computes globally optimized routing weights by minimizing expected end-to-end latency (accounting for network distance and engine-side latency under capacity constraints), from a data plane that makes fast local routing decisions using cached weights. The talk also covers production protection mechanisms: penalty-based weight reduction for outlier engines, dynamic retry budgets to avoid retry storms, and load shedding for graceful degradation under overload.

## Full article

daily.dev links to this article rather than hosting it. Read it at the original source: <https://www.youtube.com/watch?v=sOB3HSiG8vo>

## Questions this post answers

### How does OpenAI's inference load balancer decide which GPU engine should serve a request?

A control plane computes globally optimized routing weights by minimizing expected end-to-end latency across all traffic, factoring in network distance to each engine and engine-side latency under load. It respects constraints that all demand is routed, engines stay within capacity, and weights remain non-negative. A data plane in each CPU cluster then uses locally cached weights to make fast routing decisions without waiting on the control plane per request.

_Engineers designing inference routing systems can find deeper architecture breakdowns like this on daily.dev._

### Why did OpenAI move away from a feedback-loop based load balancer for LLM inference?

The original weighted consistent hashing approach, driven by a periodic feedback loop similar to a proportional controller, adapted well to observed engine performance but became hard to reason about since it combined many signals into one decision. It also caused bad oscillations: shifting traffic away from an engine made it appear underutilized, causing traffic to shift back and disrupting KV cache locality.

_Teams debugging routing oscillations in production inference systems can track these architectural tradeoffs on daily.dev._

### How do you prevent retry storms in a high-load inference serving system?

Dynamic retry budgets cap the number of allowed retries based on current system utilization, since unconstrained retries during heavy load add more traffic that causes more failures and triggers even more retries. The retry budget is more permissive during normal operation and tightens as the system approaches capacity, working alongside penalty-based weight reduction for outlier engines and proactive load shedding as a last resort.

_daily.dev surfaces production resilience patterns like retry budgeting for engineers building inference infrastructure._

## Similar posts on daily.dev

- [Multi-Provider LLM Routing Is Not a Problem, It's Your Architecture: Inference in Production Series](https://daily.dev/posts/multi-provider-llm-routing-is-not-a-problem-it-s-your-architecture-inference-in-production-series-b8ljpdbw8) · DigitalOcean Community · 0 upvotes · 0 comments
- [Optimizing LLM Inference Costs in Multi-Agent Systems with Adaptive Model Routing](https://daily.dev/posts/optimizing-llm-inference-costs-in-multi-agent-systems-with-adaptive-model-routing-jiqhng5u5) · Towards Data Science · 3 upvotes · 0 comments
- [How We Built DigitalOcean Inference Router](https://daily.dev/posts/how-we-built-digitalocean-inference-router-jpz3zpjxt) · DigitalOcean · 2 upvotes · 1 comments
- [LLM Routing Can Cost More Than Not Routing](https://daily.dev/posts/llm-routing-can-cost-more-than-not-routing-hujhqshkk) · Daily Dose of Data Science \| Avi Chawla \| Substack · 1 upvotes · 0 comments
- [Intelligent inference scheduling with llm-d on Red Hat AI](https://daily.dev/posts/intelligent-inference-scheduling-with-llm-d-on-red-hat-ai-q8lqp3mss) · Red Hat Developer · 0 upvotes · 0 comments

---

Tags: [#openai](https://daily.dev/tags/openai), [#distributed-systems](https://daily.dev/tags/distributed-systems), [#ai-inference](https://daily.dev/tags/ai-inference)

[View this post on daily.dev](https://daily.dev/posts/routing-llm-inference-in-production-from-engine-signals-to-policy-qianru-lao-lu-zhang-openai-ll852jstc)

```json
{"@context":"https://schema.org","@graph":[{"@type":"Organization","@id":"https://daily.dev/#organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180},"sameAs":["https://twitter.com/dailydotdev","https://github.com/dailydotdev","https://www.linkedin.com/company/daily-dev-ltd"]},{"@type":"WebSite","@id":"https://daily.dev/#website","url":"https://daily.dev","name":"daily.dev","publisher":{"@id":"https://daily.dev/#organization"},"potentialAction":{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https://daily.dev/search?q={search_term_string}"},"query-input":"required name=search_term_string"}}]}
{"@context":"https://schema.org","@type":"TechArticle","headline":"Routing LLM Inference in Production: From Engine Signals to Policy — Qianru Lao & Lu Zhang, OpenAI","url":"https://daily.dev/posts/routing-llm-inference-in-production-from-engine-signals-to-policy-qianru-lao-lu-zhang-openai-ll852jstc","mainEntityOfPage":{"@type":"WebPage","@id":"https://daily.dev/posts/routing-llm-inference-in-production-from-engine-signals-to-policy-qianru-lao-lu-zhang-openai-ll852jstc"},"datePublished":"2026-09-19T15:40:06.116Z","dateModified":"2026-09-19T15:40:28.885Z","description":"OpenAI inference engineers describe how their internal inference load balancer (IRB) evolved from a feedback-loop-driven weighted consistent hashing system to...","image":"https://i.ytimg.com/vi/sOB3HSiG8vo/sddefault.jpg","thumbnailUrl":"https://i.ytimg.com/vi/sOB3HSiG8vo/sddefault.jpg","isAccessibleForFree":true,"articleSection":"AI Engineer","inLanguage":"en","publisher":{"@type":"Organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180}},"author":{"@type":"Organization","name":"AI Engineer","logo":"https://media.daily.dev/image/upload/s--u5PucxNT--/f_auto/v1724338940/logos/aidotengineer","url":"https://daily.dev/sources/aidotengineer"},"commentCount":0,"discussionUrl":"https://daily.dev/posts/routing-llm-inference-in-production-from-engine-signals-to-policy-qianru-lao-lu-zhang-openai-ll852jstc","interactionStatistic":[{"@type":"InteractionCounter","interactionType":{"@type":"LikeAction"},"userInteractionCount":1},{"@type":"InteractionCounter","interactionType":{"@type":"CommentAction"},"userInteractionCount":0}],"keywords":"openai,distributed-systems,ai-inference","timeRequired":"PT18M","video":{"@type":"VideoObject","name":"Routing LLM Inference in Production: From Engine Signals to Policy — Qianru Lao & Lu Zhang, OpenAI","description":"OpenAI inference engineers describe how their internal inference load balancer (IRB) evolved from a feedback-loop-driven weighted consistent hashing system to...","thumbnailUrl":"https://i.ytimg.com/vi/sOB3HSiG8vo/sddefault.jpg","uploadDate":"2026-09-19T15:40:06.116Z","duration":"PT18M","url":"https://api.daily.dev/r/LL852JSTC","embedUrl":"https://www.youtube.com/embed/sOB3HSiG8vo"}}
{"@context":"https://schema.org","@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https://daily.dev"},{"@type":"ListItem","position":2,"name":"AI Engineer","item":"https://daily.dev/sources/aidotengineer"},{"@type":"ListItem","position":3,"name":"Routing LLM Inference in Production: From Engine Signals to Policy — Qianru Lao & Lu Zhang, OpenAI"}]}
{"@context":"https://schema.org","@type":"FAQPage","@id":"https://daily.dev/posts/routing-llm-inference-in-production-from-engine-signals-to-policy-qianru-lao-lu-zhang-openai-ll852jstc#faq","mainEntity":[{"@type":"Question","name":"How does OpenAI's inference load balancer decide which GPU engine should serve a request?","acceptedAnswer":{"@type":"Answer","text":"A control plane computes globally optimized routing weights by minimizing expected end-to-end latency across all traffic, factoring in network distance to each engine and engine-side latency under load. It respects constraints that all demand is routed, engines stay within capacity, and weights remain non-negative. A data plane in each CPU cluster then uses locally cached weights to make fast routing decisions without waiting on the control plane per request. Engineers designing inference routing systems can find deeper architecture breakdowns like this on daily.dev."}},{"@type":"Question","name":"Why did OpenAI move away from a feedback-loop based load balancer for LLM inference?","acceptedAnswer":{"@type":"Answer","text":"The original weighted consistent hashing approach, driven by a periodic feedback loop similar to a proportional controller, adapted well to observed engine performance but became hard to reason about since it combined many signals into one decision. It also caused bad oscillations: shifting traffic away from an engine made it appear underutilized, causing traffic to shift back and disrupting KV cache locality. Teams debugging routing oscillations in production inference systems can track these architectural tradeoffs on daily.dev."}},{"@type":"Question","name":"How do you prevent retry storms in a high-load inference serving system?","acceptedAnswer":{"@type":"Answer","text":"Dynamic retry budgets cap the number of allowed retries based on current system utilization, since unconstrained retries during heavy load add more traffic that causes more failures and triggers even more retries. The retry budget is more permissive during normal operation and tightens as the system approaches capacity, working alongside penalty-based weight reduction for outlier engines and proactive load shedding as a last resort. daily.dev surfaces production resilience patterns like retry budgeting for engineers building inference infrastructure."}}]}
```

