OpenAI shared first performance results for Jalapeño, its custom inference chip, showing 1.5-1.9x more AI work per watt and 1.7-3.6x lower end-to-end latency compared to leading commercial systems, tested on the InferenceX benchmark across GPT-OSS 120B, DeepSeek R1 670B, and Kimi K2.5 1T. The chip is rated at 700 watts but sustained measured power at or below 550 watts. Jalapeño's architecture minimizes data movement between prefill and decode phases through tight integration of compute, memory, and networking, and AI (including Codex with GPT-Astra) was used heavily in its design and in generating faster kernel implementations for new models. OpenAI plans to begin deploying Jalapeño within its own infrastructure by year end, with Gen 2 and Gen 3 chips already in development, while continuing to use NVIDIA and other accelerators alongside it.

8m read timeFrom openai.com
Post cover image
Table of contents
How we measured Jalapeño’s performanceArchitecting for speed and efficiency within a single chipWe used AI to design the chip, and designed the chip so AI could program itThe path ahead for efficient, ultra-fast inferenceAppendix

Questions this post answers

How much more efficient is OpenAI's Jalapeño inference chip compared to existing hardware?

Jalapeño delivers 1.5 to 1.9 times more AI work per watt at peak throughput and 1.7 to 3.6 times lower end-to-end latency than comparison systems across GPT-OSS 120B, DeepSeek R1 670B, and Kimi K2.5 1T. For highly interactive workloads it delivered 2.1 to 4.1 times higher performance. The chip is rated at 700 watts but measured sustained power stayed at or below 550 watts on tested workloads. Track how custom silicon like jalapeño is reshaping inference cost and latency tradeoffs on daily.dev.

When is OpenAI planning to deploy the Jalapeño chip in production?

OpenAI plans to begin deploying Jalapeño within its own compute infrastructure by the end of the year. It is the first generation of a multigenerational roadmap, with Gen 2 already deep in development and Gen 3 taking shape. OpenAI will continue widely deploying NVIDIA and other partner accelerators for both training and inference alongside Jalapeño. Follow rollout timelines for new AI accelerators like jalapeño as infrastructure choices evolve on daily.dev.

What benchmark was used to measure Jalapeño's inference performance against other AI systems?

Jalapeño was tested on InferenceX, a public benchmark from SemiAnalysis that measures the full process of serving an AI request, comparing throughput, power efficiency, and latency across an operating range from high-throughput serving to highly interactive, low-latency use. Results were normalized using each accelerator's published chip power rating rather than raw per-chip performance. Compare inference benchmarks like InferenceX when evaluating accelerator choices on daily.dev.

4 Impressions