---
title: "Jalapeño’s first results show industry-leading speed and efficiency in AI inference"
url: https://daily.dev/posts/jalape-o-s-first-results-show-industry-leading-speed-and-efficiency-in-ai-inference-lscwxx1vx
source_url: https://openai.com/index/jalapeno-first-results
type: article
source: "OpenAI"
published: 2026-08-25T14:46:00.450Z
updated: 2026-08-25T14:57:51.975Z
tags: ["openai", "gpu", "ai-inference"]
reading_time: 8
upvotes: 0
comments: 0
language: en
---

> ## Documentation Index
> Fetch the complete documentation index at: https://daily.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# Jalapeño’s first results show industry-leading speed and efficiency in AI inference

**[OpenAI](https://daily.dev/sources/openai)** · 8 min read · 0 upvotes · 0 comments

## Summary

OpenAI shared first performance results for Jalapeño, its custom inference chip, showing 1.5-1.9x more AI work per watt and 1.7-3.6x lower end-to-end latency compared to leading commercial systems, tested on the InferenceX benchmark across GPT-OSS 120B, DeepSeek R1 670B, and Kimi K2.5 1T. The chip is rated at 700 watts but sustained measured power at or below 550 watts. Jalapeño's architecture minimizes data movement between prefill and decode phases through tight integration of compute, memory, and networking, and AI (including Codex with GPT-Astra) was used heavily in its design and in generating faster kernel implementations for new models. OpenAI plans to begin deploying Jalapeño within its own infrastructure by year end, with Gen 2 and Gen 3 chips already in development, while continuing to use NVIDIA and other accelerators alongside it.

## Full article

daily.dev links to this article rather than hosting it. Read it at the original source: <https://openai.com/index/jalapeno-first-results>

## Questions this post answers

### How much more efficient is OpenAI's Jalapeño inference chip compared to existing hardware?

Jalapeño delivers 1.5 to 1.9 times more AI work per watt at peak throughput and 1.7 to 3.6 times lower end-to-end latency than comparison systems across GPT-OSS 120B, DeepSeek R1 670B, and Kimi K2.5 1T. For highly interactive workloads it delivered 2.1 to 4.1 times higher performance. The chip is rated at 700 watts but measured sustained power stayed at or below 550 watts on tested workloads.

_Track how custom silicon like jalapeño is reshaping inference cost and latency tradeoffs on daily.dev._

### When is OpenAI planning to deploy the Jalapeño chip in production?

OpenAI plans to begin deploying Jalapeño within its own compute infrastructure by the end of the year. It is the first generation of a multigenerational roadmap, with Gen 2 already deep in development and Gen 3 taking shape. OpenAI will continue widely deploying NVIDIA and other partner accelerators for both training and inference alongside Jalapeño.

_Follow rollout timelines for new AI accelerators like jalapeño as infrastructure choices evolve on daily.dev._

### What benchmark was used to measure Jalapeño's inference performance against other AI systems?

Jalapeño was tested on InferenceX, a public benchmark from SemiAnalysis that measures the full process of serving an AI request, comparing throughput, power efficiency, and latency across an operating range from high-throughput serving to highly interactive, low-latency use. Results were normalized using each accelerator's published chip power rating rather than raw per-chip performance.

_Compare inference benchmarks like InferenceX when evaluating accelerator choices on daily.dev._

---

Tags: [#openai](https://daily.dev/tags/openai), [#gpu](https://daily.dev/tags/gpu), [#ai-inference](https://daily.dev/tags/ai-inference)

[View this post on daily.dev](https://daily.dev/posts/jalape-o-s-first-results-show-industry-leading-speed-and-efficiency-in-ai-inference-lscwxx1vx)
