OpenAI's CFO Sarah Friar lays out the company's compute strategy as an integrated system spanning chips, data centers, models, and products. The post highlights first performance results from Jalapeño, OpenAI's first custom inference chip, which reportedly beat commercial systems on throughput per kilowatt and latency using GPT-OSS 120B, DeepSeek R1, and Kimi K2 benchmarks. It also describes a diversified hardware and cloud partner portfolio (Microsoft, NVIDIA, AWS, AMD, Broadcom, Cerebras, CoreWeave, Oracle, SB Energy, SoftBank), the Project Camellia data center in Georgia, and claims that GPT-5.6 Sol achieved a new high on the Artificial Analysis Coding Agent Index while using 54% fewer output tokens than a competing model.

4m read timeFrom openai.com
Post cover image
Table of contents
Jalapeño widens the lead at previous-best TBTBuild for breadth, own for leverageTurning efficiency into economic valueA compounding advantage

Questions this post answers

What is OpenAI's Jalapeño chip and how does it perform on inference benchmarks?

Jalapeño is OpenAI's first custom inference chip. On the InferenceX public benchmark using GPT-OSS 120B, it delivered more peak throughput per kilowatt and lower token latency than commercial systems it was compared against, and it also performed strongly on DeepSeek R1 and Kimi K2, indicating gains across different model families. Track how custom silicon like Jalapeño reshapes inference costs and latency for AI-powered apps on daily.dev.

Which cloud and hardware providers does OpenAI use for compute besides Microsoft and NVIDIA?

OpenAI's compute portfolio also includes AWS, AMD, Broadcom, Cerebras, CoreWeave, Oracle, SB Energy, and SoftBank, alongside its foundational partners Microsoft (compute) and NVIDIA (chips). Each partner contributes different strengths, such as cloud infrastructure, accelerated computing, low-latency inference, data-center development, and energy delivery, letting OpenAI direct workloads toward the best performance per dollar. Compare AI infrastructure vendor strategies like this one alongside other backend architecture coverage on daily.dev.

How much more token-efficient is GPT-5.6 Sol compared to other coding models?

GPT-5.6 Sol with max reasoning reached a new high on the Artificial Analysis Coding Agent Index while using 54% fewer output tokens than another leading model. Fewer output tokens for equal or better results translate into faster responses, fewer retries, longer completed agent workflows, and lower total cost for successful coding tasks. Follow token-efficiency gains in coding models like GPT-5.6 Sol to judge real costs on daily.dev.

3 Impressions