OpenAI on August 25, 2026 published the first performance results for Jalapeño, which it describes as its first custom inference chip. According to OpenAI, the chip and the system built around it can serve more AI work per unit of power while returning responses more quickly, delivering both higher throughput and lower latency from a single architecture. OpenAI states it tested Jalapeño on InferenceX, a public benchmark from SemiAnalysis, and compared it with what it calls leading commercially available AI systems. Across GPT-OSS 120B, DeepSeek R1 670B, and Kimi K2.5 1T, OpenAI reports Jalapeño delivered 1.5 to 1.9 times more AI work per watt at peak throughput and 1.7 to 3.6 times lower end-to-end latency than comparison systems, with 2.1 to 4.1 times higher performance on highly interactive workloads. On Kimi K2.5 1T, the largest public model tested, OpenAI cites roughly 1.5 times higher peak performance per watt and 3.4 times lower latency. The company says Jalapeño is rated at 700 watts but stayed at or below 550 watts on the workloads tested. OpenAI also states earlier and current models helped design and program the chip. The results are self-reported by OpenAI and independent verification was not available.
- OpenAI reports Jalapeño delivered 1.5-1.9x more work per watt at peak throughput across three public models
- Latency was 1.7-3.6x lower and interactive performance 2.1-4.1x higher than comparison systems, per OpenAI
- Chip is rated at 700W but ran at or below 550W in tested workloads
- Tested on the SemiAnalysis InferenceX benchmark; results are self-reported
What it means for you
OpenAI built its own chip to run AI models faster and using less electricity, and it's now sharing numbers that suggest it works. If those gains hold up, they could eventually make OpenAI's products cheaper and more responsive for you. But these are OpenAI's own test results, not independently checked, and nothing changes for customers today.
Who should care
People who follow AI infrastructure economics, and anyone building latency-sensitive agent workflows that depend on OpenAI's serving costs and speed over the next year.
Skip this if
You use ChatGPT or the OpenAI API as an end user or small business — this is a hardware announcement with no action to take and no immediate effect on your bill or experience.
Sources: OpenAI — read the original