Jalapeño’s first results show industry-leading speed and efficiency in AI inference
The 700-watt chip delivered 1.5 to 1.9 times higher throughput per watt and up to 3.6 times lower latency across open-weight models compared to commercial systems.
- OpenAI plans to begin deploying Jalapeño across its compute infrastructure by the end of 2026.
- The chip has a published TDP of 700 watts and sustained power at or below 550 watts on tested workloads.
- On Kimi K2.5 1T, Jalapeño reached 18,195 mixed TPS/kW and 1.56s latency compared to 11,862 mixed TPS/kW and 5.31s on Nvidia GB300.
- OpenAI used internal models to move from initial design to tapeout in nine months, and generated kernels that ran up to 1.8 times faster than human-written code for selected blocks.
- Successor chips Gen 2 and Gen 3 are already in development.
OpenAI will lower its operational serving costs and inference latency as it begins running production workloads on first-party silicon.

Sources
Read this as text
Back to the AI news