# Jalapeño’s first results show industry-leading speed and efficiency in AI inference

The 700-watt chip delivered 1.5 to 1.9 times higher throughput per watt and up to 3.6 times lower latency across open-weight models compared to commercial systems.

- OpenAI plans to begin deploying Jalapeño across its compute infrastructure by the end of 2026.
- The chip has a published TDP of 700 watts and sustained power at or below 550 watts on tested workloads.
- On Kimi K2.5 1T, Jalapeño reached 18,195 mixed TPS/kW and 1.56s latency compared to 11,862 mixed TPS/kW and 5.31s on Nvidia GB300.
- OpenAI used internal models to move from initial design to tapeout in nine months, and generated kernels that ran up to 1.8 times faster than human-written code for selected blocks.
- Successor chips Gen 2 and Gen 3 are already in development.

## Why it matters

OpenAI will lower its operational serving costs and inference latency as it begins running production workloads on first-party silicon.

## Sources

- [OpenAI: Jalapeño’s first results show industry-leading speed and efficiency in AI inference](https://openai.com/index/jalapeno-first-results)

---

Summarized by dstilled on 2026-08-25. https://dstilled.ai/story/37e75e57-2f24-490f-9406-aff53bd05234
