Jalapeño’s first results show industry-leading speed and efficiency in AI inference

The 700-watt chip delivered 1.5 to 1.9 times higher throughput per watt and up to 3.6 times lower latency across open-weight models compared to commercial systems.

OpenAI will lower its operational serving costs and inference latency as it begins running production workloads on first-party silicon.

Jalapeño’s first results show industry-leading speed and efficiency in AI inference

Sources

Read this as text

Back to the AI news