Megadose Built for builders and researchers.

Jalapeño’s first results show industry-leading speed and efficiency in AI inference

OpenAI Blog ·
OpenAI says its first custom inference chip returns lower-latency responses while doing more work per watt than comparison systems.

Jalapeño tested across GPT-OSS 120B, DeepSeek R1, and Kimi K2.5 1T, with reported gains of 1.5x to 1.9x more AI work per watt at peak throughput. OpenAI also says end-to-end latency was 1.7x to 3.6x lower, and highly interactive workloads saw 2.1x to 4.1x higher performance. The company attributes the gains to designing the chip, memory, networking, software, and rack-scale system together for language-model inference. OpenAI says it plans to begin deploying Jalapeño inside its compute infrastructure by year-end, while continuing to use NVIDIA and other accelerators. OpenAI's note

score 7

Categories: Money & Moves