OpenAI’s Jalapeño chip is built for fast inference at scale, benchmarks show
OpenAI says Jalapeño beat current top inference processors on speed per user and throughput per watt in early benchmarks.
The comparison cited SemiAnalysis’ InferenceX test and included an Nvidia Blackwell system. OpenAI’s Richard Ho said the chip is aimed at serving more customers with lower latency and better power efficiency. Full deployment is not near-term: Ho said very small volumes are expected at the end of 2026, with broader rollout in 2027. The chip was developed with Broadcom and is designed to reduce inference bottlenecks in prefill, communication, and KV-cache handling. TechCrunch AI's note
The comparison cited SemiAnalysis’ InferenceX test and included an Nvidia Blackwell system. OpenAI’s Richard Ho said the chip is aimed at serving more customers with lower latency and better power efficiency. Full deployment is not near-term: Ho said very small volumes are expected at the end of 2026, with broader rollout in 2027. The chip was developed with Broadcom and is designed to reduce inference bottlenecks in prefill, communication, and KV-cache handling. TechCrunch AI's note
score 8