OpenAI's custom inference chip posts record efficiency in early benchmarks
OpenAI presented benchmark results for its Jalapeño inference chip at Hot Chips, showing it delivers higher tokens per user and throughput per kilowatt than current state-of-the-art processors like Nvidia's Blackwell. The chip, developed with Broadcom, is designed to minimize delays in prefill and communication phases. Deployment is expected in small volumes by end of 2026, with larger scale in 2027.
The benchmark results were presented at the Hot Chips conference, where OpenAI's hardware chief Richard Ho highlighted the chip's dual strengths in both throughput and latency. Jalapeño's design prioritizes keeping model state, including the KV cache, local to reduce data movement during response generation. The chip emerged from a close collaboration with Broadcom, with OpenAI's own models contributing to the development process.
OpenAI positions Jalapeño as the first in a multigenerational platform, with plans to co-develop AI products, models, chips, and memory together. The company explicitly targeted bottlenecks in prefill and communication phases during inference. Deployment begins in small volumes at the end of 2026, scaling up through 2027, though competitors like Nvidia may advance their own offerings in that timeframe.
This chip could reshape the economics of AI inference, potentially lowering energy costs for large-scale AI services and enabling faster response times for end users. If Jalapeño delivers on its efficiency claims, it may pressure established chipmakers to accelerate their own designs, benefiting businesses that rely on AI infrastructure. However, the 2026-2027 timeline means near-term impact is limited, and broader adoption may depend on how quickly OpenAI scales production and whether the performance advantage persists against next-generation competitors.