MobbleOpen in Mobble ⇢
Technology · Semiconductors · published 2026-08-25 · via TechCrunch

OpenAI's custom inference chip posts record efficiency in early benchmarks

OpenAI presented benchmark results for its Jalapeño inference chip at Hot Chips, showing it delivers higher tokens per user and throughput per kilowatt than current state-of-the-art processors like Nvidia's Blackwell. The chip, developed with Broadcom, is designed to minimize delays in prefill and communication phases. Deployment is expected in small volumes by end of 2026, with larger scale in 2027.

Expanded Detail

The benchmark results were presented at the Hot Chips conference, where OpenAI's hardware chief Richard Ho highlighted the chip's dual strengths in both throughput and latency. Jalapeño's design prioritizes keeping model state, including the KV cache, local to reduce data movement during response generation. The chip emerged from a close collaboration with Broadcom, with OpenAI's own models contributing to the development process.

OpenAI positions Jalapeño as the first in a multigenerational platform, with plans to co-develop AI products, models, chips, and memory together. The company explicitly targeted bottlenecks in prefill and communication phases during inference. Deployment begins in small volumes at the end of 2026, scaling up through 2027, though competitors like Nvidia may advance their own offerings in that timeframe.

Context

This chip could reshape the economics of AI inference, potentially lowering energy costs for large-scale AI services and enabling faster response times for end users. If Jalapeño delivers on its efficiency claims, it may pressure established chipmakers to accelerate their own designs, benefiting businesses that rely on AI infrastructure. However, the 2026-2027 timeline means near-term impact is limited, and broader adoption may depend on how quickly OpenAI scales production and whether the performance advantage persists against next-generation competitors.

Expanded detail and Context are AI-generated analysis; the linked article remains the authoritative source.
Read the full article at TechCrunch →
This summary is Al-enhanced to contain extended analysis and broader social context. The original is {NAME); the linked article is the authoritative source. Original headline: “OpenAI’s Jalapeño chip is built for fast inference at scale, benchmarks show.” Browse more stories.