OpenAI's Jalapeño chip claims big efficiency and latency wins over Nvidia's GB300 in first benchmarks
OpenAI unveiled benchmark results at Hot Chips showing its custom Jalapeño inference ASIC, co-developed with Broadcom, delivers 1.5 to 1.9 times more throughput per kilowatt and 1.7 to 3.6 times lower latency than Nvidia's GB200 and GB300 rack systems on the InferenceX suite. The 700W chip outperformed Nvidia parts rated at 1,200W and 1,400W, with the widest margins at low-latency settings, though OpenAI noted the comparison used package TDP and that utility-power and multi-token prediction scenarios narrow the gap. Jalapeño is slated for deployment in OpenAI's data centers later this year, but it was not tested against Nvidia's upcoming Vera Rubin and does not handle model training.
Related stories
OpenAI claims its new AI chip beats Nvidia on inference benchmarks · Semiconductors
Waymo reveals custom 5nm chip for its next-gen robotaxi compute · Semiconductors
Also covered by: The Verge
This summary is AI-generated and original to Mobble; the linked article is the authoritative source.
Original headline: “OpenAI’s 700W Jalapeño ASIC outpaces 1,400W Nvidia flagship GPU — claims up to 1.9x throughput per kilowatt and 3.6x lower latency, co-developed with Broadcom.” Browse more stories.