The OpenAI Jalapeno chip offers higher throughput and lower response latency for artificial intelligence workloads, according to new benchmark data shared during the Hot Chips conference. Details presented at the industry event indicate that the application-specific integrated circuit (ASIC) delivers performance gains over existing server hardware during model execution phases.

OpenAI revealed initial benchmark measurements from Semianalysis’s InferenceX platform, according to TechCrunch. Across tests using GPT-OSS 120B, DeepSeek R1, and Kimi K2.5 1T models, the processor delivered 1.5 to 1.9 times more work per watt than comparison systems running Nvidia GB200 or GB300 chips. Additionally, the system demonstrated 1.7 to 3.6 times lower end-to-end latency across those three evaluation models.

Architecture and Features of the OpenAI Jalapeno chip

Engineers tailored the OpenAI Jalapeno chip specifically to address bottlenecks during the inference lifecycle. Custom optimizations reduce delays in data movement during the prefill and communication phases of model processing. The architecture allows model state, including key-value caches, to remain local while coordinating compute, memory, and networking resources for each task.

“Jalapeño can serve more AI work per unit of power, while also returning responses more quickly. It’s very efficient to serve a lot of customers, but it can also be very low latency.”

Richard Ho, Head of Hardware at OpenAI

Benchmark Results on InferenceX

Standard inference systems often face trade-offs between serving multiple concurrent requests and keeping individual query latency low. Testing showed that the custom processor registered more tokens per user and higher throughput per kilowatt than prevailing market alternatives. Consequently, this design aims to provide stable processing speeds for demanding apps and autonomous software agents as request volume increases.

Co-Design Partnership with Broadcom

OpenAI developed the hardware in collaboration with Broadcom, using its internal AI models to support the engineering workflow. The company plans to maintain a multigenerational platform where software models, computing hardware, and memory systems evolve simultaneously. This integrated development model allows hardware adjustments that directly match model algorithmic requirements.

Deployment Timeline and Compute Strategy

Production rollout of the OpenAI Jalapeno chip will begin in small volumes by late 2026, followed by a larger volume increase during 2027. Despite the internal hardware development, OpenAI stated that it will continue working with external suppliers like Nvidia as part of its broader business strategy. Engineering teams have also commenced early development work on second and third generation successors.