OpenAI shared benchmark data for its custom Jalapeno chip, stating the processor outperformed Nvidia GB300 systems in speed and energy efficiency. Developed in collaboration with Broadcom, the silicon targets AI inference workloads rather than model training. The company published test results comparing the hardware across several leading open-source models.

Jalapeno Chip Benchmark Results Across Models

Testing utilized InferenceX, an evaluation benchmark from SemiAnalysis designed to measure end-to-end request serving. OpenAI evaluated the hardware on three large language models: GPT-OSS 120B, DeepSeek R1 670B, and Kimi K2.5 1T. Across these tests, the Jalapeno chip delivered 1.5 to 1.9 times more output per watt at peak throughput compared to competing hardware.

Latency decreased by 1.7 to 3.6 times across the evaluated workloads. Specifically, on interactive agent tasks, the performance advantage reached 2.1 to 4.1 times. On the 1-trillion-parameter Kimi K2.5 model, the processor achieved roughly 1.5 times better performance per watt alongside 3.4 times lower latency.

Power Efficiency and System Architecture

Power measurements played a central role in the comparative evaluation. The hardware has a rated package power of 700 watts, compared to 1,400 watts for the GB300. Moreover, OpenAI stated that actual power draw remained at or below 550 watts during testing.

The physical design separates processing into two distinct language-model phases: compute-intensive prefill and memory-bandwidth-dependent decode. Consequently, keeping model state and cache data local within a single connected system reduces communication delays between components. This architectural approach helps modern artificial intelligence clusters operate with minimal data transfer bottlenecks.

Development Timeline and Internal Tooling

Internal software tools helped shorten the processor development cycle. Machine learning models contributed to the design phase, reducing the path from initial concept to tapeout to nine months. Furthermore, engineers applied the Codex tool to optimize software performance for three additional open-weight models within two months.

Generated code for specific model components ran up to 1.8 times faster than human-written software. These advancements influence broader computers and laptops hardware engineering methods. Meanwhile, enterprise software integration continues to expand across modern apps and software development pipelines.

Infrastructure Deployment and Commercial Strategy

OpenAI plans to deploy the Jalapeno chip across its internal data center infrastructure by the end of the year. The processor represents the first milestone in a multi-generation silicon roadmap, with follow-up versions already in development. However, the company confirmed it will continue purchasing hardware from external vendors to fulfill capacity needs.

The benchmark tests did not include Nvidia’s newer Vera Rubin processors, which recently entered distribution. Additionally, the new processor focuses strictly on inference execution rather than model training operations. Market shifts in chip design continue to impact the broader digital economy and business sector globally.