OpenAI has unveiled initial performance metrics for Jalapeño, a custom-built inference chip designed in-house. Testing across three public models—GPT-OSS 120B, DeepSeek R1 670B and Kimi K2.5 1T—using the InferenceX benchmark from SemiAnalysis showed the chip delivering between 1.5 and 1.9 times more AI work per watt at peak throughput, alongside between 1.7 and 3.6 times lower end-to-end latency compared with the systems tested.

Power efficiency takes centre stage

For data centre operations, the power consumption figures stand out. Jalapeño carries a published chip power rating of 700W, though OpenAI noted that measured sustained power consumption stayed at or below 550W during the tested workloads.

OpenAI evaluated Jalapeño by measuring useful AI work completed per unit of power while maintaining required latency standards. Across the three models tested, the chip delivered between 1.5 and 1.9 times more work per watt at peak throughput. The GPT-OSS 120B model showed the strongest result, with approximately 1.9 times higher peak mixed tokens per second per kilowatt. DeepSeek R1 demonstrated around 1.7 times improvement, while Kimi K2.5 recorded approximately 1.5 times higher peak performance per watt.

Latency improvements accompanied efficiency gains

Beyond power efficiency, OpenAI reported between 1.7 and 3.6 times lower end-to-end latency across the three public models. DeepSeek R1 showed the largest latency improvement, with Jalapeño recording 1.65 seconds versus 5.99 seconds for the comparison system. Kimi K2.5 achieved 1.56 seconds against 5.31 seconds, while GPT-OSS 120B recorded 1.03 seconds compared with 1.80 seconds.

This combination matters significantly because infrastructure optimised for throughput typically involves compromises in response speed. OpenAI designed Jalapeño to achieve higher throughput and lower latency within the same architecture, particularly for interactive and agentic workloads where response delays can compound across multiple sequential operations.

System-level design rather than standalone accelerator

OpenAI achieved these results by designing the chip, memory, networking, software and rack-scale system together around language-model inference. Different inference stages impose varying demands on infrastructure. The initial prompt processing, or prefill stage, demands significant compute resources, whereas token-by-token response generation faces constraints from memory bandwidth availability.

Data movement between chips and other system components introduces latency. OpenAI designed Jalapeño to minimise such movement and maintain model state locally where feasible, while deploying the right combination of compute, memory and networking resources for each inference stage. This positions the chip as part of a broader system-level optimisation strategy rather than a standalone accelerator.

Deployment timeline and roadmap

OpenAI intends to begin deploying Jalapeño within its own infrastructure by the end of 2026. The company characterises the chip as the first generation of a multigenerational platform, with Gen 2 already in advanced development and Gen 3 in progress. Production qualification, software development and additional performance validation continue ahead of deployment.

The company is not abandoning third-party accelerators. OpenAI stated that meeting AI demand requires compute from multiple sources and will continue deploying accelerators from NVIDIA and other partners for both training and inference. Jalapeño thus supplements rather than replaces OpenAI's existing accelerator suppliers in its infrastructure strategy.

Power efficiency reshapes scaling priorities

Jalapeño's initial results illustrate how OpenAI approaches infrastructure demands from more advanced models and agentic workloads. Rather than measuring progress solely through accelerator performance, the company now emphasises the quantity of useful AI work generated from available power and hardware.

OpenAI suggested that extracting more useful work from existing resources could enable it to meet greater demand while reducing the cost of delivering results. For data centre operators, the reported efficiency gains carry practical weight. AI infrastructure expansion remains constrained by power availability, and boosting inference work delivered per watt provides another avenue for expanding capacity alongside physical infrastructure investment.

Jalapeño's initial results remain vendor-reported benchmarks pending deployment. However, with OpenAI preparing to introduce the chip into its infrastructure this year and already developing subsequent generations, custom silicon and system-level power efficiency are becoming increasingly central to its strategy for scaling AI inference.