OpenAI Publishes First Jalapeño Chip Benchmarks Ahead of Late-2026 Deployment

OpenAI has published the first measured performance results for Jalapeño, its custom processor for artificial-intelligence inference, as the company prepares to deploy the system in production data centres beginning in late 2026.

The chip was designed for inference, the stage in which trained AI models process user requests and generate outputs. It was co-developed with Broadcom, while Celestica is supporting board, rack and system integration. OpenAI said the processor and associated systems moved from initial design to manufacturing tape-out in nine months.

Testing used InferenceX, a public, power-normalised benchmark covering the complete inference path from processing an input prompt to generating output tokens. OpenAI evaluated Jalapeño using its GPT-OSS 120B open model, DeepSeek R1 and the trillion-parameter version of Kimi K2.5.

Across those workloads, Jalapeño-based systems produced between 1.5 and 1.9 times more AI work at peak throughput and recorded between 1.7 and 3.6 times lower end-to-end latency than the strongest commercially available comparison systems used in the tests. The processor has a reported thermal-design power of 700 watts. The results are early benchmarks rather than production measurements across OpenAI’s complete proprietary model portfolio.

OpenAI designed Jalapeño to improve throughput and latency simultaneously. Conventional inference deployments can require operators to optimise systems for one of those measures at the expense of the other: high aggregate throughput lowers serving cost, while low latency improves response time for individual users. The company says its hardware, networking and software stack allows both measures to improve within a single architecture.

Jalapeño is a general large-language-model inference accelerator rather than a processor restricted to OpenAI models, as demonstrated by the tests involving DeepSeek and Kimi systems. It is not designed to train frontier models, leaving OpenAI dependent on external accelerator suppliers for a substantial part of its computing requirements. The company also does not currently plan to sell Jalapeño to third parties.

The chip forms part of a broader infrastructure strategy integrating processors, data centres, models, developer services and consumer and enterprise products. OpenAI intends to retain multiple infrastructure and hardware partners, using performance, cost and deployment availability to allocate workloads.

India is one of ChatGPT’s largest weekly-active-user markets and ranks among the top five countries for OpenAI API usage. OpenAI has separately agreed to become the first customer of Tata Consultancy Services’ HyperVault data-centre business, initially taking 100 megawatts of capacity with an option to scale to one gigawatt. The custom processor’s deployment therefore intersects with a growing India-based infrastructure and enterprise footprint, although OpenAI has not stated whether the initial Indian capacity will use Jalapeño systems.

- Advertisement -

LEAVE A REPLY

Please enter your comment!
Please enter your name here

Latest Articles

Share your details to download the Research Report 2026

Share your details to download the CISO Handbook 2026

Share your details to download the report 2026

Share your details to download the Cybersecurity Report 2025

Share your details to download the CISO Handbook 2025

Sign Up for CXO Digital Pulse Newsletters

Share your details to download the Research Report

Share your details to download the Coffee Table Book

Share your details to download the Vision 2023 Research Report

Download 8 Key Insights for Manufacturing for 2023 Report

Sign Up for CISO Handbook 2023

Download India’s Cybersecurity Outlook 2023 Report

Unlock Exclusive Insights: Access the article

Download CIO VISION 2024 Report

Share your details to download the report

Share your details to download the CISO Handbook 2024

Fill your details to Watch