Jalapeño Chip: 3 Results OpenAI Reported
OpenAI reports efficiency and latency gains, but the figures are company-reported and deployment is planned for late 2026.
Jalapeño inference chip results published by OpenAI describe the company’s first custom accelerator as a way to improve AI-serving speed and power efficiency. The new report is based on OpenAI’s own testing, not a new consumer product or an independent comparison, and the company says deployment is still planned for later in 2026.
OpenAI says Jalapeño was designed with Broadcom and systems partner Celestica for large-scale language-model inference—the process of generating responses after a prompt is submitted. The company says the system is intended to help it serve more work at lower latency and with less power, but the figures in the report are OpenAI’s results and should be read as company-reported performance claims.

What OpenAI reports
OpenAI says it tested Jalapeño with the public InferenceX benchmark across GPT‑OSS 120B, DeepSeek R1 670B and Kimi K2.5 1T. Its stated comparison focuses on matching user experience rather than a single peak-throughput number, using throughput, power and latency together.
| OpenAI’s reported result | How to read it |
|---|---|
| 1.5× to 1.9× more work per watt at peak throughput | A company-reported efficiency range for the tested models and operating points. |
| 1.7× to 3.6× lower end-to-end latency | A company-reported response-time comparison, not a promise for every ChatGPT request. |
| 700W package rating; sustained power at or below 550W in the tests | Power figures reported by OpenAI for its measured workloads, not a retail specification. |
OpenAI’s earlier June announcement said engineering samples were running ML workloads in the lab and that the platform would be developed over multiple generations with Broadcom and Celestica. The August update adds detailed company results but does not announce a retail accelerator, a public cloud instance type, or a date when an individual user can buy Jalapeño hardware.
Why inference efficiency matters
Inference is the recurring cost of running an AI product after a model is trained. For interactive tools, slower token generation can compound across multi-step tasks. A system that delivers more work with the same electricity and hardware could let an operator improve capacity, latency or cost; it does not automatically translate into faster answers or lower prices for every user.
OpenAI says Jalapeño was designed around the different bottlenecks of prompt processing and token-by-token generation, including compute, memory bandwidth and network communication. The company’s position is that keeping more work local to the system can reduce data movement and waiting between components.
What this means for ChatGPT and API users
The practical near-term takeaway is limited. OpenAI says it plans to begin deploying Jalapeño within its own compute infrastructure by the end of 2026. It has not committed to a specific ChatGPT speed increase, API price reduction, model-exclusive feature, or rollout date for customers.
Infrastructure efficiency can support those outcomes over time, but they remain separate product decisions. OpenAI’s recent GPT‑5.6 Sol API pricing cut is an example of a product-price decision rather than proof that Jalapeño has already changed customer pricing.
Questions that remain
Are the benchmark results independently verified?
No independent verification is included in OpenAI’s announcement. The company says it used InferenceX, a public benchmark, but the reported comparisons and interpretation come from OpenAI.
Will users see faster ChatGPT responses now?
OpenAI does not say that Jalapeño is serving ChatGPT today. Its stated plan is to begin deployment within its compute infrastructure by the end of 2026.
Does Jalapeño replace NVIDIA hardware?
OpenAI says it will continue to deploy accelerators from NVIDIA and other partners for training and inference. The company presents Jalapeño as one part of a broader, multi-generation infrastructure plan.
Sources: OpenAI’s August 25 results report and its June Jalapeño announcement.
Source: OpenAI
