OpenAI says its new AI chip, Jalapeño, can complete AI inference tasks more efficiently and return responses faster than competing systems. The company is positioning the chip as a way to improve speed, agent responsiveness, and access as demand for AI systems grows.
What OpenAI says Jalapeño is built to do
Jalapeño is an Application-Specific Integrated Circuit, or ASIC, made in partnership with Broadcom. OpenAI first introduced Jalapeño in June, and the chip is designed for AI inference.
Inference is the stage where a trained AI model is used to complete a task or deploy an agent. In plain terms, it is the part of the AI process that affects how quickly a system can respond after a user asks it to do something.
OpenAI hardware vice president Richard Ho said during a reporter briefing that Jalapeño offers the “best of both worlds” by combining lower latency with higher throughput. He said AI systems typically “have to make a trade-off between the two.”
That trade-off matters because throughput and latency describe different parts of performance. Throughput is about how much work a system can handle. Latency is about how long it takes to return a response. OpenAI’s claim is that Jalapeño improves both at the same time.
How the benchmark compared Jalapeño with Nvidia systems
OpenAI measured Jalapeño using InferenceX, a benchmarking platform focused on how AI systems handle inference. The comparison was made against the best results recorded at the time, which were with Nvidia’s GB200 or GB300 superchips.
According to OpenAI, Jalapeño delivered 1.5 to 1.9 times more AI work per watt across GPT-OSS 120B, DeepSeek R1, and Kimi K2.5 1T than the comparison systems. The company also says Jalapeño provided 1.7 to 3.6 times lower end-to-end latency across the three models.
Those two measurements point to different benefits. More AI work per watt suggests better efficiency for inference workloads. Lower end-to-end latency points to faster responses from the time a task starts to the time it finishes.
OpenAI’s framing is that these improvements could translate into “faster responses, more responsive agents, and more reliable access as the demand grows,” according to Ho. The company is not presenting the chip only as a speed upgrade, but also as part of a broader effort to make AI systems more available under rising usage.
Why latency and throughput both matter
For users, latency is the most visible part of performance. A lower-latency system can feel more immediate because it returns answers more quickly. That matters for general AI responses, and it matters even more for agents that may need to complete tasks through multiple steps.
Throughput matters in a different way. If a system can process more AI work at once, it can better handle demand. OpenAI’s benchmark claim ties throughput to efficiency by measuring AI work per watt, which focuses on how much output the chip can deliver for the power it uses.
The key point in OpenAI’s announcement is not just that Jalapeño was faster in one dimension. The company says the chip performed better on both efficiency and latency across the three tested models:
- GPT-OSS 120B
- DeepSeek R1
- Kimi K2.5 1T
The comparison systems were Nvidia’s GB200 or GB300 superchips, which OpenAI described as having the best recorded results at the time of the test. That makes the benchmark notable, though it remains a benchmark result rather than a full picture of deployment at scale.
Deployment will start small
OpenAI does not plan to move Jalapeño into broad use immediately. Ho said the company plans to deploy Jalapeño in “small volumes” by the end of this year, then begin to “ramp the volume up” into 2027.
The company did not say how many Jalapeño chips it expects to deploy next year. That leaves the scale of the rollout unclear, even as OpenAI describes strong benchmark performance.
The cautious rollout also shows that Jalapeño is not being treated as an instant replacement for OpenAI’s existing compute approach. Even with the stated performance gains, Ho said OpenAI does not expect to replace its entire chip lineup with Jalapeño.
Instead, OpenAI’s compute strategy will still include “very good partners,” like Nvidia. That means Jalapeño appears to be one part of the company’s AI infrastructure plans rather than a complete shift away from outside chip suppliers.
What comes next for OpenAI’s chip plans
OpenAI will continue developing the second and third generations of Jalapeño. The source does not provide technical details on those future generations, but it does make clear that OpenAI views the chip as an ongoing hardware project.
For now, the most important facts are the benchmark claims and the rollout timeline. OpenAI says Jalapeño outperformed Nvidia’s GB200 or GB300 systems on InferenceX for the tested models, with higher AI work per watt and lower end-to-end latency.
If those results carry into real deployments, Jalapeño could help OpenAI deliver faster AI responses and more responsive agents as usage increases. The next test will be how the chip performs once it moves from benchmark results into the small-volume deployment OpenAI says is planned by the end of this year.