
OpenAI announced significant performance gains from its first custom artificial intelligence inference chip, Jalapeno, demonstrating substantial improvements over commercial systems. According to the company's announcement, Jalapeno delivered 1.5 to 1.9 times more AI work per watt at peak throughput and 1.7 to 3.6 times lower end-to-end latency across three large language models. For highly interactive workloads, performance was 2.1 to 4.1 times higher than comparison systems. The chip was tested on GPT-OSS 120B, DeepSeek R1 670B, and Kimi K2.5 1T using InferenceX benchmark from SemiAnalysis. OpenAI CEO Sam Altman confirmed the breakthrough, stating "We made a chip and it is fast" in a post on X. The results demonstrate that OpenAI's custom chip architecture can work across models developed both by OpenAI and external developers, with the company emphasizing that greater inference efficiency could improve operating leverage by allowing useful work and revenue to grow faster than the cost of serving users.
On Kimi K2.5 1T, the largest public model in the test, OpenAI reported 1.5 times higher peak performance per watt and 3.4 times lower end-to-end latency than the comparison system. Jalapeno is rated at 700 watts, though its measured sustained power remained at or below 550 watts during the workloads tested. The chip is an application-specific integrated circuit (ASIC) designed in collaboration with Broadcom and manufactured by TSMC. Unlike Nvidia's general-purpose GPUs that handle everything from training massive models to running inference at scale, Jalapeno is purpose-built exclusively for AI inference. The company designed the chip, memory, networking and software as a single system, describing this as a 'full-stack advantage' that allows integrated design of models, products, serving software, chips, memory, networking and systems together. The architecture is meant to stay useful as the balance between prompt work and response generation changes, with the mechanism placing model state—including the KV cache—locally while activating necessary compute, memory and network resources for each stage.
OpenAI confirmed that it used its own AI models, including Codex, to help design Jalapeno, taking the project from initial design to tapeout in approximately nine months. According to the company, AI-generated attention and mixture-of-experts kernels reportedly ran 1.5 to 1.8 times faster than versions hand-written by human experts on select blocks, suggesting AI-assisted chip design is moving into shipping silicon rather than remaining a demonstration. The company emphasized that it will continue to widely deploy accelerators from Nvidia and other partners for both training and inference workloads, particularly for model training where Nvidia still dominates. OpenAI noted that producing more useful AI work from the same amount of power and hardware could allow the company to serve more demand without costs rising at the same pace. The company is also presenting Jalapeno as a software-development experiment, noting that AI helped the team reach tapeout in nine months, and that Codex with GPT-Astra brought three open-weight models outside the original production plan to high performance within two months.
OpenAI plans to begin deploying Jalapeno within its own computing infrastructure by the end of 2026, with a very small deployment expected at that time, followed by more significant deployment in 2027. The company reportedly went from schematic to tape-out in approximately nine months, with the company's own AI models assisting in the chip design process. According to OpenAI, the chip could mean faster ChatGPT responses, more responsive Codex coding sessions and AI agents, while also helping handle growing demand for AI services. The company emphasized that it will continue to widely deploy accelerators from Nvidia and other partners for both training and inference workloads, particularly for model training where Nvidia still dominates. OpenAI noted that producing more useful AI work from the same amount of power and hardware could allow the company to serve more demand without costs rising at the same pace. The company is also presenting Jalapeno as a software-development experiment, noting that AI helped the team reach tapeout in nine months, and that Codex with GPT-Astra brought three open-weight models outside the original production plan to high performance within two months.
Jalapeno represents the first generation of what OpenAI described as a multigenerational custom silicon roadmap, with Gen 2 deep in development and Gen 3 taking shape. The company emphasized that each generation will build on what they learn and further advance both efficiency and speed. OpenAI noted that meeting growing demand will require compute from multiple sources, with the company continuing to deploy accelerators from Nvidia and other partners for both AI training and inference. The current results are based on ongoing testing and the company is continuing to prepare Jalapeno for operation at scale. Broadcom CEO Hock Tan went further, claiming Jalapeno matches the performance of Nvidia's Blackwell architecture and Google's TPU while offering roughly a 50% cost advantage on a per-token and per-kilowatt basis. The results could improve economics for interactive agents, but they remain controlled benchmarks as tests used short, single-turn 8k/1k workloads and excluded longer-context AgentX scenarios.