OpenAI and Broadcom Unveil Jalapeño, a Custom LLM Inference Accelerator
OpenAI and Broadcom announced Jalapeño, a purpose-built AI chip for language model inference, designed to improve efficiency and reduce costs at scale.
Last verified:
OpenAI and Broadcom Announce Custom Inference Chip
OpenAI and Broadcom have unveiled Jalapeño, a custom-built AI accelerator designed specifically for large language model inference. According to the OpenAI Blog, the chip represents the first step in a multi-generation compute platform designed to make advanced AI faster, cheaper, and more accessible. Engineering samples are already running production workloads in OpenAI’s labs, including the company’s GPT-5.3-Codex-Spark model, at target frequency and power consumption levels.
Nine-Month Development Cycle and Full-Stack Integration
The Jalapeño chip was developed from design to production in nine months, a compressed timeline enabled by OpenAI’s internal model roadmap and inference requirements. According to OpenAI, the accelerator was architected around the company’s “deep understanding of LLM fundamentals,” drawing on insights from its serving systems, kernel designs, and product needs. Broadcom President and CEO Hock Tan and President Charlie Kawwas delivered the first engineering samples to OpenAI CEO Sam Altman and President Greg Brockman, signaling the partnership’s maturity.
Broadcom handled silicon implementation and high-performance networking, including its Tomahawk networking technology, while Celestica contributed to board design and rack-level system integration for production scale.
Performance Gains and Architecture Focus
Early testing shows Jalapeño will deliver “substantially better” performance per watt than current state-of-the-art inference hardware, according to OpenAI’s announcement. The accelerator reduces data movement and balances compute, memory, and networking resources to approach theoretical peak utilization more closely than existing solutions. A detailed technical report is promised in the coming months.
The chip is designed to work with all language models, not just OpenAI’s products, guided by the company’s analysis of industry-wide inference patterns. This flexibility suggests OpenAI intends to position Jalapeño as a general-purpose LLM accelerator, though initial deployment will prioritize internal workloads.
Why This Matters
OpenAI’s move into custom silicon reshapes the economics of LLM inference at scale. By controlling the full stack—from models through serving systems to hardware—the company can eliminate inefficiencies that arise from general-purpose processors. Deployment at “gigawatt scale” with multiple data center partners implies OpenAI plans both internal use and potential licensing to third parties, potentially creating a new revenue stream while reducing inference costs for its products like ChatGPT and enterprise offerings.
For competitors and cloud providers, Jalapeño signals that custom silicon is becoming table stakes for inference cost leadership. The nine-month development cycle also demonstrates that specialized AI hardware can reach production faster than assumed, compressing the timeline for similar initiatives at rivals like Google, Meta, and Amazon.
Frequently Asked Questions
What is Jalapeño and what does it do?
Jalapeño is OpenAI's first custom inference accelerator, co-developed with Broadcom. It is designed to run large language model inference workloads with improved energy efficiency and performance per watt compared to existing hardware.
When will Jalapeño be available?
OpenAI has not announced a public availability date. The chip is currently in engineering samples running production workloads in the lab. Deployment is planned at 'gigawatt scale' with data center partners over multiple generations.
Why is OpenAI building its own chips?
According to OpenAI President Greg Brockman, custom silicon allows OpenAI to serve 'more intelligence with greater efficiency' and advance its full-stack infrastructure strategy, reducing costs and improving access to advanced AI.