OpenAI partnering with Broadcom on custom 'Jalapeño' inference chip to bypass Nvidia supply chain
The nine-month hardware development cycle signals a dramatic shift toward vertically integrated data center infrastructure for enterprise AI.
The astronomical cost of serving large language models has been forcing a fundamental shift in the artificial intelligence supply chain. OpenAI, in a direct collaboration with Broadcom, has unveiled its first proprietary Intelligence Processor, codenamed Jalapeño. The application-specific integrated circuit (ASIC) is built exclusively for large language model (LLM) inference—the heavy computational process of generating live user responses.
The development represents a massive pivot away from generic accelerators toward vertically integrated, application-specific silicon. By designing its own hardware, OpenAI aims to directly counter the steep capital expenditures associated with third-party components, where chipmakers like Nvidia historically command profit margins as high as 75%.
Jalapeño was designed by OpenAI to match its internal roadmap of models, software kernels, and serving systems. Broadcom handled the silicon engineering and high-performance networking integration, while TSMC was tasked with the physical manufacturing of the chip. Celestica completes the production loop by managing the engineering of the board and rack systems.
From a technical perspective, the chip addresses the primary bottleneck of modern interactive AI: data movement. Unlike general-purpose graphics processors adapted from legacy workloads, Jalapeño balances compute and memory resources to maximize data throughput. To scale these capabilities across clustered environments, the platform integrates Broadcom’s Tomahawk networking silicon directly into the design, preparing the hardware for massive data center environments.
Early engineering samples are already benchmarking frontier workloads, including an unreleased GPT-5.3-Codex-Spark model, at target production frequency and power. According to OpenAI, initial lab testing indicates a substantial improvement in performance-per-watt compared to current state-of-the-art accelerators.
Remaking the infrastructure flywheel
While competitors like Google have built proprietary hardware since 2015, OpenAI's entry into the silicon market marks a distinct business evolution. By controlling the entire technology stack, from the underlying chip architecture up to the ChatGPT application layer, OpenAI is implementing a highly optimized, closed-loop infrastructure strategy.
This vertical integration creates a distinct operational advantage. By pairing proprietary algorithms with custom silicon, the efficiency gains lower the cost of both training models and serving users at scale. The resulting cost reductions allow for continuous reinvestment into more capable model generations.
The operational timeline for the new silicon moved unusually fast. OpenAI and Broadcom transitioned the design from a blank slate to manufacturing tape-out in nine months. The acceleration was driven by OpenAI using its own advanced language models to design, simulate, and optimize portions of the hardware architecture.
Enterprise data center deployment is slated to begin by the end of 2026. Broadcom CEO Hock Tan confirmed that the hardware rollout will scale to gigawatt levels alongside cloud infrastructure partners, including Microsoft, fundamentally altering how enterprise AI compute is provisioned and priced in the coming years.