OpenAI just shipped its first custom inference chip. The company used its own AI models to help design the chip that will now run its AI models.
OpenAI released initial performance results for the chip, called Jalapeño, this week. The company says it delivers higher throughput and lower latency at the same time. Most existing hardware has to choose between the two.
The Numbers, With Appropriate Context
These figures come from OpenAI's own testing, using InferenceX, a public benchmark built by the analytics firm SemiAnalysis. OpenAI ran the comparisons and reported the results itself.
OpenAI tested Jalapeño across three models: GPT-OSS 120B, DeepSeek R1, and Kimi K2.5. It reports the chip delivered 1.5 to 1.9 times more AI work per watt at peak throughput. It claims 1.7 to 3.6 times lower end-to-end latency than the comparison systems.
Those comparison systems were Nvidia's GB200 and GB300 chips, rated at 1,200 and 1,400 watts respectively. Jalapeño is rated at 700 watts, and OpenAI says its measured sustained power stayed at or below 550 watts during testing.
That power gap is worth sitting with. OpenAI is comparing a chip using roughly half the rated power against Nvidia's flagship hardware.
That makes the per-watt efficiency claims the more meaningful number here, not raw throughput alone.
Kimi K2.5 was the largest model tested. On that model, OpenAI reported roughly 1.5 times higher performance per watt and 3.4 times lower latency than the comparison system.
How AI Helped Design the Chip That Runs AI
OpenAI says earlier versions of its own models helped design and bring up Jalapeño. Its newest models are now helping optimize how the chip gets programmed.
The company says AI assistance helped take the chip from initial design to tapeout in just nine months. AI also helped optimize the chip's arithmetic circuits.
One specific claim stands out. For select attention and mixture-of-experts blocks, OpenAI says AI-generated code ran faster than versions written by human engineers. The gain was 1.5 to 1.8 times.
That figure applies only to those specific blocks, not the full chip's software stack. OpenAI is clear about that limitation itself.
It's a meaningful proof of concept regardless. AI-assisted chip programming is replacing hand-written expert code, even on a limited scale.
That points toward a real shift in how future chips get built and optimized.
Why This Isn't a Break From Nvidia
Despite the competitive framing against Nvidia's chips, OpenAI isn't walking away from Nvidia at all. The company says it will keep widely deploying Nvidia accelerators for both training and inference.
Jalapeño is better understood as OpenAI adding a second option, not replacing its main supplier. OpenAI plans to begin deploying it within its own infrastructure by the end of the year. A second and third generation are already in development.
What This Means for Miami
This adds a genuinely new angle to the AI infrastructure economics MAIN has tracked all year. That includes the neocloud market and the Delta versus Vertiv cooling comparison. A major AI lab building its own chips changes the calculus for every company renting GPU capacity from someone else.
For Miami's AI infrastructure investors, the real question isn't whether Jalapeño's specific performance claims hold up under independent testing. It's whether more labs follow OpenAI toward custom silicon. That shift would directly affect neocloud providers betting their whole business on renting out someone else's Nvidia chips.