AI RadarWe read first, then explain what changed
OpenAI's chip counterattack

OpenAI used models running on Nvidia GPUs to build a chip that beat Blackwell in its first tests

Last updated 2026-08-26Editorial synthesis: signals connected before judgementNot a wire dump; facts, judgement, and unknowns are separated
Original graphic comparing OpenAI's first-generation Jalapeño inference chip with Nvidia GB200 and GB300 on efficiency and latency
Editorial graphic: in three InferenceX tests published by OpenAI, Jalapeño delivered 1.5–1.9x more work per watt than the comparison systems. It remains engineering silicon, without a complete Rubin or long-context Agent comparison.
Bottom line

On August 25, OpenAI published the first measurements for Jalapeño, its custom inference chip. Across fixed 8k/1k InferenceX tests on GPT-OSS 120B, DeepSeek R1 and Kimi K2.5, peak work per watt was 1.5–1.9x higher than Nvidia GB200/GB300, with 1.7–3.6x lower end-to-end latency. The results are strong, but this is still engineering silicon, the figures came from OpenAI, and complete Rubin, long-context and multi-turn Agent testing is still missing.

Conclusion

OpenAI has built its first custom inference chip, and the first report card is unusually strong.

In results published on August 25, Jalapeño ran GPT-OSS 120B, DeepSeek R1 and Kimi K2.5 while delivering 1.5–1.9 times more work per watt at peak than the Nvidia GB200 and GB300 comparison systems. End-to-end latency was 1.7–3.6 times lower. DeepSeek R1 reached roughly 700 tokens per second for one user at low concurrency, while GPT-OSS reached about 1,459 tokens per second.

The claim that OpenAI has already beaten every comparable Nvidia, AMD and Google chip goes too far. The public appendix directly compares GB200 and GB300. It is not a complete, matched contest against Rubin, AMD accelerators and Google TPUs. The test also uses fixed 8,000-token inputs and 1,000-token outputs, and Jalapeño remains engineering silicon.

The real news is that a model company is reaching into Nvidia's richest layer of the stack. The tool helping it do so is AI that first grew up running on Nvidia GPUs.

June showed the chip, August delivered the report card

OpenAI and Broadcom first unveiled Jalapeño in June. The announcement made one bold claim: early testing suggested substantially better performance per watt than current advanced hardware. It provided no detailed specification sheet or readable benchmark, so the market had to keep the claim on account.

Two months later, OpenAI put three public models through InferenceX. The benchmark evaluates the complete serving system, not only theoretical arithmetic on one chip. It maps throughput, power and the time a user waits for tokens.

The gaps were not small. GPT-OSS 120B reached 85,448 peak mixed tokens per kilowatt versus 44,960 for GB200. DeepSeek R1 reached 19,641 versus 11,781 for GB300. Kimi K2.5 reached 18,195 versus 11,862 for GB300.

Jalapeño has a 700W rating, while measured sustained power stayed at or below 550W on the three workloads. Data centers are often constrained less by the money for one more accelerator than by the electricity they can deliver to the room. Tokens from each megawatt eventually become revenue and API cost.

Seven hundred tokens per second is fast, but one number cannot describe production

For DeepSeek R1, minimum time between tokens fell from 5.90ms on GB300 to 1.43ms on Jalapeño. That corresponds to generation speed rising from 169 to about 700 tokens per second for one user. Kimi K2.5 reached about 694 tokens per second, while GPT-OSS reached 1,459.

Those speeds came from low concurrency, fixed sequence lengths and single-token prediction, without multi-token prediction or speculative decoding. They show that Jalapeño is very good at interactive inference. They do not prove that every user on a busy API will see 700 tokens per second throughout a day of messy production traffic.

SemiAnalysis researchers watched benchmark runs in OpenAI's lab, but documented important limits. OpenAI supplied the original numbers; the researchers did not run the entire InferenceX suite and did not see AgentX, which is closer to long-context, multi-turn Agent work. The 8k/1k result is an impressive sprint, not yet a marathon.

Running DeepSeek and Kimi makes this more than an OpenAI-only socket

Jalapeño was designed for language-model inference, but it is not locked to OpenAI models. The three tests cover GPT-OSS 120B, DeepSeek R1 670B and Kimi K2.5 1T, built by different teams with different architectures.

That does not mean any model can be dropped onto the chip at full speed. OpenAI says each new model family still needs kernels and model-specific optimization. The difference is that the hardware and programming model appear to make that work faster. SemiAnalysis even saw Doom ported to the chip through Codex prompts. Running a game has little commercial value, but it is a memorable demonstration that the device is not a black box that recognizes only ChatGPT.

The dramatic part: AI is helping build the furnace that feeds it

OpenAI says Jalapeño took nine months from formal design to tapeout. SemiAnalysis's roughly 16-month figure begins with team hiring in mid-2024. The numbers describe different starting lines rather than a contradiction.

AI did more than take meeting notes. OpenAI used models to explore circuit implementations, shorten design and verification loops, and optimize arithmetic circuits. After tapeout, the team used Codex with GPT-Astra to bring three open-weight models that were not in the original production plan to high performance within two months.

For selected GPT-OSS attention and mixture-of-experts blocks, AI-generated implementations ran 1.5–1.8 times faster than the existing human-expert versions. That figure applies to selected blocks, not the entire chip or complete model.

The image is darkly funny. OpenAI's models grew up consuming Nvidia GPUs, then helped engineers design and program an ASIC intended to reduce dependence on Nvidia.

Has Nvidia's CUDA moat been blown up

Not yet. Jalapeño shows that OpenAI can build a strong first-generation chip for selected inference workloads. It does not show that the device can replace the full range of work handled by GPUs.

GPUs train models, serve them and absorb constantly changing software and computation. Jalapeño currently focuses on inference. OpenAI itself says it will continue deploying accelerators from Nvidia and other partners at large scale. The immediate prize is choice and bargaining power, not ejecting Nvidia from the data center tomorrow.

Rubin also cannot be dismissed with one sentence. SemiAnalysis argues that some Jalapeño efficiency results exceed Nvidia's published early Rubin figures, while acknowledging that the comparison is incomplete. Rubin has started shipping to customers; Jalapeño remains engineering silicon. Software maturity, prediction techniques and test conditions are not fully aligned.

The contest has moved from model leaderboards to the other side of the power meter

OpenAI used to buy compute. It is now designing the chip, memory path, network, kernels and rack system. Jalapeño does not need to win every workload. Lowering the cost of frequent inference is enough to let ChatGPT, Codex and the API take many more steps.

That is why performance per watt matters more than one chip's theoretical peak. An Agent repeatedly calls models, reads and writes caches, and operates tools. The electricity saved on one answer is invisible. Across hundreds of millions of calls, it changes service prices, response times and gross margins.

OpenAI plans to begin deploying Jalapeño in its own infrastructure by the end of 2026. SemiAnalysis says production should ramp gradually through 2027, with most output toward year-end, and describes 100MW as the next target. Reproducing a beautiful laboratory curve across thousands of chips is the next exam.

Our judgment

Jalapeño has not defeated Nvidia's empire. It has weakened an old assumption: a first-generation custom ASIC must pay years of tuition, and a model company can never get around CUDA.

OpenAI moved from design to tapeout in nine months, produced measurable silicon after three months of bring-up, and used AI to accelerate model support. Even if only part of its traffic moves, the company gains leverage in supplier negotiations and controls more of the feedback loop between models, software and hardware.

Ordinary users will not be buying a Jalapeño graphics card soon. The practical effect could be faster or cheaper ChatGPT and Codex, or Agents that can take more steps at the same price. For Nvidia, the near-term problem is not suddenly losing ten thousand sales. It is that one of its largest customers now has a credible second path.

What to watch

  • Complete results for Rubin, AMD and Google TPU on the same public suite, models and prediction method;
  • Performance on long context, multi-turn Agents, changing cache-hit rates and high-concurrency production traffic;
  • Yield, HBM4 supply, rack reliability and progress toward 100MW during the 2027 ramp;
  • Whether OpenAI passes the cost advantage to API users or retains it mainly as margin;
  • Whether Gen 2 and Gen 3 can sustain a roughly nine-month development cadence;
  • Whether Nvidia's software ecosystem answers with faster model support and new inference methods.

Sources and evidence boundary

  • Primary source: OpenAI's performance report and appendix support the three InferenceX results, power figures, AI-assisted development and deployment plan;
  • Independent chip analysis: SemiAnalysis witnessed runs in the lab and supplied architecture details, the Doom demonstration, the 16-month program timeline, the 2027 ramp and the 100MW target;
  • “Beat Blackwell” applies to the published fixed 8k/1k inference tests, not every chip, model, workload or production system.