AI news and explainers at 7 AM weekdays, plus a Sunday weekly at 8Get it in your inbox

METAL LAB

OpenAI's First Inference Chip Jalapeño Outpaced Blackwell

First benchmarks unveiled at Hot Chips show throughput per watt up to 1.9x that of the best existing systems

이미지: METAL LAB 생성

Summary

  • OpenAI unveiled the first benchmark results for its own inference chip, Jalapeño, at the Hot Chips conference.
  • In tests using SemiAnalysis's InferenceX benchmark, throughput per watt came in up to 1.9 times higher than NVIDIA's Blackwell systems.
  • The chip, built with Broadcom, is set for limited deployment in late 2026 with a full-scale rollout in 2027.
발표 계기
미국 팔로알토 Hot Chips 콘퍼런스, 2026년 8월 25일(화)
칩 이름
Jalapeño, 오픈AI 첫 자체 추론 전용 칩
개발 파트너
브로드컴과 공동 개발, 2024년 중반 설계 시작·2025년 11월 양산 착수
테스트 모델
GPT-OSS 120B·DeepSeek R1 670B·Kimi K2.5 1T
핵심 성능
와트당 AI 작업량 1.5~1.9배, 종단 지연 1.7~3.6배 감소, 대화형 워크로드 2.1~4.1배
GPT-OSS 처리량
85,448 tokens/sec/kW (기존 최고 44,960 대비 1.9배)
배포 일정
2026년 말 소량, 2027년 본격 확대 (오픈AI 하드웨어 총괄 리처드 호)

OpenAI has released the first performance figures for Jalapeño, its in-house inference-only chip. The numbers came out of the Hot Chips conference in Palo Alto, and according to measurements using SemiAnalysis's public InferenceX benchmark, throughput per watt came in up to 1.9 times higher than the best systems currently available.

Three nodes sit side by side. A solid arrow runs from Jalapeño to Blackwell labeled "1.9x ahead," with Blackwell shown as a thick filled circle representing the weight of the incumbent leader. A dotted arrow runs from Jalapeño to Rubin, orbiting on a dashed path, labeled "not yet tested." The point: the rival it just beat isn't the rival that actually matters.

The Chip First Announced Back in June

Jalapeño's name first surfaced back in June. What's new now is the first round of actual test results since then. On a press call, OpenAI's head of hardware, Richard Ho, said the chip handles more AI workload per unit of power while also responding faster. But it's worth noting that the comparison here is against Blackwell, NVIDIA's previous-generation system, not the newer Rubin. By the time Jalapeño actually ships at scale, the competition will likely have moved well past where it stands today.

The Numbers, Across Three Models

The InferenceX benchmark released by SemiAnalysis was run across three models: GPT-OSS 120B, DeepSeek R1 670B, and Kimi K2.5 1T. OpenAI supplied the figures, and SemiAnalysis reportedly verified some of the runs on-site.

Model (Parameters)JalapeñoBest Existing SystemMultiplier
GPT-OSS · 120B85,44844,9601.9× 100
DeepSeek R1 · 670B19,64111,7811.7× 89
Kimi K2.5 · 1T18,19511,8621.5× 79

The unit here is mixed tokens per second divided by kilowatts. In more concrete terms, OpenAI's announcement says GPT-OSS delivered roughly 1,400 tokens per second per user, while DeepSeek R1 produced more than 700 tokens per second on a single concurrent request. What stands out is that Jalapeño hit these numbers without using acceleration techniques like multi-token prediction or speculative decoding — techniques some of the comparison systems were already using. That suggests there's still room for improvement.

A First-Generation Chip, 16 Months in the Making

Jalapeño was built by OpenAI in partnership with Broadcom. Design work began in mid-2024, the final design moved to the production line in November 2025, and it reportedly took nine months from the first chip design to finished blueprints heading to the fab. OpenAI said it used its own models throughout the process — older-generation models helped with the chip design itself, while its newest models handled programming and optimization.

In a tweet from SemiAnalysis CEO Dylan Patel, he wrote that first-generation chips are usually uncompetitive, but OpenAI managed to beat both NVIDIA's Blackwell and Rubin. SemiAnalysis found that even against NVIDIA's Vera Rubin, which uses the same memory spec, Jalapeño came out ahead on output tokens per megawatt. That said, when measured by total cost of ownership per token, the two systems came out roughly even.

A Rival That's Still a Customer

Jalapeño isn't a chip tuned specifically for OpenAI's own models — it's a general-purpose LLM inference accelerator. It can't handle training, only inference. OpenAI CFO Sarah Friar said the chip is part of OpenAI's stated compute strategy, which ties together data centers, chips, models, developer platforms, products, and devices into a single system. The company's position is that this complements existing partnerships with NVIDIA, AMD, AWS, Cerebras, and CoreWeave rather than replacing them. Since NVIDIA, AMD, and AWS are also investors and compute partners of OpenAI, the announcement effectively captures a relationship that's simultaneously cooperative and competitive.

Still at the Engineering-Sample Stage

NVIDIA and AMD have already published results on larger models like DeepSeek V4 Pro and Kimi K3, but Jalapeño hasn't been tested on those yet. Rubin systems are already shipping to customers, while Jalapeño is reportedly still at the engineering-sample stage. Richard Ho said limited deployment will begin later this year, with a full scale-up planned for 2027.

Editor's Take

OpenAI announced its plan to build its own chips last year, but this is the first time it's actually put real numbers on the table. And the fact that those numbers are measured against a rival's previous generation, not its latest, is exactly what puts this announcement in proper perspective. A first-generation chip beating a shipping commercial system is genuinely significant. But the real contest will be decided by where NVIDIA's Rubin stands in 2027, when Jalapeño is set to actually deploy at scale.

What's more interesting here than the benchmark numbers is the development process itself. OpenAI says it used older models to design the chip and newer models to optimize that design — and if that's accurate, it means the company is accumulating chip-design know-how through its own AI, without relying on external toolchains. That's also what SemiAnalysis was getting at when it wrote that "NVIDIA's CUDA moat may be dead." The real benchmark for future chip competition may come down to how quickly a new model can be made to run on a new chip.

For companies watching from outside this race, the takeaway is straightforward. There's no way to actually rent Jalapeño right now. But for service companies where inference costs eat up a meaningful share of revenue, this should read as a signal: starting around 2027, the cost gap could start widening between big tech firms with their own chips and everyone else still dependent on NVIDIA. For now, the practical move is simply to keep track of which cloud-and-model combinations manage cost per token best.

In the coming months, a second round of benchmarks — likely pitting Jalapeño against NVIDIA's Rubin on larger models like DeepSeek V4 Pro or Kimi K3 — is a strong possibility. Those results will be what actually confirms what this first-generation report card really means.

Comments