AI news and explainers at 7 AM weekdays, plus a Sunday weekly at 8Get it in your inbox

METAL LAB

NVIDIA pairs Vera Rubin with Groq 3 LPX, quadrupling token speed

Agent-focused inference infrastructure already adopted by SpaceXAI, CoreWeave, and Nebius unveiled at Hot Chips

검은 배경 위에 서버 랙 다섯 대가 나란히 놓여 있다

이미지: NVIDIA

Summary

  • NVIDIA announced that it has entered full-scale production combining its Vera Rubin NVL72 rack-scale system with the Groq 3 LPX inference accelerator.
  • In benchmarks from Artificial Analysis, running the open-source agentic model Gemma 4 31B with a 100,000-token context produced 3,400 tokens per second — four times faster than the closest competing platform.
  • The announcement also confirmed that SpaceXAI is the first to adopt the Vera CPU, CoreWeave the first to deploy Spectrum-X Multiplane networking, and Nebius the first to bring in Groq 3 LPX.
발표 계기
미국 팔로알토 Hot Chips 콘퍼런스, 2026년 8월 24일
시스템
NVIDIA 베라 루빈 NVL72 랙스케일 시스템 + Groq 3 LPX, 완전 양산(full production) 단계
벤치마크
아티피셜 애널리시스 측정, Gemma 4 31B 기준 10만 토큰 문맥에서 초당 3,400 출력 토큰 (대안 대비 4배)
파트너
SpaceXAI(베라 CPU 채택), CoreWeave(Spectrum-X Multiplane 상용 배포), Nebius(Groq 3 LPX 첫 채택 클라우드)
네트워크 확장성
Spectrum-X Multiplane, 3계층 추가 없이 최대 GPU 51만2,000개까지 확장
네트워크 복원력
8플레인 구성 기준 한 플레인 장애 시 대역폭 약 90% 유지, 하드웨어 복구가 소프트웨어 방식보다 11배 빠름
스위치 사양
Spectrum-X SN6000 시리즈, 102.4Tb/s Spectrum-6 이더넷 ASIC + ConnectX-9 SuperNIC, GPU당 최대 1,600Gb/s
신규 네트워킹 축
다섯 번째 축 'Scale-In', NVIDIA BlueField-4 프로세서·DOCA 소프트웨어 기반

The time it takes an agent to answer a single question is really just the sum of every token generated along the way. As that process stretches across tool calls and hand-offs between multiple agents, the delays stack up — and the lag a person actually feels can balloon fast. NVIDIA's answer, unveiled this week at the Hot Chips conference in Palo Alto, is a new inference accelerator. On August 24, the company said it has entered full production combining the Vera Rubin NVL72 rack-scale system with Groq 3 LPX.

Three nodes connect in sequence. A thickly filled circle labeled "Rubin GPU" represents holding a long context in heavy memory, and a solid line passes that context to the next node, a satellite-orbit-shaped "LPX." LPX handles the repetitive job of generating tokens one after another, feeding into a node shown as expanding dots labeled "4x speed" — representing the result of 3,400 tokens per second over a 100,000-token context.

In the agent era, token generation speed is the bottleneck

As AI shifts from the training phase to the inference phase, the job infrastructure has to do is changing too. Agents don't stop at a single response — they keep generating tokens as they reason, call tools, and exchange information with other AI systems. Context windows, the amount an AI can hold in memory and read at once, keep growing as well. NVIDIA argues this calls for a new kind of infrastructure, distinct from training infrastructure, that can simultaneously deliver throughput, responsiveness, and cost efficiency.

What NVIDIA is emphasizing here isn't optimizing individual components in isolation — it's what the company calls "extreme co-design." That means designing Vera Rubin's compute, Spectrum-X Ethernet networking, and Groq 3 LPX's inference acceleration together as one unified system. The Rubin GPU handles large-scale context processing, while LPX accelerates the latency-sensitive work of decoding — that is, generating tokens.

3,400 tokens per second at a 100,000-token context

The actual performance numbers come from Artificial Analysis, a firm that independently measures model speed and pricing. Running the open-source agentic model Gemma 4 31B in a long-context setting of 100,000 tokens produced 3,400 output tokens per second. That's a particularly meaningful figure for agentic systems, and NVIDIA says it's four times faster than the closest alternative platform.

At rack scale, a Groq 3 LPX deployment can link up to 256 LPU accelerators through direct chip-to-chip connections. NVIDIA says that cluster of LPUs then behaves like a single massive processor optimized for deterministic inference — meaning it delivers the same latency for the same input every time.

SpaceXAI, CoreWeave, and Nebius are already on board

Three companies that have already deployed the platform were named alongside the announcement.

PartnerDeployment
SpaceXAIPlans to deploy Vera CPUs across next-generation agentic AI architecture, spanning ground-based data centers to orbital satellites
CoreWeaveDeployed Spectrum-X Multiplane, which connects Vera Rubin racks, into production
NebiusFirst AI cloud to adopt Groq 3 LPX, integrating it into its Token Factory

SpaceXAI said it plans to use Vera CPUs to accelerate CPU-heavy agentic workloads like orchestration, tool use, code execution, data processing, and simulation. CoreWeave has already put Spectrum-X Multiplane into production — a system that links multiple switches in parallel to build a high-bandwidth, lossless network. Nebius said it's adding Groq 3 LPX to its Token Factory inference service to boost responsiveness for real-time applications like coding agents.

The network is changing too — Spectrum-X Multiplane

As AI factories scale up, the bottleneck tends to hit the network before it hits the chips. Scaling a cluster the conventional way means adding a third network tier, which increases latency and drives up costs for cabling, optical transceivers, and power. Spectrum-X Multiplane sidesteps this by splitting a single server's network connections into multiple independent "planes," each operating as its own two-tier network. NVIDIA says this makes it possible to build a flat network reaching up to 512,000 GPUs without adding that third tier.

Distribution and failure handling are managed automatically by a dedicated hardware engine inside the ConnectX SuperNIC. In an eight-plane configuration, NVIDIA says that if one plane goes down, roughly 90% of total bandwidth is preserved, and recovery is 11 times faster than software-based approaches. Translated into overall AI factory throughput, that works out to a 1.6x increase in output. The hardware behind this includes the SN6000 series of switches, built on 102.4Tb/s Spectrum-6 Ethernet ASICs, and ConnectX-9 SuperNICs supporting up to 1,600Gb/s per GPU. Add Spectrum-XGS Ethernet, which links multiple data centers together as a single "super factory," and NVIDIA says multi-site collaborative computation (NCCL collectives) speeds up by 1.9x.

A fifth axis: Scale-In

NVIDIA also introduced a new axis of networking at the event, called Scale-In, which runs on Spectrum-X Ethernet atop the BlueField-4 processor and DOCA software platform. The company describes it as folding what used to be treated as north-south access networking — traffic moving in and out of a data center — into a single accelerated infrastructure domain.

NVIDIA's reasoning is that as agents constantly interact with data, storage, and services, security, operations, and storage now need to be accelerated right alongside compute. Scale-In is designed to provide multi-tenant networking, high-performance storage access, silicon-level security, elastic provisioning, and real-time observability, while keeping infrastructure processing separate from host computing resources.

Editor's take

Boil this announcement down to one line, and it's NVIDIA in the middle of a transition — from a company that sells training GPUs to one that sells the entire inference pipeline. Rolling out Groq 3 LPX, Spectrum-X Multiplane, and Scale-In together at the same conference on the same day isn't a coincidence. In the training era, all that mattered was GPU count and bandwidth. Agentic inference, though, is a problem that hits latency, context length, and network fault tolerance all at once — which means NVIDIA has essentially concluded it has nothing to sell unless it co-designs the entire infrastructure stack itself.

There was a time when scaling GPU clusters was synonymous with scaling performance. This announcement suggests that equation no longer holds. The 3,400 tokens-per-second figure itself matters less than the condition under which it was achieved: a 100,000-token context. Being fast on a short prompt is easy — the real bottleneck shows up once conversations run long and tool calls start stacking on top of each other. Teams building agentic services domestically shouldn't take comfort in benchmarks run only on short contexts. The first step is measuring latency at the actual context lengths your service will use in production.

The networking numbers are also worth sitting with from a practical standpoint. Preserving 90% of bandwidth even when one plane fails means you don't need to buy separate redundant infrastructure just to handle failures. For Korean startups that rent GPUs through cloud providers rather than running their own infrastructure, it's worth asking — before signing a contract — whether this kind of resilience actually shows up in pricing or uptime guarantees.

Expect more clouds to start offering Groq 3 LPX in the coming weeks. NVIDIA itself has framed this announcement as "just the beginning," with more models and optimizations to follow — so these benchmark numbers will likely be superseded again by the next generation.

Comments