METAL

vLLM

Infrastructure · chipsUnited States

vLLM is an open-source inference engine that improves inference speed and memory efficiency so that large language models can serve many users simultaneously. Its signature technology is PagedAttention, a technique that manages GPU memory in page units to boost throughput. Rather than being a commercial product sold by a specific company, vLLM is an open-source project whose development and distribution are led by community maintainers. In August 2026, vLLM's lead maintainer shared operational experience running open models at a scale of 500,000 GPUs, offering a glimpse into large-scale inference infrastructure operations.

Official site ↗

Current rank (1M)

25

Change
Prev.
17

Rank over the last 30 days

2026.08.202026.09.18 · High 16 · Low 25 · Now 25

vLLM · Related stories

AI

Qwen3.8-27B released as open weights under Apache 2.0

Alibaba's Qwen has opened up the full weights of its 27B multimodal model. At 4-bit precision it's 17.9GB — small enough to fit on a single graphics card — yet the official chart shows it beating Claude Opus 4.6 Max on several metrics.

By 김현국

110

vLLM · Recent announcements

  1. 344Congrats to @SparkLLM on releasing Spark-X2.5-4B and 1.7B! Compact on-device agent models: 200+ languages, native 1M context.2026-09-01
  2. 310Congrats to @deepseek_ai on DeepSeek-V4-Flash-Vision-Exp, the first multimodal model in the V4 family! vLLM serves it now.2026-09-01
  3. 183미니맥스 H3, 영상 생성이 재생시간보다 빨라졌다기사2026-09-01
  4. 196We worked alongside @inferact and @FireworksAI_HQ on the investigation.2026-08-30
  5. 189Congrats to @Zai_org on opening the GLM-5.3 weights, the largest model in the GLM-5.3 line. Day-0 support in vLLM.2026-08-28
  6. 176Huge congrats to the FastVideo team at @haoailab on the FastH3 release!영상2026-08-28
  7. 83On Tuesday, capping off a full day of speakers at the first vLLM Conference, the vLLM community gathered for a rooftop happy hour at sunset. The event hit capacity, and sunset over the city was a great backdrop.2026-08-29
  8. 73vLLM core maintainer and Inferact CEO @simon_mo_ is featured in the new PyTorch Conference promo video, talking about the goal for vLLM: the easiest to use and most efficient inference engine.영상2026-08-28