METAL

vLLM

基础设施 · 芯片美国

vLLM是一款开源推理引擎,旨在提升推理速度和内存效率,使大语言模型能够同时为众多用户提供服务。其代表性技术是PagedAttention(分页注意力),该技术以页为单位管理GPU内存,从而提升吞吐量。vLLM并非某家企业销售的商业产品,而是由社区维护者主导开发与发布的开源项目。2026年8月,vLLM的首席维护者公开了在50万块GPU规模下运行开源模型的经验,分享了大规模推理基础设施的运营案例。

官方网站 ↗

当前排名 (1M)

25

变化
上期
17

最近 30 天排名走势

2026.08.202026.09.18 · 最高 16 · 最低 25 · 当前 25

vLLM · 相关文章

AI

Qwen3.8-27B,以Apache 2.0协议开放权重

阿里巴巴通义千问将27B多模态模型的权重完整开放。4比特量化后仅17.9GB——一张显卡就能装下,而在官方对比表中,该模型在多个项目上超过了Claude Opus 4.6 Max。

作者 김현국

110

vLLM · 最新发布

  1. 344Congrats to @SparkLLM on releasing Spark-X2.5-4B and 1.7B! Compact on-device agent models: 200+ languages, native 1M context.2026-09-01
  2. 310Congrats to @deepseek_ai on DeepSeek-V4-Flash-Vision-Exp, the first multimodal model in the V4 family! vLLM serves it now.2026-09-01
  3. 183미니맥스 H3, 영상 생성이 재생시간보다 빨라졌다기사2026-09-01
  4. 196We worked alongside @inferact and @FireworksAI_HQ on the investigation.2026-08-30
  5. 189Congrats to @Zai_org on opening the GLM-5.3 weights, the largest model in the GLM-5.3 line. Day-0 support in vLLM.2026-08-28
  6. 176Huge congrats to the FastVideo team at @haoailab on the FastH3 release!영상2026-08-28
  7. 83On Tuesday, capping off a full day of speakers at the first vLLM Conference, the vLLM community gathered for a rooftop happy hour at sunset. The event hit capacity, and sunset over the city was a great backdrop.2026-08-29
  8. 73vLLM core maintainer and Inferact CEO @simon_mo_ is featured in the new PyTorch Conference promo video, talking about the goal for vLLM: the easiest to use and most efficient inference engine.영상2026-08-28