METAL

vLLM

基础设施 · 芯片美国

vLLM是一款开源推理引擎,旨在提升推理速度和内存效率,使大语言模型能够同时为众多用户提供服务。其代表性技术是PagedAttention(分页注意力),该技术以页为单位管理GPU内存,从而提升吞吐量。vLLM并非某家企业销售的商业产品,而是由社区维护者主导开发与发布的开源项目。2026年8月,vLLM的首席维护者公开了在50万块GPU规模下运行开源模型的经验,分享了大规模推理基础设施的运营案例。

官方网站 ↗

当前排名 (1M)

23

变化
上期
14

最近 30 天排名走势

2026.08.182026.09.16 · 最高 14 · 最低 23 · 当前 23

vLLM · 相关文章

已经是最后一篇了。

vLLM · 最新发布

  1. 608New blog! MTP, EAGLE-3, DFlash or DSpark, which speculative decoding method should you actually use?영상2026-08-28
  2. 510Qwen3.8-Flash-Next from @Alibaba_Qwen has day-0 support in vLLM, verified on NVIDIA and AMD GPUs.2026-08-26
  3. 441vLLM v0.28.0 is out! 584 commits from 270 contributors (76 new).2026-08-27
  4. 344Congrats to @SparkLLM on releasing Spark-X2.5-4B and 1.7B! Compact on-device agent models: 200+ languages, native 1M context.2026-09-01
  5. 351Congrats to @Zai_org on GLM-5.3-Flash, the first natively multimodal model in the GLM-5 series, and their first hybrid of sparse and linear attention.2026-08-26
  6. 310Congrats to @deepseek_ai on DeepSeek-V4-Flash-Vision-Exp, the first multimodal model in the V4 family! vLLM serves it now.2026-09-01
  7. 216This weight transfer engine is native to vLLM. Any Ray-based trainer can adopt it with a single WeightSource iterator.2026-08-26
  8. 183미니맥스 H3, 영상 생성이 재생시간보다 빨라졌다기사2026-09-01
  9. 196We worked alongside @inferact and @FireworksAI_HQ on the investigation.2026-08-30
  10. 189Congrats to @Zai_org on opening the GLM-5.3 weights, the largest model in the GLM-5.3 line. Day-0 support in vLLM.2026-08-28
  11. 176Huge congrats to the FastVideo team at @haoailab on the FastH3 release!영상2026-08-28
  12. 83On Tuesday, capping off a full day of speakers at the first vLLM Conference, the vLLM community gathered for a rooftop happy hour at sunset. The event hit capacity, and sunset over the city was a great backdrop.2026-08-29
  13. 76@TencentHunyuan's Hy4-preview runs in vLLM from day 0, verified on NVIDIA GPUs.2026-08-28
  14. 73vLLM core maintainer and Inferact CEO @simon_mo_ is featured in the new PyTorch Conference promo video, talking about the goal for vLLM: the easiest to use and most efficient inference engine.영상2026-08-28
  15. 61Monday afternoon we hosted a a vLLM x @googlecloud x @anyscalecompute happy hour in San Francisco , featuring core maintainers of vLLM and the builders of TPU. There were drinks, food, and a room full of people working on inference.2026-08-26
  16. 56On Monday before the vLLM Conference, we co-hosted the vLLM x NVIDIA Dynamo meetup. We had room for 300 people and got 1,600 signups. Incredible turnout from this community.2026-08-26