vLLM
vLLM是一款开源推理引擎,旨在提升推理速度和内存效率,使大语言模型能够同时为众多用户提供服务。其代表性技术是PagedAttention(分页注意力),该技术以页为单位管理GPU内存,从而提升吞吐量。vLLM并非某家企业销售的商业产品,而是由社区维护者主导开发与发布的开源项目。2026年8月,vLLM的首席维护者公开了在50万块GPU规模下运行开源模型的经验,分享了大规模推理基础设施的运营案例。
当前排名 (1M)
23位
- 变化
- 上期
- 14
最近 30 天排名走势
2026.08.18 – 2026.09.16 · 最高 14位 · 最低 23位 · 当前 23位
vLLM · 相关文章
已经是最后一篇了。
vLLM · 最新发布
- 608New blog! MTP, EAGLE-3, DFlash or DSpark, which speculative decoding method should you actually use?
- 510Qwen3.8-Flash-Next from @Alibaba_Qwen has day-0 support in vLLM, verified on NVIDIA and AMD GPUs.
- 441vLLM v0.28.0 is out! 584 commits from 270 contributors (76 new).
- 344Congrats to @SparkLLM on releasing Spark-X2.5-4B and 1.7B! Compact on-device agent models: 200+ languages, native 1M context.
- 351Congrats to @Zai_org on GLM-5.3-Flash, the first natively multimodal model in the GLM-5 series, and their first hybrid of sparse and linear attention.
- 310Congrats to @deepseek_ai on DeepSeek-V4-Flash-Vision-Exp, the first multimodal model in the V4 family! vLLM serves it now.
- 216This weight transfer engine is native to vLLM. Any Ray-based trainer can adopt it with a single WeightSource iterator.
- 183미니맥스 H3, 영상 생성이 재생시간보다 빨라졌다
- 196We worked alongside @inferact and @FireworksAI_HQ on the investigation.
- 189Congrats to @Zai_org on opening the GLM-5.3 weights, the largest model in the GLM-5.3 line. Day-0 support in vLLM.
- 176Huge congrats to the FastVideo team at @haoailab on the FastH3 release!
- 83On Tuesday, capping off a full day of speakers at the first vLLM Conference, the vLLM community gathered for a rooftop happy hour at sunset. The event hit capacity, and sunset over the city was a great backdrop.
- 76@TencentHunyuan's Hy4-preview runs in vLLM from day 0, verified on NVIDIA GPUs.
- 73vLLM core maintainer and Inferact CEO @simon_mo_ is featured in the new PyTorch Conference promo video, talking about the goal for vLLM: the easiest to use and most efficient inference engine.
- 61Monday afternoon we hosted a a vLLM x @googlecloud x @anyscalecompute happy hour in San Francisco , featuring core maintainers of vLLM and the builders of TPU. There were drinks, food, and a room full of people working on inference.
- 56On Monday before the vLLM Conference, we co-hosted the vLLM x NVIDIA Dynamo meetup. We had room for 300 people and got 1,600 signups. Incredible turnout from this community.


