vLLM
vLLM是一款开源推理引擎,旨在提升推理速度和内存效率,使大语言模型能够同时为众多用户提供服务。其代表性技术是PagedAttention(分页注意力),该技术以页为单位管理GPU内存,从而提升吞吐量。vLLM并非某家企业销售的商业产品,而是由社区维护者主导开发与发布的开源项目。2026年8月,vLLM的首席维护者公开了在50万块GPU规模下运行开源模型的经验,分享了大规模推理基础设施的运营案例。
当前排名 (1M)
25位
- 变化
- 上期
- 17
最近 30 天排名走势
2026.08.20 – 2026.09.18 · 最高 16位 · 最低 25位 · 当前 25位
vLLM · 相关文章
美国利率上升动摇AI热潮,AMD发布新款EPYC、GPT-5.6降价
美国国债收益率急剧上升,威胁到亚洲AI股价的上涨势头,与此同时,AMD发布新一代服务器芯片,GPT-5.6大幅降价,开源模型与基础设施的讨论也贯穿全天。
10
Qwen3.8-27B,以Apache 2.0协议开放权重
阿里巴巴通义千问将27B多模态模型的权重完整开放。4比特量化后仅17.9GB——一张显卡就能装下,而在官方对比表中,该模型在多个项目上超过了Claude Opus 4.6 Max。
110
vLLM · 最新发布
- 344Congrats to @SparkLLM on releasing Spark-X2.5-4B and 1.7B! Compact on-device agent models: 200+ languages, native 1M context.
- 310Congrats to @deepseek_ai on DeepSeek-V4-Flash-Vision-Exp, the first multimodal model in the V4 family! vLLM serves it now.
- 183미니맥스 H3, 영상 생성이 재생시간보다 빨라졌다
- 196We worked alongside @inferact and @FireworksAI_HQ on the investigation.
- 189Congrats to @Zai_org on opening the GLM-5.3 weights, the largest model in the GLM-5.3 line. Day-0 support in vLLM.
- 176Huge congrats to the FastVideo team at @haoailab on the FastH3 release!
- 83On Tuesday, capping off a full day of speakers at the first vLLM Conference, the vLLM community gathered for a rooftop happy hour at sunset. The event hit capacity, and sunset over the city was a great backdrop.
- 73vLLM core maintainer and Inferact CEO @simon_mo_ is featured in the new PyTorch Conference promo video, talking about the goal for vLLM: the easiest to use and most efficient inference engine.












