vLLM
vLLM은 대규모 언어모델을 여러 사용자에게 동시에 서비스할 수 있도록 추론 속도와 메모리 효율을 높이는 오픈소스 추론 엔진이다. 대표 기술로는 GPU 메모리를 페이지 단위로 관리해 처리량을 끌어올리는 페이지드어텐션 기법이 꼽힌다. vLLM은 특정 기업이 판매하는 상용 제품이 아니라 커뮤니티 메인테이너들이 개발과 배포를 이끄는 오픈소스 프로젝트다. 2026년 8월 vLLM의 리드 메인테이너는 50만 GPU 규모의 오픈모델 운영 경험을 공개하며 대규모 추론 인프라 운영 사례를 알렸다.
현재 순위 (1M)
23위
- 변동
- 전 순위
- 14
최근 30일 순위 추이
2026.08.18 – 2026.09.16 · 최고 14위 · 최저 23위 · 현재 23위
vLLM · 관련 기사
마지막 기사입니다.
vLLM · 최근 발표
- 608New blog! MTP, EAGLE-3, DFlash or DSpark, which speculative decoding method should you actually use?
- 510Qwen3.8-Flash-Next from @Alibaba_Qwen has day-0 support in vLLM, verified on NVIDIA and AMD GPUs.
- 441vLLM v0.28.0 is out! 584 commits from 270 contributors (76 new).
- 344Congrats to @SparkLLM on releasing Spark-X2.5-4B and 1.7B! Compact on-device agent models: 200+ languages, native 1M context.
- 351Congrats to @Zai_org on GLM-5.3-Flash, the first natively multimodal model in the GLM-5 series, and their first hybrid of sparse and linear attention.
- 310Congrats to @deepseek_ai on DeepSeek-V4-Flash-Vision-Exp, the first multimodal model in the V4 family! vLLM serves it now.
- 216This weight transfer engine is native to vLLM. Any Ray-based trainer can adopt it with a single WeightSource iterator.
- 183미니맥스 H3, 영상 생성이 재생시간보다 빨라졌다
- 196We worked alongside @inferact and @FireworksAI_HQ on the investigation.
- 189Congrats to @Zai_org on opening the GLM-5.3 weights, the largest model in the GLM-5.3 line. Day-0 support in vLLM.
- 176Huge congrats to the FastVideo team at @haoailab on the FastH3 release!
- 83On Tuesday, capping off a full day of speakers at the first vLLM Conference, the vLLM community gathered for a rooftop happy hour at sunset. The event hit capacity, and sunset over the city was a great backdrop.
- 76@TencentHunyuan's Hy4-preview runs in vLLM from day 0, verified on NVIDIA GPUs.
- 73vLLM core maintainer and Inferact CEO @simon_mo_ is featured in the new PyTorch Conference promo video, talking about the goal for vLLM: the easiest to use and most efficient inference engine.
- 61Monday afternoon we hosted a a vLLM x @googlecloud x @anyscalecompute happy hour in San Francisco , featuring core maintainers of vLLM and the builders of TPU. There were drinks, food, and a room full of people working on inference.
- 56On Monday before the vLLM Conference, we co-hosted the vLLM x NVIDIA Dynamo meetup. We had room for 300 people and got 1,600 signups. Incredible turnout from this community.


