METAL

vLLM

인프라 · 칩미국

vLLM은 대규모 언어모델을 여러 사용자에게 동시에 서비스할 수 있도록 추론 속도와 메모리 효율을 높이는 오픈소스 추론 엔진이다. 대표 기술로는 GPU 메모리를 페이지 단위로 관리해 처리량을 끌어올리는 페이지드어텐션 기법이 꼽힌다. vLLM은 특정 기업이 판매하는 상용 제품이 아니라 커뮤니티 메인테이너들이 개발과 배포를 이끄는 오픈소스 프로젝트다. 2026년 8월 vLLM의 리드 메인테이너는 50만 GPU 규모의 오픈모델 운영 경험을 공개하며 대규모 추론 인프라 운영 사례를 알렸다.

공식 사이트 ↗

현재 순위 (1M)

23

변동
전 순위
14

최근 30일 순위 추이

2026.08.182026.09.16 · 최고 14 · 최저 23 · 현재 23

vLLM · 관련 기사

AI

큐원3.8-27B, 아파치 2.0 오픈웨이트 공개

알리바바 큐원이 27B 멀티모달 모델의 가중치를 통째로 풀었다. 4비트로 17.9GB — 그래픽카드 한 장에 들어가는 크기인데, 공식 표에서는 클로드 오퍼스 4.6 맥스를 여러 항목에서 앞선다.

By 김현국

90

vLLM · 최근 발표

  1. 608New blog! MTP, EAGLE-3, DFlash or DSpark, which speculative decoding method should you actually use?영상2026-08-28
  2. 510Qwen3.8-Flash-Next from @Alibaba_Qwen has day-0 support in vLLM, verified on NVIDIA and AMD GPUs.2026-08-26
  3. 441vLLM v0.28.0 is out! 584 commits from 270 contributors (76 new).2026-08-27
  4. 344Congrats to @SparkLLM on releasing Spark-X2.5-4B and 1.7B! Compact on-device agent models: 200+ languages, native 1M context.2026-09-01
  5. 351Congrats to @Zai_org on GLM-5.3-Flash, the first natively multimodal model in the GLM-5 series, and their first hybrid of sparse and linear attention.2026-08-26
  6. 310Congrats to @deepseek_ai on DeepSeek-V4-Flash-Vision-Exp, the first multimodal model in the V4 family! vLLM serves it now.2026-09-01
  7. 216This weight transfer engine is native to vLLM. Any Ray-based trainer can adopt it with a single WeightSource iterator.2026-08-26
  8. 183미니맥스 H3, 영상 생성이 재생시간보다 빨라졌다기사2026-09-01
  9. 196We worked alongside @inferact and @FireworksAI_HQ on the investigation.2026-08-30
  10. 189Congrats to @Zai_org on opening the GLM-5.3 weights, the largest model in the GLM-5.3 line. Day-0 support in vLLM.2026-08-28
  11. 176Huge congrats to the FastVideo team at @haoailab on the FastH3 release!영상2026-08-28
  12. 83On Tuesday, capping off a full day of speakers at the first vLLM Conference, the vLLM community gathered for a rooftop happy hour at sunset. The event hit capacity, and sunset over the city was a great backdrop.2026-08-29
  13. 76@TencentHunyuan's Hy4-preview runs in vLLM from day 0, verified on NVIDIA GPUs.2026-08-28
  14. 73vLLM core maintainer and Inferact CEO @simon_mo_ is featured in the new PyTorch Conference promo video, talking about the goal for vLLM: the easiest to use and most efficient inference engine.영상2026-08-28
  15. 61Monday afternoon we hosted a a vLLM x @googlecloud x @anyscalecompute happy hour in San Francisco , featuring core maintainers of vLLM and the builders of TPU. There were drinks, food, and a room full of people working on inference.2026-08-26
  16. 56On Monday before the vLLM Conference, we co-hosted the vLLM x NVIDIA Dynamo meetup. We had room for 300 people and got 1,600 signups. Incredible turnout from this community.2026-08-26