vLLM
Infrastructure · chips·United States
vLLM is an open-source inference engine that improves inference speed and memory efficiency so that large language models can serve many users simultaneously. Its signature technology is PagedAttention, a technique that manages GPU memory in page units to boost throughput. Rather than being a commercial product sold by a specific company, vLLM is an open-source project whose development and distribution are led by community maintainers. In August 2026, vLLM's lead maintainer shared operational experience running open models at a scale of 500,000 GPUs, offering a glimpse into large-scale inference infrastructure operations.
Official site ↗