vllm-project/vllm

A high-throughput and memory-efficient inference and serving engine for LLMs

View on GitHub
Python
Stars 91.9k
Forks 22.3k
License Apache-2.0
Open Issues 8055
Updated 1d ago