vllm-project/vllm

A high-throughput and memory-efficient inference and serving engine for LLMs

View on GitHub
Python
Stars 87.5k
Forks 20.0k
License Apache-2.0
Open Issues 6119
Updated 1d ago