vllm-project/vllmA high-throughput and memory-efficient inference and serving engine for LLMs
A high-throughput and memory-efficient inference and serving engine for LLMs
licence Apache-2.0since 2023branch mainvllm.ai ↗github.com/vllm-project/vllm ↗
Outside contributors get merged here, though some wait a while for a reply.
16 of 67outside PRs merged (24%)
checked
02What happened to outsiders
03Where newcomer work lands
04The evidence