open-source · since 2023
vLLM
vLLM community
A high-throughput, memory-efficient inference and serving engine for large language models, with an OpenAI-compatible API and deployment patterns for single-node and distributed model serving.
AIinferencemodel servingLLM
desk notes
Verified 2026-08-31 against the project site, GitHub repository metadata and releases. Repository is active; latest GitHub release is v0.28.0.