vLLM
A high-performance GPU serving engine with continuous batching and LoRA support; the default for serious serving.
A high-performance GPU serving engine with continuous batching and LoRA support; the default for serious serving. (M14)
The most popular open-source LLM inference/serving engine; introduced PagedAttention and bundles the whole optimization toolbox plus an OpenAI-compatible server.