vLLM

Appears in 2 tutorials

A high-performance GPU serving engine with continuous batching and LoRA support; the default for serious serving.

As used in Fine-Tuning & Model Customization →

A high-performance GPU serving engine with continuous batching and LoRA support; the default for serious serving. (M14)

As used in LLM Infrastructure →

The most popular open-source LLM inference/serving engine; introduced PagedAttention and bundles the whole optimization toolbox plus an OpenAI-compatible server.