vLLM is the leading open-source LLM inference server, offering 24× higher throughput than HuggingFace Transformers through PagedAttention. This guide covers installation, serving a model, and configuring for production workloads.
HuggingFace Transformers is the standard for model loading and fine-tuning, but its inference is slow under load because each request allocates KV cache memory statically. vLLM's PagedAttention manages the KV cache like virtual memory — dynamically allocating pages across requests. This allows continuous batching of hundreds of concurrent requests on a single GPU.
For a 7B model on an A100, vLLM achieves around 3000–4000 tokens/second throughput — 10–20× more than a naive HuggingFace setup. For APIs serving multiple users, this dramatically reduces cost per query.
vLLM requires an NVIDIA GPU with CUDA 12.1+. Install via pip: `pip install vllm`. The package is large (~2 GB) as it includes compiled CUDA kernels. For the latest features, install from source: `pip install git+https://github.com/vllm-project/vllm.git`.
On AMD GPUs (MI250/MI300), use the ROCm build: `pip install vllm --extra-index-url https://download.pytorch.org/whl/rocm6.1`. Apple Silicon is not supported — use Ollama or llama.cpp for local Mac inference.
Start the server: `vllm serve meta-llama/Llama-3.3-70B-Instruct --tensor-parallel-size 2 --gpu-memory-utilization 0.9 --port 8000`. The `--tensor-parallel-size 2` flag splits the model across 2 GPUs. `--gpu-memory-utilization 0.9` allows vLLM to use 90% of VRAM for the KV cache.
The server exposes an OpenAI-compatible API at `http://localhost:8000/v1`. Any OpenAI SDK client works without modification. The `/v1/models` endpoint lists the available model. The model name in requests should match the HuggingFace model ID.
For maximum throughput (batch processing): increase `--max-num-seqs 256` (max concurrent sequences) and enable continuous batching (default). Set `--max-num-batched-tokens 16384` to allow larger batches. Throughput improves linearly with batch size up to GPU saturation.
For minimum latency (interactive chat): use a smaller `--max-num-seqs 8`, enable speculative decoding with `--speculative-model <small-draft-model>`, and use `--enforce-eager` to skip CUDA graph capture for faster startup. Speculative decoding reduces latency by 2–3× for chat-length generations.
For a 70B model on a single 80GB GPU, use AWQ quantisation: `vllm serve meta-llama/Meta-Llama-3-70B-Instruct-AWQ --quantization awq`. AWQ (Activation-aware Weight Quantisation) reduces model size to ~4 bits with minimal quality loss and is faster than GPTQ.
For models larger than your VRAM, use pipeline parallelism across multiple nodes: `--pipeline-parallel-size 2` splits the model's layers across 2 nodes. Combined with tensor parallelism (`--tensor-parallel-size 4` per node), you can serve 405B+ models across a GPU cluster.
Ollama is the fastest way to run open-source LLMs locally. This guide walks you through installation on macOS, Linux, and Windows, pulling your first model, and running it from the terminal or via its REST API.
Read guideLM Studio is a desktop app for running local AI models without any terminal commands. This guide covers installation, downloading a model from Hugging Face, and starting a local inference server.
Read guide