vLLM Quickstart: High-Performance LLM Serving - in 2026

Install, serve and tune vLLM in 2026

Page content

vLLM is a high-throughput, memory-efficient inference and serving engine for Large Language Models (LLMs) developed by UC Berkeley’s Sky Computing Lab.

Its PagedAttention algorithm treats the KV cache like OS virtual memory pages and pairs it with continuous batching — the combination that made vLLM the default engine for production LLM serving. To see how vLLM fits among Ollama, Docker Model Runner, LocalAI and cloud providers—including cost and infrastructure trade-offs—see LLM Hosting: Local, Self-Hosted & Cloud Infrastructure Compared.

vllm logo

This guide is current as of vLLM v0.31.0 (October 2026). Two things changed recently if you are arriving from older tutorials: python -m vllm.entrypoints.openai.api_server is deprecated and prints a warning — vllm serve is the standard CLI now — and prebuilt wheels target CUDA 12.9 by default, not 11.8.

What is vLLM?

vLLM (virtual LLM) is an open-source library for fast LLM inference and serving that has quickly become the industry standard for production deployments. Released in 2023, it introduced PagedAttention, a groundbreaking memory management technique that dramatically improves serving efficiency.

Key Features

High Throughput Performance: vLLM’s launch-era benchmarks reported 14-24x higher throughput than HuggingFace Transformers on the same hardware, and the engine has kept that lead through continuous batching, optimized CUDA kernels, and the PagedAttention algorithm that eliminates memory fragmentation.

OpenAI API Compatibility: vLLM includes a built-in API server that’s fully compatible with OpenAI’s format. This allows seamless migration from OpenAI to self-hosted infrastructure without changing application code. Simply point your API client to vLLM’s endpoint and it works transparently.

PagedAttention Algorithm: The core innovation behind vLLM’s performance is PagedAttention, which applies the concept of virtual memory paging to attention mechanisms. Instead of allocating contiguous memory blocks for KV caches (which leads to fragmentation), PagedAttention divides memory into fixed-size blocks that can be allocated on-demand. This reduces memory waste by up to 4x and enables much larger batch sizes.

Continuous Batching: Unlike static batching where you wait for all sequences to complete, vLLM uses continuous (rolling) batching. As soon as one sequence finishes, a new one can be added to the batch. This maximizes GPU utilization and minimizes latency for incoming requests.

Multi-GPU Support: vLLM supports tensor parallelism and pipeline parallelism for distributing large models across multiple GPUs. It can efficiently serve models that don’t fit in a single GPU’s memory, supporting configurations from 2 to 8+ GPUs.

Wide Model Support: Compatible with popular model architectures including LLaMA, Mistral, Mixtral, Qwen, Phi, Gemma, and many others. Supports both instruction-tuned and base models from HuggingFace Hub.

When to Use vLLM

vLLM excels in specific scenarios where its strengths shine:

Production API Services: When you need to serve an LLM to many concurrent users via API, vLLM’s high throughput and efficient batching make it the best choice. Companies running chatbots, code assistants, or content generation services benefit from its ability to handle hundreds of requests per second.

High-Concurrency Workloads: If your application has many simultaneous users making requests, vLLM’s continuous batching and PagedAttention enable serving more users with the same hardware compared to alternatives.

Cost Optimization: When GPU costs are a concern, vLLM’s superior throughput means you can serve the same traffic with fewer GPUs, directly reducing infrastructure costs. The 4x memory efficiency from PagedAttention also allows using smaller, cheaper GPU instances.

Kubernetes Deployments: vLLM’s stateless design and container-friendly architecture make it ideal for Kubernetes clusters. Its consistent performance under load and straightforward resource management integrate well with cloud-native infrastructure.

When NOT to Use vLLM: For local development, experimentation, or single-user scenarios, tools like Ollama or llama.cpp provide better user experience with simpler setup. vLLM’s complexity is justified when you need its performance advantages for production workloads.

How to Install vLLM

Prerequisites

Before installing vLLM, ensure your system meets these requirements:

  • OS: Linux with glibc 2.35 or newer (Ubuntu 22.04+, Debian 12+, RHEL 9+)
  • GPU: NVIDIA GPU with compute capability 7.5 or higher (T4, A10, A100, L4, H100, B200, RTX 20/30/40/50 series). Older 7.0-generation cards such as the V100 are no longer in the supported list. AMD GPUs are supported through separate ROCm wheels — see ROCm vs Vulkan for AMD Local LLM Hosting for the AMD-specific image, device flags, and kernel-coverage caveats.
  • CUDA: The default pip wheels are compiled against CUDA 12.9; 12.8 and 13.0 variants are published alongside them. Blackwell GPUs (B200, GB200) require CUDA 12.8+. Your NVIDIA driver must be new enough for the CUDA version you install.
  • Python: 3.10 to 3.13 — 3.12 is the recommended version (ROCm wheels are 3.12-only)
  • VRAM: Minimum 16GB for 7B models, 24GB+ for 13B, 40GB+ for larger models

Installation via pip

The recommended install uses uv to create a Python 3.12 environment and select the matching PyTorch CUDA backend automatically (--torch-backend=auto inspects the installed driver). For a deeper look at uv itself, see uv: The New Python Package Project and Environment Manager:

uv venv --python 3.12 --seed
source .venv/bin/activate
uv pip install vllm --torch-backend=auto

# Verify installation
python -c "import vllm; print(vllm.__version__)"

Plain pip works too:

python3 -m venv vllm-env
source vllm-env/bin/activate
pip install vllm

The default wheels are compiled against CUDA 12.9. If you need a different CUDA version, pull the matching PyTorch wheel index — for example, CUDA 12.8:

pip install vllm --extra-index-url https://download.pytorch.org/whl/cu128

The old pinned wheels such as vllm==0.4.2+cu121 no longer exist in this form. Mixing vLLM into an existing environment with a different PyTorch build is a common source of install failures; the project recommends a fresh environment, and a source build when you need a custom CUDA target.

Installation with Docker

Docker provides the most reliable deployment method, especially for production. Pin the image to a release version instead of latest so upgrades are deliberate, and note that variant images exist for CUDA 12.9 (-cu129 tag suffix), AMD ROCm (vllm-openai-rocm), CPU (vllm-openai-cpu) and Intel XPU (vllm-openai-xpu):

# Pull the official vLLM image
docker pull vllm/vllm-openai:v0.31.0

# Run vLLM with GPU support
docker run --gpus all \
    -v ~/.cache/huggingface:/root/.cache/huggingface \
    -p 8000:8000 \
    --ipc=host \
    vllm/vllm-openai:v0.31.0 \
    --model Qwen/Qwen3-8B

The --ipc=host flag is important for multi-GPU setups as it enables proper inter-process communication.

Building from Source

For the latest features or custom modifications, build from source:

git clone https://github.com/vllm-project/vllm.git
cd vllm
pip install -e .

vLLM Quickstart Guide

Running Your First Model

Start vLLM with a model using the command-line interface:

# Download and serve a model with the OpenAI-compatible API
vllm serve Qwen/Qwen3-8B \
    --port 8000

vllm serve is the current CLI for the OpenAI-compatible server. The older python -m vllm.entrypoints.openai.api_server invocation still runs but prints a deprecation warning and may be removed in a future release, so update any scripts or Kubernetes manifests that still use it.

vLLM will automatically download the model from HuggingFace Hub (if not cached) and start the server. You’ll see output indicating the server is ready:

INFO:     Started server process [12345]
INFO:     Waiting for application startup.
INFO:     Application startup complete.
INFO:     Uvicorn running on http://0.0.0.0:8000

Making API Requests

Once the server is running, you can make requests using the OpenAI Python client or curl:

Using curl:

curl http://localhost:8000/v1/completions \
    -H "Content-Type: application/json" \
    -d '{
        "model": "Qwen/Qwen3-8B",
        "prompt": "Explain what vLLM is in one sentence:",
        "max_tokens": 100,
        "temperature": 0.7
    }'

Using OpenAI Python Client:

from openai import OpenAI

# Point to your vLLM server
client = OpenAI(
    base_url="http://localhost:8000/v1",
    api_key="not-needed"  # vLLM ignores this unless you start the server with --api-key
)

response = client.completions.create(
    model="Qwen/Qwen3-8B",
    prompt="Explain what vLLM is in one sentence:",
    max_tokens=100,
    temperature=0.7
)

print(response.choices[0].text)

Chat Completions API:

response = client.chat.completions.create(
    model="Qwen/Qwen3-8B",
    messages=[
        {"role": "system", "content": "You are a helpful assistant."},
        {"role": "user", "content": "What is PagedAttention?"}
    ],
    max_tokens=200
)

print(response.choices[0].message.content)

Advanced Configuration

vLLM offers numerous parameters to optimize performance:

vllm serve Qwen/Qwen3-8B \
    --port 8000 \
    --gpu-memory-utilization 0.95 \  # Use 95% of GPU memory
    --max-model-len 32k \            # Maximum sequence length (8192 and 32k both work)
    --tensor-parallel-size 2 \       # Use 2 GPUs with tensor parallelism
    --dtype float16 \                # Use FP16 precision
    --max-num-seqs 256               # Maximum batch size

Key Parameters Explained:

  • --gpu-memory-utilization: Fraction of GPU memory for the model executor (default 0.92). Higher values allow larger batches but leave less margin for memory spikes.
  • --max-model-len: Maximum context length. Accepts token counts or human-readable values like 32k. Reducing this saves memory for larger batches.
  • --tensor-parallel-size: Number of GPUs to split the model across.
  • --dtype: Data type for weights (auto, float16, bfloat16, float32). auto follows the checkpoint’s dtype, which is usually what you want.
  • --max-num-seqs: Maximum number of sequences in a batch.
  • --max-num-active-seqs: Caps RUNNING admission independently of max_num_seqs (new in v0.31.0) — useful to bound latency without shrinking the scheduler’s batch window.

vLLM vs Ollama

vLLM is engineered for high-throughput, multi-user production serving with continuous batching, PagedAttention, and multi-GPU support. Ollama optimizes for fast local setup, single-user convenience, and simple model management.

For a detailed decision guide covering migration signals, planning steps, Docker Compose setup, and a practical checklist, see Ollama to vLLM: When to Migrate Your Local LLM Server.

vLLM vs Docker Model Runner

Docker’s Model Runner is their official solution for local AI model deployment. How does it compare to vLLM?

Architecture Philosophy

Docker Model Runner aims to be the “Docker for AI” – a simple, standardized way to run AI models locally with the same ease as running containers. It abstracts away complexity and provides a consistent interface across different models and frameworks.

vLLM is a specialized inference engine focused solely on LLM serving with maximum performance. It’s a lower-level tool that you containerize with Docker, rather than a complete platform.

Setup and Getting Started

Docker Model Runner installation is straightforward for Docker users:

docker model pull llama3:8b
docker model run llama3:8b

This similarity to Docker’s image workflow makes it instantly familiar to developers already using containers; the full command set is in the Docker Model Runner cheatsheet.

vLLM requires more initial setup (Python, CUDA, dependencies) or using pre-built Docker images:

docker pull vllm/vllm-openai:latest
docker run --runtime nvidia --gpus all vllm/vllm-openai:latest --model <model-name>

Performance Characteristics

vLLM delivers superior throughput for multi-user scenarios due to PagedAttention and continuous batching. For production API services handling hundreds of requests per second, vLLM’s optimizations provide 2-5x better throughput than generic serving approaches.

Docker Model Runner focuses on ease of use rather than maximum performance. It’s suitable for local development, testing, and moderate workloads, but doesn’t implement the advanced optimizations that make vLLM excel at scale.

Model Support

Docker Model Runner provides a curated model library with one-command access to popular models. It supports multiple frameworks (not just LLMs) including Stable Diffusion, Whisper, and other AI models, making it more versatile for different AI workloads.

vLLM specializes in LLM inference with deep support for transformer-based language models. It supports any HuggingFace-compatible LLM but doesn’t extend to other AI model types like image generation or speech recognition.

Production Deployment

vLLM is battle-tested in production at companies like Anthropic, Replicate, and many others serving billions of tokens daily. Its performance characteristics and stability under heavy load make it the de facto standard for production LLM serving.

Docker Model Runner is newer and positions itself more for development and local testing scenarios. While it could serve production traffic, it lacks the proven track record and performance optimizations that production deployments require.

Integration Ecosystem

vLLM integrates with production infrastructure tools: Kubernetes operators, Prometheus metrics, Ray for distributed serving, and extensive OpenAI API compatibility for existing applications.

Docker Model Runner integrates naturally with Docker’s ecosystem and Docker Desktop. For teams already standardized on Docker, this integration provides a cohesive experience but fewer specialized LLM serving features.

When to Use Each

Use vLLM for:

  • Production LLM API services
  • High-throughput, multi-user deployments
  • Cost-sensitive cloud deployments needing maximum efficiency
  • Kubernetes and cloud-native environments
  • When you need proven scalability and performance

Use Docker Model Runner for:

  • Local development and testing
  • Running various AI model types (not just LLMs)
  • Teams heavily invested in Docker ecosystem
  • Quick experimentation without infrastructure setup
  • Learning and educational purposes

Hybrid Approach: Many teams develop with Docker Model Runner locally for convenience, then deploy with vLLM in production for performance. The Docker Model Runner images can also be used to run vLLM containers, combining both approaches.

Production Deployment Best Practices

Docker Deployment

Create a production-ready Docker Compose configuration:

version: '3.8'

services:
  vllm:
    image: vllm/vllm-openai:latest
    runtime: nvidia
    environment:
      - CUDA_VISIBLE_DEVICES=0,1
    volumes:
      - ~/.cache/huggingface:/root/.cache/huggingface
      - ./logs:/logs
    ports:
      - "8000:8000"
    command: >
      --model Qwen/Qwen3-8B
      --tensor-parallel-size 2
      --gpu-memory-utilization 0.90
      --max-num-seqs 256
      --max-model-len 8192
    restart: unless-stopped
    shm_size: '16gb'
    deploy:
      resources:
        reservations:
          devices:
            - driver: nvidia
              count: 2
              capabilities: [gpu]

Kubernetes Deployment

Deploy vLLM on Kubernetes for production scale:

apiVersion: apps/v1
kind: Deployment
metadata:
  name: vllm-server
spec:
  replicas: 2
  selector:
    matchLabels:
      app: vllm
  template:
    metadata:
      labels:
        app: vllm
    spec:
      containers:
      - name: vllm
        image: vllm/vllm-openai:latest
        args:
          - --model
          - Qwen/Qwen3-8B
          - --tensor-parallel-size
          - "2"
          - --gpu-memory-utilization
          - "0.90"
        resources:
          limits:
            nvidia.com/gpu: 2
        ports:
        - containerPort: 8000
        volumeMounts:
        - name: cache
          mountPath: /root/.cache/huggingface
      volumes:
      - name: cache
        hostPath:
          path: /mnt/huggingface-cache
---
apiVersion: v1
kind: Service
metadata:
  name: vllm-service
spec:
  selector:
    app: vllm
  ports:
  - port: 80
    targetPort: 8000
  type: LoadBalancer

Monitoring and Observability

The OpenAI-compatible server exposes Prometheus metrics at /metrics by default — no extra flags needed:

curl -s http://localhost:8000/metrics | grep '^vllm:'

Key metrics to monitor:

  • vllm:num_requests_running / vllm:num_requests_waiting / vllm:num_requests_swapped — requests in the RUNNING, WAITING and SWAPPED scheduler states
  • vllm:kv_cache_usage_perc — fraction of used KV cache blocks (0–1)
  • vllm:time_to_first_token_seconds — TTFT histogram
  • vllm:time_per_output_token_seconds — per-token generation speed
  • vllm:num_preemptions_total — cumulative preemptions; a rising counter means the KV cache is oversubscribed and requests are being recomputed

For dashboards, PromQL examples and alert thresholds around LLM serving, see Monitor LLM Inference in Production (2026): Prometheus & Grafana for vLLM, TGI, llama.cpp.

Performance Tuning

Optimize GPU Memory Utilization: The default is 0.92; adjust based on observed behavior. Higher values allow larger batches but risk OOM errors during traffic spikes.

Tune Max Sequence Length: If your use case doesn’t need full context length, reduce --max-model-len. This frees memory for larger batches. For example, if you only need 4K context, set --max-model-len 4096 instead of using the model’s maximum (often 8K-32K).

Choose Appropriate Quantization: For models that support it, use quantized versions (8-bit, 4-bit) to reduce memory and increase throughput:

--quantization awq  # For AWQ quantized models
--quantization gptq # For GPTQ quantized models

Prefix Caching: Prefix caching is enabled by default in the current engine, so requests sharing a prompt prefix (a fixed system prompt, a RAG context template) reuse its KV cache without any flag. Disable it with --no-enable-prefix-caching for workloads with many unique long prompts where cache bookkeeping buys nothing.

Troubleshooting Common Issues

Out of Memory Errors

Symptoms: Server crashes with CUDA out of memory errors.

Solutions:

  • Reduce --gpu-memory-utilization to 0.85 or 0.80
  • Decrease --max-model-len if your use case allows
  • Lower --max-num-seqs to reduce batch size
  • Use a quantized model version
  • Enable tensor parallelism to distribute across more GPUs

Two distinct OOM shapes need different fixes. A startup OOM happens before any request: vLLM profiles memory and fails to allocate even the minimum KV cache, so lowering --max-num-seqs does not help — reduce --max-model-len or raise --gpu-memory-utilization if you have headroom. A runtime OOM under load can come from CUDA-graph capture memory, which scales with --max-num-seqs but sits outside the KV-cache accounting; run once with --enforce-eager to check whether CUDA graphs are the missing memory, and re-test the exact same config before blaming the cache. If boots on identical hardware keep succeeding with the same allocation, pass back the --kv-cache-memory-bytes value vLLM logs at startup to skip the profiling pass on subsequent boots.

On a single 16 GB card, most of these OOM errors trace back to the KV-cache budget rather than the weights alone — KV Cache on 16 GB GPUs covers the exact formula, FP8 KV-cache dtype, and prefix-caching trade-offs behind --max-model-len and --max-num-seqs before you resort to a smaller model.

Slow Startup

Symptoms: The server takes minutes from launch to “Application startup complete”.

Solutions:

  • Weight loading, compilation and CUDA-graph capture dominate startup — --enforce-eager skips both compile and capture and tells you how much of the boot they cost, at the price of steady-state decode speed
  • Reuse the logged --kv-cache-memory-bytes value to skip the memory-profiling and CUDA-graph estimation passes on repeated boots
  • On v0.30.0+, the vllm preload command runs a persistent weight-cache daemon that keeps post-quantized weights resident in GPU memory, so engine restarts map weights over CUDA IPC instead of reloading from disk (--load-format ipc_cache)
  • For NVFP4 checkpoints that OOM during startup compilation, limiting build threads with MAX_JOBS=4 and NVCC_THREADS=4 is a widely used fix — it slows the JIT build but stops the OOM

Low Throughput

Symptoms: Server handles fewer requests than expected.

Solutions:

  • Increase --max-num-seqs to allow larger batches
  • Raise --gpu-memory-utilization if you have headroom
  • Check if CPU is bottlenecked with htop – consider faster CPUs
  • Verify GPU utilization with nvidia-smi – should be 95%+
  • Enable FP16 if using FP32: --dtype float16

Slow First Token Time

Symptoms: High latency before generation starts.

Solutions:

  • Use smaller models for latency-critical applications
  • Prefix caching is on by default — verify it is active rather than re-adding the flag
  • Reduce --max-num-seqs to prioritize latency over throughput
  • Consider speculative decoding for supported models
  • Optimize tensor parallelism configuration

Model Loading Failures

Symptoms: Server fails to start, can’t load model.

Solutions:

  • Verify model name matches HuggingFace format exactly
  • Check network connectivity to HuggingFace Hub
  • Ensure sufficient disk space in ~/.cache/huggingface
  • For gated models, set HF_TOKEN environment variable
  • Try manually downloading with hf download <model> (or huggingface-cli download <model> on older installs)

Advanced Features

Speculative Decoding

vLLM supports speculative decoding, where a smaller draft model proposes tokens that a larger target model verifies. This can accelerate generation by 1.5-2x. For a comprehensive guide to speculative decoding methods — draft models, EAGLE-3, P-EAGLE, and n-gram — see Speculative Decoding: Faster Inference Without Quality Loss.

vllm serve meta-llama/Llama-3.1-8B-Instruct \
    --speculative-config '{"method": "draft_model", "model": "meta-llama/Llama-3.2-1B-Instruct", "num_speculative_tokens": 4}'

The old --speculative-model and --num-speculative-tokens flags are gone — configuration moved into the single --speculative-config JSON argument, which also selects the method (draft_model, ngram, EAGLE-family, or MTP for models with native multi-token-prediction heads).

LoRA Adapters

Serve multiple LoRA adapters on top of a base model without loading multiple full models:

vllm serve meta-llama/Llama-3.1-8B-Instruct \
    --enable-lora \
    --lora-modules sql-lora=/path/to/sql-adapter \
                   code-lora=/path/to/code-adapter

Then specify which adapter to use per request by putting its name in the model field — the same mechanism scales to dozens of task-specific adapters (customer-specific or per-domain variants) with only the base model resident in memory:

response = client.completions.create(
    model="sql-lora",  # Use the SQL adapter
    prompt="Convert this to SQL: Show me all users created this month"
)

Prefix Caching

Prefix caching avoids recomputing the KV cache for shared prompt prefixes and is enabled by default in the current engine.

It pays off most for:

  • Chatbots with fixed system prompts
  • RAG applications with consistent context templates
  • Few-shot learning prompts repeated across requests

For requests sharing a prefix, time-to-first-token drops substantially because prefill for the shared part is skipped.

Integration Examples

LangChain Integration

from langchain_community.llms import VLLMOpenAI

llm = VLLMOpenAI(
    openai_api_key="EMPTY",
    openai_api_base="http://localhost:8000/v1",
    model_name="Qwen/Qwen3-8B",
    max_tokens=512,
    temperature=0.7,
)

response = llm("Explain PagedAttention in simple terms")
print(response)

FastAPI Application

from fastapi import FastAPI
from openai import AsyncOpenAI

app = FastAPI()
client = AsyncOpenAI(
    base_url="http://localhost:8000/v1",
    api_key="not-needed"
)

@app.post("/generate")
async def generate(prompt: str):
    response = await client.completions.create(
        model="Qwen/Qwen3-8B",
        prompt=prompt,
        max_tokens=200
    )
    return {"result": response.choices[0].text}

Performance Benchmarks

Real-world performance data helps illustrate vLLM’s advantages. The figures below come from the project’s launch-era documentation and early community measurements — treat them as order-of-magnitude indications of the architecture’s effect, not as current results for today’s model generation:

Throughput Comparison (Mistral-7B on A100 GPU):

  • vLLM: ~3,500 tokens/second with 64 concurrent users
  • HuggingFace Transformers: ~250 tokens/second with same concurrency
  • Ollama: ~1,200 tokens/second with same concurrency
  • Result: vLLM provides 14x improvement over basic implementations

Memory Efficiency (LLaMA-2-13B):

  • Standard implementation: 24GB VRAM, 32 concurrent sequences
  • vLLM with PagedAttention: 24GB VRAM, 128 concurrent sequences
  • Result: 4x more concurrent requests with same memory

Latency Under Load (Mixtral-8x7B on 2xA100):

  • vLLM: P50 latency 180ms, P99 latency 420ms at 100 req/s
  • Standard serving: P50 latency 650ms, P99 latency 3,200ms at 100 req/s
  • Result: vLLM maintains consistent latency under high load

Cost Analysis

Understanding the cost implications of choosing vLLM:

Scenario: Serving 1M requests/day

With Standard Serving:

  • Required: 8x A100 GPUs (80GB)
  • AWS cost: ~$32/hour × 24 × 30 = $23,040/month
  • Cost per 1M tokens: ~$0.75

With vLLM:

  • Required: 2x A100 GPUs (80GB)
  • AWS cost: ~$8/hour × 24 × 30 = $5,760/month
  • Cost per 1M tokens: ~$0.19
  • Savings: $17,280/month (75% reduction)

This cost advantage grows with scale. Organizations serving billions of tokens monthly save hundreds of thousands of dollars by using vLLM’s optimized serving instead of naive implementations.

Security Considerations

Authentication

vLLM ships a built-in --api-key flag, but per the official docs it only authenticates endpoints under the /v1, /v2 and /inference path prefixes — other endpoints on the same server remain unauthenticated. Treat it as one layer, not the whole answer; for full coverage, implement authentication at the reverse proxy level:

# Nginx configuration
location /v1/ {
    auth_request /auth;
    proxy_pass http://vllm-backend:8000;
}

location /auth {
    proxy_pass http://auth-service:8080/verify;
    proxy_pass_request_body off;
    proxy_set_header Content-Length "";
    proxy_set_header X-Original-URI $request_uri;
}

Or use API gateways like Kong, Traefik, or AWS API Gateway for enterprise-grade authentication and rate limiting.

Network Isolation

Run vLLM in private networks, not directly exposed to the internet:

# Kubernetes NetworkPolicy example
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
  name: vllm-access
spec:
  podSelector:
    matchLabels:
      app: vllm
  policyTypes:
  - Ingress
  ingress:
  - from:
    - podSelector:
        matchLabels:
          role: api-gateway
    ports:
    - protocol: TCP
      port: 8000

Rate Limiting

Implement rate limiting to prevent abuse:

# Example using Redis for rate limiting
from fastapi import FastAPI, HTTPException
from fastapi.middleware.cors import CORSMiddleware
import redis
from datetime import datetime, timedelta

app = FastAPI()
redis_client = redis.Redis(host='localhost', port=6379)

@app.middleware("http")
async def rate_limit_middleware(request, call_next):
    client_ip = request.client.host
    key = f"rate_limit:{client_ip}"
    
    requests = redis_client.incr(key)
    if requests == 1:
        redis_client.expire(key, 60)  # 60 second window
    
    if requests > 60:  # 60 requests per minute
        raise HTTPException(status_code=429, detail="Rate limit exceeded")
    
    return await call_next(request)

Model Access Control

For multi-tenant deployments, control which users can access which models:

ALLOWED_MODELS = {
    "user_tier_1": ["Qwen/Qwen3-8B"],
    "user_tier_2": ["Qwen/Qwen3-8B", "meta-llama/Llama-2-13b-chat-hf"],
    "admin": ["*"]  # All models
}

def verify_model_access(user_tier: str, model: str) -> bool:
    allowed = ALLOWED_MODELS.get(user_tier, [])
    return "*" in allowed or model in allowed

Migration Guide

From OpenAI to vLLM

Migrating from OpenAI to self-hosted vLLM is straightforward thanks to API compatibility:

Before (OpenAI):

from openai import OpenAI

client = OpenAI(api_key="sk-...")
response = client.chat.completions.create(
    model="gpt-3.5-turbo",
    messages=[{"role": "user", "content": "Hello"}]
)

After (vLLM):

from openai import OpenAI

client = OpenAI(
    base_url="https://your-vllm-server.com/v1",
    api_key="your-internal-key"  # If you added authentication
)
response = client.chat.completions.create(
    model="Qwen/Qwen3-8B",
    messages=[{"role": "user", "content": "Hello"}]
)

Only two changes needed: update base_url and model name. All other code remains identical.

From Ollama to vLLM

Ollama uses a different API format. The basic client-side change is switching from Ollama’s REST endpoint to vLLM’s OpenAI-compatible API:

Ollama API:

import requests

response = requests.post('http://localhost:11434/api/generate',
    json={'model': 'llama2', 'prompt': 'Why is the sky blue?'})

vLLM Equivalent:

from openai import OpenAI

client = OpenAI(base_url="http://localhost:8000/v1", api_key="not-needed")
response = client.completions.create(
    model="meta-llama/Llama-2-7b-chat-hf",
    prompt="Why is the sky blue?"
)

The Ollama API is plain REST with its own request shape, while vLLM speaks the OpenAI protocol — that API-format difference is the main client-side change, and the staged approach (run both servers, shift traffic endpoint by endpoint) is covered in the migration guide linked earlier in this article.

From HuggingFace Transformers to vLLM

Direct Python usage migration:

HuggingFace:

from transformers import AutoModelForCausalLM, AutoTokenizer

model = AutoModelForCausalLM.from_pretrained("Qwen/Qwen3-8B")
tokenizer = AutoTokenizer.from_pretrained("Qwen/Qwen3-8B")

inputs = tokenizer("Hello", return_tensors="pt")
outputs = model.generate(**inputs, max_new_tokens=100)
result = tokenizer.decode(outputs[0])

vLLM:

from vllm import LLM, SamplingParams

llm = LLM(model="Qwen/Qwen3-8B")
sampling_params = SamplingParams(max_tokens=100)

outputs = llm.generate("Hello", sampling_params)
result = outputs[0].outputs[0].text

vLLM’s Python API is simpler and much faster for batch inference.

Recent Releases: What Changed in 2026

Three releases in September–October 2026 (v0.29.0, v0.30.0, v0.31.0) show where development is heading:

Fast restarts: vllm preload launches a persistent per-GPU weight-cache daemon that keeps post-quantized, TP-sharded weights resident in GPU memory, so restarted engines map them over CUDA IPC (--load-format ipc_cache) instead of reloading from disk. Experimental vllm snapshot create/restore snapshots a fully initialized engine with CRIU.

Speculative decoding on Model Runner V2: draft-model speculative decoding, custom logits processors and adaptive verification moved onto the new model runner, with the LiLiCorr drafter added and async scheduling for the DFlash drafter.

Large-scale MoE serving: new balanced expert-parallel all2all backends (--all2all-backend moonep), DeepEPv2 with sequence parallelism, EPLB with shared-expert overlap, and prefill context parallelism combined with data parallelism.

Scheduling control: --max-num-active-seqs caps RUNNING admission independently of max_num_seqs, and the waiting queue now schedules requests that already hold KV blocks first.

Roadmap items from older vLLM guides have long since shipped — GGUF support, multimodal serving, disaggregated prefill/decode and multi-node inference are all part of current releases, so treat any tutorial that lists them as “coming soon” as outdated. Release notes for every version are on the vLLM releases page.

  • Ollama Cheatsheet - Complete Ollama command reference and cheatsheet covering installation, model management, API usage, and best practices for local LLM deployment. Essential for developers using Ollama alongside or instead of vLLM.

  • Docker Model Runner vs Ollama: Which to Choose? - In-depth comparison of Docker’s Model Runner and Ollama for local LLM deployment, analyzing performance, GPU support, API compatibility and use cases. Helps understand the competitive landscape vLLM operates in.

External Resources and Documentation

  • vLLM GitHub Repository - Official vLLM repository with source code, comprehensive documentation, installation guides, and active community discussions. Essential resource for staying current with latest features and troubleshooting issues.

  • vLLM Documentation - Official documentation covering all aspects of vLLM from basic setup to advanced configuration. Includes API references, performance tuning guides, and deployment best practices.

  • PagedAttention Paper - Academic paper introducing PagedAttention algorithm that powers vLLM’s efficiency. Essential reading for understanding the technical innovations behind vLLM’s performance advantages.

  • vLLM Blog - Official vLLM blog featuring release announcements, performance benchmarks, technical deep dives, and community case studies from production deployments.

  • HuggingFace Model Hub - Comprehensive repository of open-source LLMs that work with vLLM. Search for models by size, task, license, and performance characteristics to find the right model for your use case.

  • Ray Serve Documentation - Ray Serve framework documentation for building scalable, distributed vLLM deployments. Ray provides advanced features like autoscaling, multi-model serving, and resource management for production systems.

  • NVIDIA TensorRT-LLM - NVIDIA’s TensorRT-LLM for highly optimized inference on NVIDIA GPUs. Alternative to vLLM with different optimization strategies, useful for comparison and understanding the inference optimization landscape.

  • OpenAI API Reference - Official OpenAI API documentation that vLLM’s API is compatible with. Reference this when building applications that need to work with both OpenAI and self-hosted vLLM endpoints interchangeably.

Subscribe

Get new posts on AI systems, Infrastructure, and AI engineering.