vllm

High-throughput and memory-efficient LLM serving engine

LLM ToolsOpen sourceOpen sourceAPI available
Visit tool

vllm.ai

About vllm

vLLM is a fast and easy-to-use library for LLM inference and serving. It supports various decoding algorithms, tensor, pipeline, data, expert, and context parallelism for distributed inference.

Description summarised by AI from the sources listed below.

Key features

  • State-of-the-art serving throughput
  • Efficient management of attention key and value memory
  • Efficient management of attention key and value memory with PagedAttention
  • Continuous batching of incoming requests
  • Continuous batching of incoming requests, chunked prefill, prefix caching
  • Quantization
  • Fast and flexible model execution with piecewise and full CUDA/HIP graphs
  • Quantization: FP8, MXFP8/MXFP4, NVFP4, INT8, INT4, GPTQ/AWQ, GGUF, compressed-tensors, ModelOpt, TorchAO, and more
  • Optimized attention kernels
  • Speculative decoding
  • Optimized attention kernels including FlashAttention, FlashInfer, TRTLLM-GEN, FlashMLA, and Triton
  • Automatic kernel generation and graph-level transformations
  • Optimized GEMM/MoE kernels for various precisions using CUTLASS, TRTLLM-GEN, CuTeDSL
  • Speculative decoding including n-gram, suffix, EAGLE, DFlash
  • Automatic kernel generation and graph-level transformations using torch.compile
  • Disaggregated prefill, decode, and encode
  • Seamless integration with popular Hugging Face models
  • Tensor, pipeline, data, expert, and context parallelism for distributed inference
  • Streaming outputs
  • High-throughput serving with various decoding algorithms

Use cases

  • LLM inference and serving
  • Distributed inference
  • Model execution

Pricing

Pricing model: Open source. Detailed plans are not recorded; check the official website for current prices.

Pricing summarised by AI from the sources listed below.