Model-Optimizer

A unified library of SOTA model optimization techniques.

Machine LearningAI ModelsOpen sourceOpen source
Visit tool

nvidia.github.io

About Model-Optimizer

Model-Optimizer is a unified library of SOTA model optimization techniques like quantization, distillation, pruning, neural architecture search, speculative decoding, etc. It compresses deep learning models for downstream deployment frameworks like TensorRT-LLM, TensorRT, vLLM, etc. to optimize inference speed.

Description summarised by AI from the sources listed below.

Key features

  • Quantization
  • Distillation
  • Pruning
  • Neural Architecture Search
  • Speculative Decoding

Use cases

  • Deep learning model compression
  • Inference speed optimization

Pricing

Pricing model: Open source. Detailed plans are not recorded; check the official website for current prices.

Pricing summarised by AI from the sources listed below.