TensorRT-LLM

A Python API for defining Large Language Models and performing inference on NVIDIA GPUs.

AI ModelsLLM ToolsUnknownOpen sourceAPI available
Visit tool

nvidia.github.io

About TensorRT-LLM

TensorRT LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and supports state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs.

Description summarised by AI from the sources listed below.

Key features

  • easy-to-use Python API
  • Easy-to-use Python API
  • state-of-the-art optimizations
  • State-of-the-art optimizations for inference on NVIDIA GPUs
  • inference execution

Use cases

  • Large Language Model inference
  • NVIDIA GPU optimization

Pricing

Pricing model: Unknown — we have not been able to confirm pricing from the official website, so nothing is stated here.

Pricing from the tool's own website.