TensorRT-LLM
A Python API for defining Large Language Models and performing inference on NVIDIA GPUs.
Visit tool
nvidia.github.io
About TensorRT-LLM
TensorRT LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and supports state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs.
Description summarised by AI from the sources listed below.
Key features
- easy-to-use Python API
- Easy-to-use Python API
- state-of-the-art optimizations
- State-of-the-art optimizations for inference on NVIDIA GPUs
- inference execution
Use cases
- Large Language Model inference
- NVIDIA GPU optimization
Pricing
Pricing model: Unknown — we have not been able to confirm pricing from the official website, so nothing is stated here.
Pricing from the tool's own website.
