
Visit tool
llama.app
About Llama.cpp
Llama.cpp is an open-source project that enables LLM inference with minimal setup and state-of-the-art performance on a wide range of hardware. It is a plain C/C++ implementation without any dependencies and supports various backends, including BLAS, BLIS, CANN, CUDA, HIP, and Metal.
Description summarised by AI from the sources listed below.
Key features
- LLM inference
- Support for various backends
- Optimized for different hardware
- Custom CUDA kernels for running LLMs on NVIDIA GPUs
Use cases
- Enabling LLM inference with minimal setup
- State-of-the-art performance on a wide range of hardware
Pricing
Pricing model: Open source. Detailed plans are not recorded; check the official website for current prices.
Pricing from the tool's own website.
