Llama.cpp

LLM inference in C/C++

LLM ToolsOpen Source AIOpen sourceOpen source
Visit tool

llama.app

About Llama.cpp

Llama.cpp is an open-source project that enables LLM inference with minimal setup and state-of-the-art performance on a wide range of hardware. It is a plain C/C++ implementation without any dependencies and supports various backends, including BLAS, BLIS, CANN, CUDA, HIP, and Metal.

Description summarised by AI from the sources listed below.

Key features

  • LLM inference
  • Support for various backends
  • Optimized for different hardware
  • Custom CUDA kernels for running LLMs on NVIDIA GPUs

Use cases

  • Enabling LLM inference with minimal setup
  • State-of-the-art performance on a wide range of hardware

Pricing

Pricing model: Open source. Detailed plans are not recorded; check the official website for current prices.

Pricing from the tool's own website.