memra
Rust + CUDA inference engine for serving GGUF models on NVIDIA GPUs.
Visit tool
huggingface.co
About memra
memra is a Rust + CUDA inference engine for serving GGUF models on NVIDIA GPUs. It is optimized for Blackwell (sm_120a), with a separately gated Hopper/H100 (sm_90a) lane.
Description summarised by AI from the sources listed below.
Pricing
Pricing model: Unknown — we have not been able to confirm pricing from the official website, so nothing is stated here.
Pricing summarised by AI from the sources listed below.
