
kserve
A standardized distributed generative and predictive AI inference platform for scalable, multi-framework deployment on Kubernetes.
Visit tool
kserve.github.io
About kserve
KServe is an open-source standard for self-hosted AI, providing a unified platform for both Generative and Predictive AI inference on Kubernetes. It encapsulates the complexity of autoscaling, networking, health checking, and server configuration to bring cutting-edge serving features to ML deployments.
Description summarised by AI from the sources listed below.
Key features
- Autoscaling
- autoscaling
- GPU Acceleration
- GPU acceleration
- Model Caching
- model caching
- KV cache offloading
- KV Cache Offloading
- Hugging Face ready
- Intelligent Routing
- intelligent routing
- predictive AI
- advanced deployments
- Advanced Deployments
- multi-framework support
- Model Explainability
- model explainability
- canary rollouts
- advanced monitoring
- Advanced Monitoring
Use cases
- generative AI
- predictive AI
- model serving
- machine learning model deployment
Pricing
Pricing model: Open source. Detailed plans are not recorded; check the official website for current prices.
Pricing from the tool's own website.
