kserve

A standardized distributed generative and predictive AI inference platform for scalable, multi-framework deployment on Kubernetes.

Machine LearningModel ServingOpen sourceOpen sourceAPI availableDemo available
Visit tool

kserve.github.io

About kserve

KServe is an open-source standard for self-hosted AI, providing a unified platform for both Generative and Predictive AI inference on Kubernetes. It encapsulates the complexity of autoscaling, networking, health checking, and server configuration to bring cutting-edge serving features to ML deployments.

Description summarised by AI from the sources listed below.

Key features

  • Autoscaling
  • autoscaling
  • GPU Acceleration
  • GPU acceleration
  • Model Caching
  • model caching
  • KV cache offloading
  • KV Cache Offloading
  • Hugging Face ready
  • Intelligent Routing
  • intelligent routing
  • predictive AI
  • advanced deployments
  • Advanced Deployments
  • multi-framework support
  • Model Explainability
  • model explainability
  • canary rollouts
  • advanced monitoring
  • Advanced Monitoring

Use cases

  • generative AI
  • predictive AI
  • model serving
  • machine learning model deployment

Pricing

Pricing model: Open source. Detailed plans are not recorded; check the official website for current prices.

Pricing from the tool's own website.