Get in Touch
 Duration 14 hours

Course Outline

Preparing Machine Learning Models for Deployment

  • Containerizing models using Docker
  • Exporting models from TensorFlow and PyTorch ecosystems
  • Best practices for versioning and storage

Serving Models on Kubernetes

  • Introduction to inference server architectures
  • Deployment strategies for TensorFlow Serving and TorchServe
  • Configuration of dedicated model endpoints

Techniques for Inference Optimization

  • Implementation of batching strategies
  • Managing concurrent request handling
  • Tuning for optimal latency and throughput

Autoscaling ML Workloads

  • Utilizing the Horizontal Pod Autoscaler (HPA)
  • Applying the Vertical Pod Autoscaler (VPA)
  • Implementing Kubernetes Event-Driven Autoscaling (KEDA)

GPU Allocation and Resource Control

  • Setup and configuration of GPU-enabled nodes
  • Overview of the NVIDIA device plugin
  • Defining resource requests and limits for ML workloads

Strategies for Model Rollout and Release

  • Blue/green deployment techniques
  • Adopting canary rollout patterns
  • Conducting A/B testing for model performance evaluation

Monitoring and Observability for Production ML

  • Tracking metrics specific to inference workloads
  • Establishing robust logging and tracing practices
  • Creating dashboards and configuring alerting systems

Security and Reliability Best Practices

  • Hardening model endpoints against threats
  • Implementing network policies and access controls
  • Ensuring high availability and system resilience

Wrap-Up and Recommended Next Steps

Requirements

  • Working knowledge of containerized application workflows
  • Practical experience with Python-based machine learning models
  • Basic proficiency in Kubernetes core concepts

Target Audience

  • ML Engineers
  • DevOps Engineers
  • Platform Engineering Teams

Number of participants


Price per participant

Testimonials (4)

Upcoming Courses

Related Categories