Thank you for sending your enquiry! One of our team members will contact you shortly.
Thank you for sending your booking! One of our team members will contact you shortly.
Duration 14 hours
Course Outline
Introduction to AIOps with Open Source Tools
- Key concepts and benefits of AIOps.
- The role of Prometheus and Grafana in the observability stack.
- The place of ML in AIOps: predictive vs. reactive analytics.
Setting Up Prometheus and Grafana
- Installing and configuring Prometheus for time series data collection.
- Building Grafana dashboards utilizing real-time metrics.
- Exploring exporters, relabeling, and service discovery mechanisms.
Data Preprocessing for ML
- Extracting and transforming metrics from Prometheus.
- Preparing datasets specifically for anomaly detection and forecasting tasks.
- Utilizing Grafana’s transformation capabilities or Python-based pipelines.
Applying Machine Learning for Anomaly Detection
- Fundamental ML models for outlier detection (e.g., Isolation Forest, One-Class SVM).
- Training and evaluating models on time series datasets.
- Visualizing detected anomalies within Grafana dashboards.
Forecasting Metrics with ML
- Developing simple forecasting models (Introduction to ARIMA, Prophet, and LSTM).
- Predicting system load or resource utilization patterns.
- Leveraging predictions for early alerting and scaling decisions.
Integrating ML with Alerting and Automation
- Defining alert rules based on ML outputs or predefined thresholds.
- Configuring Alertmanager and notification routing strategies.
- Triggering scripts or automation workflows upon anomaly detection.
Scaling and Operationalizing AIOps
- Integrating with external observability tools (e.g., ELK stack, Moogsoft, Dynatrace).
- Operationalizing ML models within observability pipelines.
- Best practices for implementing AIOps at scale.
Summary and Next Steps
Requirements
- A solid grasp of system monitoring and observability principles.
- Practical experience utilizing Grafana or Prometheus.
- Proficiency in Python and a fundamental understanding of machine learning concepts.
Target Audience
- Observability engineers.
- Infrastructure and DevOps teams.
- Monitoring platform architects and site reliability engineers (SREs).