Thank you for sending your enquiry! One of our team members will contact you shortly.
Thank you for sending your booking! One of our team members will contact you shortly.
Course Outline
Introduction to Agentic AI in Operations
- Transitioning from static runbooks to reasoning agents: the evolution of IT automation
- Anatomy of an agent: reasoning loops, tool utilization, memory management, and planning
- Determining when to automate versus when to maintain human involvement
Agent Frameworks and Architectural Patterns
- Single-agent patterns: ReAct, Plan-and-Execute, and tool-calling cycles
- Multi-agent architectures: supervisor, hierarchical, and swarm models
- Framework analysis: LangGraph, CrewAI, AutoGen, and custom agent implementations
- Constructing your first operational agent: querying monitoring, diagnosing issues, and proposing solutions
Tool Integration for IT Operations
- Connecting agents to Prometheus, Grafana, Datadog, and PagerDuty APIs
- Agent-based log querying: integration with Elasticsearch, Loki, and Splunk
- Leveraging infrastructure tools: kubectl, Terraform, and Ansible via agent actions
- Designing secure tool interfaces with parameter validation and idempotency
Incident Response Automation
- Automated incident triage: severity classification and routing mechanisms
- Generating root cause hypotheses and collecting supporting evidence
- Automated remediation: executing restart, scale, rollback, and failover actions
- Developing an incident runbook agent with progressive levels of autonomy
Safety, Guardrails, and Human-in-the-Loop Mechanisms
- Action classification: read-only, low-risk, high-risk, and destructive operations
- Establishing approval gates and escalation policies for critical operations
- Guardrail patterns: implementing action allowlists, blast radius limits, and rollback guarantees
- Maintaining audit trails and decision provenance for compliance purposes
Multi-Agent Orchestration for Complex Incidents
- Coordinating specialist agents: triage, diagnosis, and remediation roles
- Managing inter-agent communication and shared context
- Resolving conflicts when agents propose contradictory actions
- Simulating end-to-end major incidents with multi-agent response strategies
Observability and Evaluation Metrics
- Tracing agent reasoning chains for effective debugging and auditing
- Evaluating agent decision quality: measuring precision, recall, and time-to-resolution
- Implementing feedback loops: learning from operator overrides and operational outcomes
- Tracking costs and analyzing token economics for operational agents
Production Deployment and Operations
- Deploying agents as services: utilizing APIs, webhooks, and scheduled jobs
- Gradual autonomy rollout: transitioning from shadow mode to full auto-remediation
- Runbooks for agent failures: handling scenarios where the agent itself encounters issues
- Building a business case and measuring ROI for autonomous operations
Requirements
- Professional experience with IT operations, DevOps, or SRE practices.
- Proficiency in Python scripting and REST API interactions.
- Fundamental understanding of LLM capabilities and prompt engineering techniques.
Target Audience
- SRE and DevOps engineers exploring AI-driven automation strategies.
- Platform engineers focused on building self-healing infrastructure.
- IT operations leaders evaluating agentic AI for enhanced incident management.
14 Hours