Get in Touch

Course Outline

Production Foundations of Agentic Systems

  • Agentic architectures: loops, tools, memory, and orchestration layers
  • Agent lifecycle: from development and deployment to continuous operation
  • Key challenges in managing agents at production scale

Infrastructure and Deployment Strategies

  • Deploying agents within containerized and cloud environments
  • Scaling approaches: horizontal vs. vertical scaling, concurrency, and throttling
  • Multi-agent orchestration and workload distribution

Monitoring and Observability

  • Critical metrics: latency, success rates, memory consumption, and agent call depth
  • Tracing agent activities and visualizing call graphs
  • Implementing observability with Prometheus, OpenTelemetry, and Grafana

Logging, Auditing, and Compliance

  • Centralized logging and structured event aggregation
  • Ensuring compliance and auditability in agentic workflows
  • Creating audit trails and replay mechanisms for effective debugging

Performance Tuning and Resource Efficiency

  • Minimizing inference overhead and refining agent orchestration cycles
  • Utilizing model caching and lightweight embeddings for enhanced retrieval speed
  • Conducting load testing and stress analysis for AI pipelines

Cost Governance and Control

  • Identifying cost drivers: API calls, memory, compute, and external integrations
  • Monitoring agent-specific costs and establishing chargeback models
  • Implementing automation policies to curb agent sprawl and idle resource usage

CI/CD and Rollout Tactics for Agents

  • Incorporating agent pipelines into CI/CD frameworks
  • Testing, versioning, and rollback strategies for iterative agent enhancements
  • Progressive rollouts and secure deployment methods

Failure Recovery and Reliability Engineering

  • Architecting for fault tolerance and graceful degradation
  • Applying retry, timeout, and circuit breaker patterns for agent stability
  • Incident response and post-mortem frameworks for AI operations

Capstone Project

  • Construct and deploy an agentic AI system with comprehensive monitoring and cost tracking
  • Simulate load, assess performance, and refine resource consumption
  • Present the final architecture and monitoring dashboard to peers

Recap and Future Steps

Requirements

  • Solid grasp of MLOps and production-grade machine learning systems
  • Proficiency in containerized deployments (Docker/Kubernetes)
  • Knowledge of cloud cost optimization and observability tools

Intended Audience

  • MLOps Engineers
  • Site Reliability Engineers (SREs)
  • Engineering Managers responsible for AI infrastructure
 21 Hours

Number of participants


Price per participant

Testimonials (3)

Upcoming Courses

Related Categories