Thank you for sending your enquiry! One of our team members will contact you shortly.
Thank you for sending your booking! One of our team members will contact you shortly.
Duration 35 hours
Course Outline
Introduction, Learning Objectives, and Migration Strategy
- Defining course goals, aligning with participant profiles, and establishing success criteria.
- Exploring high-level migration approaches and associated risk considerations.
- Setting up workspaces, repositories, and preparing lab datasets.
Day 1 — Migration Fundamentals and Architecture
- Overview of Lakehouse concepts, Delta Lake, and Databricks architecture.
- Analyzing the differences between SMP and MPP and their implications for migration.
- Introducing the Medallion (Bronze→Silver→Gold) design pattern and Unity Catalog basics.
Day 1 Lab — Translating a Stored Procedure
- Hands-on migration of a sample stored procedure to a Databricks notebook.
- Mapping temporary tables and cursors to equivalent DataFrame transformations.
- Validating and comparing the new output against the original results.
Day 2 — Advanced Delta Lake & Incremental Loading
- Understanding ACID transactions, commit logs, versioning, and time travel features.
- Implementing Auto Loader, MERGE INTO patterns, upserts, and schema evolution.
- Optimizing storage with OPTIMIZE, VACUUM, Z-ORDER, and partitioning strategies.
Day 2 Lab — Incremental Ingestion & Optimization
- Building Auto Loader ingestion processes and MERGE workflows.
- Applying OPTIMIZE, Z-ORDER, and VACUUM commands, followed by result validation.
- Measuring and analyzing improvements in read and write performance.
Day 3 — SQL in Databricks, Performance & Debugging
- Leveraging analytical SQL features, including window functions, higher-order functions, and JSON/array handling.
- Interpreting Spark UI metrics, DAGs, shuffles, stages, and tasks to diagnose bottlenecks.
- Applying query tuning patterns such as broadcast joins, hints, caching, and reducing data spilling.
Day 3 Lab — SQL Refactoring & Performance Tuning
- Refactoring a complex SQL process into optimized Spark SQL.
- Utilizing Spark UI traces to identify and resolve skew and shuffle issues.
- Benchmarking performance before and after tuning and documenting the steps taken.
Day 4 — Tactical PySpark: Replacing Procedural Logic
- Exploring the Spark execution model: drivers, executors, lazy evaluation, and partitioning strategies.
- Converting loops and cursors into efficient, vectorized DataFrame operations.
- Implementing modularization, UDFs/pandas UDFs, widgets, and reusable libraries.
Day 4 Lab — Refactoring Procedural Scripts
- Refactoring a procedural ETL script into modular PySpark notebooks.
- Incorporating parametrization, unit-style testing, and reusable functions.
- Conducting code reviews and applying best-practice checklists.
Day 5 — Orchestration, End-to-End Pipeline & Best Practices
- Designing Databricks Workflows: job structures, task dependencies, triggers, and error handling.
- Creating incremental Medallion pipelines with integrated quality rules and schema validation.
- Integrating with Git (GitHub/Azure DevOps), CI, and establishing testing strategies for PySpark logic.
Day 5 Lab — Build a Complete End-to-End Pipeline
- Assembling a Bronze→Silver→Gold pipeline orchestrated through Workflows.
- Implementing logging, auditing, retry mechanisms, and automated validations.
- Executing the full pipeline, validating outputs, and preparing deployment documentation.
Operationalization, Governance, and Production Readiness
- Best practices for Unity Catalog governance, data lineage, and access controls.
- Managing costs, cluster sizing, autoscaling, and job concurrency patterns.
- Creating deployment checklists, rollback strategies, and operational runbooks.
Final Review, Knowledge Transfer, and Next Steps
- Participant presentations showcasing migration work and key lessons learned.
- Gap analysis, recommended follow-up activities, and handover of training materials.
- Providing references, further learning paths, and support options.
Requirements
- A foundational understanding of data engineering concepts.
- Practical experience with SQL and stored procedures, specifically within Synapse or SQL Server environments.
- Familiarity with ETL orchestration principles, such as those used in ADF or similar tools.
Target Audience
- Technology managers possessing a background in data engineering.
- Data engineers in the process of migrating procedural OLAP logic to Lakehouse patterns.
- Platform engineers tasked with overseeing the adoption of Databricks.