Get in Touch
 Duration 35 hours

Course Outline

Introduction, Learning Objectives, and Migration Strategy

  • Defining course goals, aligning with participant profiles, and establishing success criteria.
  • Exploring high-level migration approaches and associated risk considerations.
  • Setting up workspaces, repositories, and preparing lab datasets.

Day 1 — Migration Fundamentals and Architecture

  • Overview of Lakehouse concepts, Delta Lake, and Databricks architecture.
  • Analyzing the differences between SMP and MPP and their implications for migration.
  • Introducing the Medallion (Bronze→Silver→Gold) design pattern and Unity Catalog basics.

Day 1 Lab — Translating a Stored Procedure

  • Hands-on migration of a sample stored procedure to a Databricks notebook.
  • Mapping temporary tables and cursors to equivalent DataFrame transformations.
  • Validating and comparing the new output against the original results.

Day 2 — Advanced Delta Lake & Incremental Loading

  • Understanding ACID transactions, commit logs, versioning, and time travel features.
  • Implementing Auto Loader, MERGE INTO patterns, upserts, and schema evolution.
  • Optimizing storage with OPTIMIZE, VACUUM, Z-ORDER, and partitioning strategies.

Day 2 Lab — Incremental Ingestion & Optimization

  • Building Auto Loader ingestion processes and MERGE workflows.
  • Applying OPTIMIZE, Z-ORDER, and VACUUM commands, followed by result validation.
  • Measuring and analyzing improvements in read and write performance.

Day 3 — SQL in Databricks, Performance & Debugging

  • Leveraging analytical SQL features, including window functions, higher-order functions, and JSON/array handling.
  • Interpreting Spark UI metrics, DAGs, shuffles, stages, and tasks to diagnose bottlenecks.
  • Applying query tuning patterns such as broadcast joins, hints, caching, and reducing data spilling.

Day 3 Lab — SQL Refactoring & Performance Tuning

  • Refactoring a complex SQL process into optimized Spark SQL.
  • Utilizing Spark UI traces to identify and resolve skew and shuffle issues.
  • Benchmarking performance before and after tuning and documenting the steps taken.

Day 4 — Tactical PySpark: Replacing Procedural Logic

  • Exploring the Spark execution model: drivers, executors, lazy evaluation, and partitioning strategies.
  • Converting loops and cursors into efficient, vectorized DataFrame operations.
  • Implementing modularization, UDFs/pandas UDFs, widgets, and reusable libraries.

Day 4 Lab — Refactoring Procedural Scripts

  • Refactoring a procedural ETL script into modular PySpark notebooks.
  • Incorporating parametrization, unit-style testing, and reusable functions.
  • Conducting code reviews and applying best-practice checklists.

Day 5 — Orchestration, End-to-End Pipeline & Best Practices

  • Designing Databricks Workflows: job structures, task dependencies, triggers, and error handling.
  • Creating incremental Medallion pipelines with integrated quality rules and schema validation.
  • Integrating with Git (GitHub/Azure DevOps), CI, and establishing testing strategies for PySpark logic.

Day 5 Lab — Build a Complete End-to-End Pipeline

  • Assembling a Bronze→Silver→Gold pipeline orchestrated through Workflows.
  • Implementing logging, auditing, retry mechanisms, and automated validations.
  • Executing the full pipeline, validating outputs, and preparing deployment documentation.

Operationalization, Governance, and Production Readiness

  • Best practices for Unity Catalog governance, data lineage, and access controls.
  • Managing costs, cluster sizing, autoscaling, and job concurrency patterns.
  • Creating deployment checklists, rollback strategies, and operational runbooks.

Final Review, Knowledge Transfer, and Next Steps

  • Participant presentations showcasing migration work and key lessons learned.
  • Gap analysis, recommended follow-up activities, and handover of training materials.
  • Providing references, further learning paths, and support options.

Requirements

  • A foundational understanding of data engineering concepts.
  • Practical experience with SQL and stored procedures, specifically within Synapse or SQL Server environments.
  • Familiarity with ETL orchestration principles, such as those used in ADF or similar tools.

Target Audience

  • Technology managers possessing a background in data engineering.
  • Data engineers in the process of migrating procedural OLAP logic to Lakehouse patterns.
  • Platform engineers tasked with overseeing the adoption of Databricks.

Number of participants


Price per participant

Upcoming Courses

Related Categories