Skip to main content
End-to-End Platform Development & Hands-On Labs

Development Arena

Build production data engineering pipelines from the ground up across modern lakehouse and warehouse architectures. Architect Bronze, Silver, and Gold layers, configure Delta tables, integrate external APIs, design dimensional data marts, implement CDC pipelines, and tackle interactive Hands-On Engineering Labs.

High Priority
200 XPDEV-001

Build Customer Lakehouse Pipeline

Build a production-ready customer data pipeline from source ingestion through Bronze, Silver, and Gold layers for analytics consumption.

DatabricksPySparkDelta LakeSnowflake
Award: ๐Ÿ—๏ธ Lakehouse Builder
Difficulty:Intermediate
Start Building โ†’
High Priority
250 XPDEV-002

Incremental Sales Pipeline

Build a reliable incremental sales pipeline that processes cloud storage transaction files, handles duplicates and late data, and delivers trusted analytics datasets.

DatabricksPySparkDelta LakeSnowflakeCloud Storage
Award: โšก Incremental Engineer
Difficulty:Intermediate
Start Building โ†’
High Priority
300 XPDEV-003

REST API Integration Pipeline

Build an ingestion pipeline consuming third-party REST API customer activity with pagination, transient error handling, rate limiting, and Silver/Gold curation.

DatabricksPySparkREST APIJSONDelta LakeSnowflake
Award: ๐Ÿ”Œ Integration Engineer
Difficulty:Intermediate
Start Building โ†’
High Priority
400 XPDEV-004

Enterprise Sales Reporting Data Mart

Design and implement a dimensional reporting data mart with explicit grain definition, fact/dimension relationships, SCD handling, and Snowflake reconciliation.

SnowflakeDatabricksPySparkSQLDelta Lake
Award: ๐Ÿ›๏ธ Data Mart Architect
Difficulty:Intermediate โ†’ Advanced
Start Building โ†’
Critical Priority
500 XPDEV-005

CDC-Based Customer Data Pipeline

Design a Change Data Capture (CDC) processing pipeline handling INSERT, UPDATE, and DELETE event streams with out-of-order sequencing, replay, and current-state materialization.

DatabricksPySparkDelta LakeSnowflakeCDC Concepts
Award: ๐Ÿ”„ Change Data Engineer
Difficulty:Advanced
Start Building โ†’
Critical Priority
600 XPDEV-006

Real-Time Sales Streaming Pipeline

Build a low-latency streaming pipeline continuously ingesting real-time sales events, handling watermarking, late data, stateful micro-batch aggregations, and Snowflake synchronization.

DatabricksPySparkStructured StreamingDelta LakeSnowflake
Award: ๐ŸŒŠ Streaming Engineer
Difficulty:Advanced
Start Building โ†’
High Priority
700 XPDEV-007

Customer ML Feature Engineering Pipeline

Design a production-grade ML feature engineering pipeline transforming customer and transaction history into reusable feature stores with strict point-in-time leakage prevention.

DatabricksPySparkDelta LakeMLflowSnowflake
Award: ๐Ÿง  Feature Engineer
Difficulty:Advanced
Start Building โ†’
High Priority
800 XPDEV-008

Production CI/CD Data Platform

Implement automated GitHub Actions CI/CD workflows for multi-environment data pipelines with linting, unit/integration testing, secret management, and rollback mechanisms.

GitHub ActionsDatabricksPySparkSnowflakeCloud CI/CD
Award: ๐Ÿš€ Production Data Engineer
Difficulty:Expert
Start Building โ†’
Critical Priority
900 XPDEV-009

Multi-Source Enterprise Data Platform

Build a unified enterprise platform integrating Customer DB, Sales Files, Product REST API, and Store Reference Data with entity resolution, SLA freshness, and cross-source reconciliation.

DatabricksPySparkSnowflakeREST APICloud StorageDelta Lake
Award: ๐Ÿ›๏ธ Data Platform Architect
Difficulty:Expert
Start Building โ†’
Critical Priority
1000 XPDEV-010

Enterprise Data Engineering Capstone

Architect and build an end-to-end enterprise lakehouse platform combining batch ingestion, real-time streaming, CDC, REST APIs, ML feature stores, CI/CD, and Snowflake data marts.

DatabricksPySparkStructured StreamingCDCMLflowGitHub ActionsSnowflake
Award: ๐Ÿ† Enterprise Data Engineer
Difficulty:Expert / Capstone
Start Building โ†’
Critical Priority
600 XPDEV-011
โšก HANDS-ON LAB

Build a Production-Ready Sales Pipeline

Hands-On Lab: Ingest raw retail sales CSV files, clean corrupt/duplicate data, construct Bronze, Silver, and Gold Delta layers, and calculate executive business metrics.

PySparkDatabricksDelta LakeSQLSnowflake
Award: ๐Ÿ› ๏ธ Pipeline Builder
Difficulty:Advanced Hands-On
Start Lab โ†’
Critical Priority
700 XPDEV-012
โšก HANDS-ON LAB

Production Pipeline Debugging Challenge

Hands-On Lab: Receive a broken production PySpark pipeline producing corrupted data, diagnose Cartesian joins and null handling bugs, fix the implementation, and validate reconciliation.

PySparkDatabricksDelta LakeSQLDebugging
Award: ๐Ÿž Pipeline Debugger
Difficulty:Advanced Hands-On
Start Lab โ†’
Critical Priority
800 XPDEV-013
โšก HANDS-ON LAB

Optimize a Slow Spark Pipeline

Hands-On Lab: Profile a deliberately inefficient 48-minute Spark workload, eliminate shuffle bottlenecks, handle key skew, tune partitions, and benchmark measurable runtime improvement.

PySparkSpark SQLDatabricksDelta LakeProfiling
Award: โšก Spark Performance Engineer
Difficulty:Expert Hands-On
Start Lab โ†’
Critical Priority
900 XPDEV-014
โšก HANDS-ON LAB

Production Data Pipeline Incident Recovery

Hands-On Lab: Respond to an urgent 06:30 AM production pipeline failure, inspect error logs, quarantine corrupt files, execute Delta time-travel recovery, and author a Post-Incident Review.

DatabricksPySparkDelta LakeSnowflakeMonitoring
Award: ๐Ÿšจ Data Reliability Engineer
Difficulty:Expert Hands-On
Start Lab โ†’
Critical Priority
1000 XPDEV-015
โšก HANDS-ON LAB

Build an End-to-End Data Engineering Platform

Hands-On Master Challenge: Given a 5-source enterprise starter dataset, build the complete platform code: ingestion, quarantine, Silver conformed layer, Gold star schema, testing, and telemetry.

PySparkDelta LakeDatabricksSnowflakeCI/CDTesting
Award: ๐Ÿ† Hands-On Data Engineer
Difficulty:Expert / Master Hands-On
Start Lab โ†’