Skip to main content
PYSPARK • LESSON 168

Production Gold Layer Partitioned Storage Pipeline

How do we derive partition keys and write compressed Snappy Parquet files to cloud object storage?

Expert3 Minutes1000 XP
🤔 THE QUESTION

How do we derive partition keys and write compressed Snappy Parquet files to cloud object storage?

💡 WHAT IS IT?

A production pipeline deriving year/month keys, enabling Snappy compression, and writing partitioned Parquet.

🎯 WHAT IS IT USED FOR?

Enterprise Lakehouse gold-layer publication to AWS S3, Google Cloud Storage, or Azure ADLS Gen2.

💻 EXAMPLE
from pyspark.sql.functions import col, month, year

gold_df = silver_df \
    .withColumn("year", year(col("transaction_date"))) \
    .withColumn("month", month(col("transaction_date")))

gold_df.write \
    .mode("overwrite") \
    .partitionBy("year", "month") \
    .option("compression", "snappy") \
    .parquet("s3://lakehouse/gold/fact_transactions")

🎯 Mission Objectives

Practice typing production-grade PySpark code for Production Gold Layer Partitioned Storage Pipeline.

  • Derive year and month partition columns
  • Configure Snappy compression
  • Write partitioned Parquet to cloud storage