Skip to main content
PYSPARK • LESSON 190

Partition-Pruned Incremental Daily Ingestion

How do we filter for a specific calendar partition (year, month, day) to minimize S3 data transfer costs?

Production2 Minutes1440 XP
🤔 THE QUESTION

How do we filter for a specific calendar partition (year, month, day) to minimize S3 data transfer costs?

💡 WHAT IS IT?

Filtering directly on partition directory keys to trigger Spark partition pruning and avoid reading non-target directories.

🎯 WHAT IS IT USED FOR?

Daily reconciliation jobs, targeted backfills, and reducing cloud storage egress costs.

💻 EXAMPLE
from pyspark.sql.functions import col

daily_partition = spark.read \
    .parquet("lakehouse/events") \
    .filter((col("year") == 2026) & (col("month") == 8) & (col("day") == 30))

🎯 Mission Objectives

Practice typing production-grade PySpark code for Partition-Pruned Incremental Daily Ingestion.

  • Filter on partition columns (year, month, day)
  • Trigger physical partition pruning
  • Minimize cloud storage I/O