Skip to main content
PYSPARK • LESSON 191

Salting Large Joins to Mitigate Data Skew

How do we add random salt keys to skewed fact records and explode the dimension to eliminate join bottlenecks?

Production3 Minutes1480 XP
🤔 THE QUESTION

How do we add random salt keys to skewed fact records and explode the dimension to eliminate join bottlenecks?

💡 WHAT IS IT?

Salting splits high-cardinality hot keys into multiple sub-keys across cluster executors to eliminate join skew.

🎯 WHAT IS IT USED FOR?

Fixing long-tail straggler tasks in massive joins where 1% of keys contain 90% of row volume.

💻 EXAMPLE
from pyspark.sql.functions import concat, lit, rand

skewed_sales = sales.withColumn("salt", (rand() * 10).cast("int"))
salted_dim = dim.withColumn("salt", explode(array([lit(i) for i in range(10)])))
balanced_join = skewed_sales.join(salted_dim, ["key", "salt"], "inner").drop("salt")

🎯 Mission Objectives

Practice typing production-grade PySpark code for Salting Large Joins to Mitigate Data Skew.

  • Add randomized salt to skewed records
  • Explode dimension table across salt range
  • Execute balanced, non-skewed join