Skip to main content
PYSPARK • LESSON 287

Scenario: User Retention Cohort Analysis

How can we group users by signup cohort month and evaluate how many return to place orders in month 1, month 2, and month 3?

Production3 Minutes860 XP
🤔 THE QUESTION

How can we group users by signup cohort month and evaluate how many return to place orders in month 1, month 2, and month 3?

💡 WHAT IS IT?

Joining signup dates with order activity dates computes cohort month offsets for retention grid matrices.

🎯 WHAT IS IT USED FOR?

SaaS and subscription analytics measuring customer retention and product-market fit over time.

💻 EXAMPLE
df_cohorts = df_orders.join(df_users.select("user_id", col("signup_date").alias("cohort_start")), "user_id").withColumn("cohort_month", date_format("cohort_start", "yyyy-MM")).withColumn("activity_month", date_format("order_date", "yyyy-MM")).withColumn("month_offset", months_between(col("order_date"), col("cohort_start")).cast("int"))

🎯 Mission Objectives

Practice typing production-grade PySpark code for Scenario: User Retention Cohort Analysis.

  • Cohort analysis calculation
  • Months between signup and order
  • Retention matrix aggregation