Skip to main content
PYSPARK • LESSON 158

Tuning Vectorized Batch Memory Limits

How do we configure maxRecordsPerBatch to avoid Out-Of-Memory (OOM) errors in Arrow conversions?

Expert2 Minutes880 XP
🤔 THE QUESTION

How do we configure maxRecordsPerBatch to avoid Out-Of-Memory (OOM) errors in Arrow conversions?

💡 WHAT IS IT?

spark.sql.execution.arrow.maxRecordsPerBatch limits the number of rows processed per Arrow batch.

🎯 WHAT IS IT USED FOR?

Preventing executor memory exhaustion when handling wide tables or heavy ML feature vectors.

💻 EXAMPLE
spark.conf.set("spark.sql.execution.arrow.maxRecordsPerBatch", 10000)

🎯 Mission Objectives

Practice typing production-grade PySpark code for Tuning Vectorized Batch Memory Limits.

  • Set maxRecordsPerBatch configuration
  • Tune Arrow batch size to 10,000 rows
  • Prevent executor memory exhaustion