Which repartition strategy is the best option for Spark code?
💡 Model Answer
In Spark, repartitioning is used to redistribute data across partitions to improve parallelism or reduce skew. The main strategies are: 1) repartition(n) which performs a full shuffle to create n evenly sized partitions; 2) coalesce(n) which collapses existing partitions without a shuffle, useful when reducing the number of partitions; 3) partitionBy on a DataFrame or RDD to hash‑partition by a key; and 4) bucketBy for Hive tables to pre‑partition data. The best option depends on the use case. If you need a balanced distribution for a join or aggregation, repartition is appropriate, but it incurs shuffle cost. If you are only reducing partitions after a wide transformation, coalesce is cheaper. For join performance, partitioning by the join key (partitionBy) can eliminate shuffle. In practice, start with repartition for large, skewed datasets, then consider coalesce when shrinking partitions. Always monitor shuffle metrics and partition sizes to avoid data skew.
This answer was generated by AI for study purposes. Use it as a starting point — personalize it with your own experience.
🎤 Get questions like this answered in real-time
Assisting AI listens to your interview, captures questions live, and gives you instant AI-powered answers on a discreet on-screen overlay.
Get Assisting AI — Starts at ₹500