How can you optimize the conversion of a 100‑200 GB CSV file to Parquet using Spark, including partitioning and Adaptive Query Execution (AQE)?
💡 Model Answer
For a 100‑200 GB CSV, the goal is to minimize I/O, shuffle, and file‑system overhead. First, read the CSV with a predefined schema to avoid schema inference overhead and enable predicate push‑down. Set spark.sql.files.maxPartitionBytes to a value that balances the number of input splits (e.g., 128 MB) and the number of executors. Use spark.sql.shuffle.partitions to match the cluster size (e.g., 200–400). Enable AQE (spark.sql.adaptive.enabled=true) so Spark can merge small shuffle partitions and prune partitions at runtime. When writing Parquet, use partitionBy on a high‑cardinality key (date, region) to create many small, query‑prunable files. Choose a fast compression codec such as snappy or zstd; zstd gives better compression at a modest CPU cost. If the cluster has many cores, consider coalescing the final output to a smaller number of files (e.g., 50–100) to avoid the “small file” problem. Finally, monitor the job with Spark UI, look for skewed partitions, and adjust the number of partitions or use dynamic partition pruning. The overall complexity remains O(n) for the data, but careful tuning reduces constant factors and improves throughput.
This answer was generated by AI for study purposes. Use it as a starting point — personalize it with your own experience.
🎤 Get questions like this answered in real-time
Assisting AI listens to your interview, captures questions live, and gives you instant AI-powered answers on a discreet on-screen overlay.
Get Assisting AI — Starts at ₹500