How would you solve the small file problem?
💡 Model Answer
The small file problem arises in Hadoop/Spark when a dataset is split into many tiny files (often < 64 MB). This leads to high metadata overhead, increased job scheduling time, and poor compression. To solve it, you can: 1) Combine small files into larger ones using coalesce or repartition before writing; 2) Use mapreduce.fileoutputcommitter.algorithm.version=2 to merge files during write; 3) Enable spark.sql.files.maxPartitionBytes to control partition size; 4) Use partitioning schemes that produce fewer files per partition; 5) Employ tools like Hadoop DistCp or HDFS concat to merge files post‑write; 6) Store data in columnar formats (Parquet, ORC) that support efficient compression and predicate pushdown. Additionally, consider using a data lakehouse approach where small files are periodically compacted by a background job. These strategies reduce the number of files, lower metadata load, and improve query performance.
This answer was generated by AI for study purposes. Use it as a starting point — personalize it with your own experience.
🎤 Get questions like this answered in real-time
Assisting AI listens to your interview, captures questions live, and gives you instant AI-powered answers on a discreet on-screen overlay.
Get Assisting AI — Starts at ₹500