HomeInterview QuestionsEtl, Csv, Big Data

A CSV file is around 10 GB. How would you read, transform, and load it?

🟡 Medium Conceptual Mid level
1 Times asked
Jul 2026 Last seen
Jul 2026 First seen

💡 Model Answer

For a 10 GB CSV, loading it into memory is impractical. I would use a distributed processing framework like Apache Spark or Dask. First, I would read the file in a streaming or chunked manner, specifying the correct schema to avoid type inference overhead. Spark’s DataFrame API can read CSV with options such as header, inferSchema, and delimiter. After loading, I would perform transformations using Spark SQL or DataFrame operations—filtering, aggregations, joins, or UDFs—leveraging Catalyst’s optimization. Finally, I would write the result to a scalable storage format like Parquet or ORC, partitioned by a key (e.g., date) to improve query performance. If Spark isn’t available, I’d use Python’s pandas with read_csv in chunks, process each chunk, and append to a Parquet file using pyarrow. This approach keeps memory usage bounded and scales horizontally.

Sign in to unlock the rest of this answer

This answer was generated by AI for study purposes. Use it as a starting point — personalize it with your own experience.

🎤 Get questions like this answered in real-time

Assisting AI listens to your interview, captures questions live, and gives you instant AI-powered answers on a discreet on-screen overlay.

Get Assisting AI — Starts at ₹500