Suppose you have a 100‑200 GB CSV file stored on S3. How would you use AWS Glue and Spark to transform the data and load the final dataset into a target data warehouse (e.g., Redshift or Snowflake)?
1Times asked
Sep 2026Last seen
Sep 2026First seen
💡 Model Answer
- Create a Glue Crawler to scan the S3 bucket and infer the schema of the CSV. The crawler creates a table in the Glue Data Catalog. 2. Write a Glue ETL job (Python or Scala) that uses Spark. Load the CSV into a DynamicFrame via
glueContext.create_dynamic_frame.from_catalog. 3. Transform: Convert the DynamicFrame to a Spark DataFrame, apply any required transformations (filter, cast, UDFs, aggregations). If the data is large, repartition by a key that will be used for downstream queries. 4. Write to a staging location: Write the transformed data back to S3 in Parquet or ORC, partitioned by a useful key (e.g., year/month). 5. Load into the warehouse: Use the Glue job to copy the Parquet files into Redshift via the COPY command (specify IAM role, S3 path, and compression). For Snowflake, use the Snowflake connector or Snowpipe to ingest the Parquet files. 6. Performance tuning: Enable Glue job bookmarks to avoid re‑processing, set--enable-continuous-cloudwatch-logfor monitoring, and usespark.sql.adaptive.enabled=trueto reduce shuffle overhead. This pipeline ensures scalable ingestion, transformation, and loading of 100‑200 GB of data.
This answer was generated by AI for study purposes. Use it as a starting point — personalize it with your own experience.
🎤 Get questions like this answered in real-time
Assisting AI listens to your interview, captures questions live, and gives you instant AI-powered answers on a discreet on-screen overlay.
Get Assisting AI — Starts at ₹500