Home › Interview Questions › Explain how Amazon Kinesis works for real‑time dat…

Explain how Amazon Kinesis works for real‑time data streaming and how you can run Spark code on Kinesis streams.

🟡 Medium Conceptual Junior level
1Times asked
Aug 2026Last seen
Aug 2026First seen

💡 Model Answer

Amazon Kinesis is a fully managed service that lets you ingest, buffer, and process streaming data in real time. A Kinesis Data Stream is a set of shards; each shard can ingest up to 1 MB/s of data and read up to 2 MB/s. Producers write records to the stream, and consumers read them in parallel. To run Spark on Kinesis, you use Spark Structured Streaming with the Kinesis connector. The connector reads records from the stream as a streaming DataFrame, applies transformations, and writes results to a sink such as S3, Redshift, or another stream. The typical workflow is: 1) Create a Kinesis stream and configure a producer (e.g., Kinesis Agent, SDK). 2) In Spark, use spark.readStream.format("kinesis") with the stream name, region, and credentials. 3) Apply Spark SQL or DataFrame APIs to process the data. 4) Write the output using writeStream. Spark handles checkpointing and fault tolerance automatically. This integration allows you to build real‑time analytics pipelines that scale horizontally with the number of shards and Spark executors.

This answer was generated by AI for study purposes. Use it as a starting point — personalize it with your own experience.

🎤 Get questions like this answered in real-time

Assisting AI listens to your interview, captures questions live, and gives you instant AI-powered answers on a discreet on-screen overlay.

Get Assisting AI — Starts at ₹500