InterviewStack.io LogoInterviewStack.io

Batch, Streaming, and Real-Time Serving Trade-offs Questions

Reasoning about when to use batch, micro-batch, or continuous streaming and how to serve low-latency analytics: latency, cost, complexity, and correctness trade-offs; lambda vs kappa architectures; and reprocessing semantics. Covers real-time aggregation, freshness vs consistency trade-offs, and reconciling streaming results with batch ground truth, including geospatial and high-throughput real-time workloads under eventual consistency. The data-systems judgment topic for choosing and reconciling batch versus real-time approaches, distinct from the hands-on streaming transport itself.

EasyTechnical
34 practiced

Explain the real differences between batch processing and stream processing for a production data platform: latency, throughput, cost, operational complexity, and correctness. Give one concrete workload that clearly favors each approach, and describe a scenario where a hybrid of the two is the right call.

MediumTechnical
29 practiced

Compare batch processing and stream processing as general models, then bring Lambda and Kappa architecture into the comparison. For a concrete analytics workflow of your choosing, walk through why you would pick pure batch, pure streaming, Lambda, or Kappa.

MediumTechnical
31 practiced

Your team processes 100M events/day through a batch pipeline and retrains models every few hours. Someone proposes adding real-time features for personalization. Before committing to that, how would you evaluate whether the need is real: what minimal experiment or prototype would you run, and what objective success and cost criteria would decide whether it's worth building?

MediumTechnical
29 practiced

Design an approach to compute per-city, per-minute aggregates (trip count, median fare, median ETA) updated every minute for dashboards. Would you use materialized views, streaming pre-aggregations, or periodic batch aggregation, and why? Justify the trade-offs in latency, cost, and accuracy, calling out anything (like a median) that doesn't aggregate as cleanly as a sum or count.

HardTechnical
37 practiced

A team needs both retrainable ML models (which want reproducible, accurate historical data) and low-latency online scoring. Which would you recommend, Lambda or Kappa, and why? Sketch how you would migrate from the other architecture with minimal risk, and describe how you'd keep the online features and the periodic batch snapshots used for training reconciled with each other.

Unlock Full Question Bank

Get access to all 7 Batch, Streaming, and Real-Time Serving Trade-offs interview questions and detailed answers.

Sign in to Continue

Join thousands of developers preparing for their dream job.