InterviewStack.io LogoInterviewStack.io

Cloud Data Platforms and Managed Services Questions

Evaluating and choosing among managed cloud data platform PRODUCTS: cloud data warehouses (Snowflake, BigQuery, Redshift, Synapse) as vendor options, the storage-and-compute-separation model as a purchasing and operating decision, serverless versus provisioned compute models, warehouse and streaming-service sizing and capacity planning, concurrency and workload management as a platform operating concern, pricing-model comparison and platform-level cost trade-offs, vendor lock-in and portability, platform-to-platform migration, and the recurring managed-versus-self-managed decision applied to warehouses, databases, streaming, and ETL/orchestration services. Focuses on platform SELECTION and operation as a product, not designing the ingestion pipelines, ETL transform patterns, or streaming processing logic that run on top of a chosen platform, and not a single vendor's certification trivia.

MediumTechnical
90 practiced

Compare open-source distributed query engines (Spark, Presto/Trino) with managed cloud data warehouses (Snowflake, BigQuery) for typical analytics workloads: ad-hoc SQL, batch ETL, streaming ETL, and dashboards. Discuss the trade-offs in cost, latency, concurrency, and maintenance burden, and explain when you would choose each in a data platform.

MediumTechnical
83 practiced

You are ingesting 10,000 events per second, averaging 2KB each, into a managed streaming service such as Kinesis Data Streams. Calculate how many shards you need, showing your assumptions and arithmetic, and describe how you would scale shard count up without losing data or disrupting consumers.

MediumTechnical
74 practiced

When would you recommend a fully managed cloud streaming service (such as AWS MSK Serverless, Kinesis Data Streams, Confluent Cloud, or GCP Pub/Sub) over a self-managed Kafka cluster running on VMs or Kubernetes? Give at least five decision criteria covering operational overhead, feature parity, ordering guarantees, scaling, and cost predictability.

MediumTechnical
80 practiced

Compare running Spark workloads on three managed services: AWS EMR, AWS Glue, and GCP Dataproc. For each, discuss operational overhead, pricing model, startup latency, autoscaling capability, and integration with cloud storage and IAM, and describe common scenarios where one is preferred over the others.

HardTechnical
73 practiced

Design a benchmarking methodology to compare Redshift, Snowflake, and BigQuery on cost-per-query, concurrency handling, and latency at scale. Specify the dataset sizes, sample queries, concurrency profiles, and cluster/compute sizing you would use for a fair comparison, and explain how you would make the results reproducible across the three platforms.

Unlock Full Question Bank

Get access to all 30 Cloud Data Platforms and Managed Services interview questions and detailed answers.

Sign in to Continue

Join thousands of developers preparing for their dream job.