InterviewStack.io LogoInterviewStack.io

Cloud Data Platforms and Managed Services Questions

Evaluating and choosing among managed cloud data platform PRODUCTS: cloud data warehouses (Snowflake, BigQuery, Redshift, Synapse) as vendor options, the storage-and-compute-separation model as a purchasing and operating decision, serverless versus provisioned compute models, warehouse and streaming-service sizing and capacity planning, concurrency and workload management as a platform operating concern, pricing-model comparison and platform-level cost trade-offs, vendor lock-in and portability, platform-to-platform migration, and the recurring managed-versus-self-managed decision applied to warehouses, databases, streaming, and ETL/orchestration services. Focuses on platform SELECTION and operation as a product, not designing the ingestion pipelines, ETL transform patterns, or streaming processing logic that run on top of a chosen platform, and not a single vendor's certification trivia.

MediumTechnical
83 practiced

You are ingesting 10,000 events per second, averaging 2KB each, into a managed streaming service such as Kinesis Data Streams. Calculate how many shards you need, showing your assumptions and arithmetic, and describe how you would scale shard count up without losing data or disrupting consumers.

MediumTechnical
88 practiced

You're evaluating managed cloud data warehouse platforms (Snowflake, BigQuery, and Redshift) for a fast-growing analytics team. Walk through the criteria you would use to compare them (architecture model, concurrency handling, pricing model, storage format support, and operational overhead) and make a recommendation for a specific team size and query pattern.

HardTechnical
80 practiced

As a senior data engineer, you are asked to lead a cross-functional migration from an on-prem data warehouse to BigQuery. Describe a plan that covers stakeholder alignment, training for analysts and engineers, cost governance, phased data migration and validation, and how you would handle resistance from teams and demonstrate business impact.

MediumTechnical
76 practiced

How would you troubleshoot an Azure Data Factory pipeline that intermittently fails while copying large files to ADLS with a 403 Forbidden error? List the diagnostic steps, logs to check, and remediation actions.

HardTechnical
78 practiced

You need to scale ADF copy pipelines to perform 1000 parallel file copies between storage accounts without overwhelming network egress or hitting service limits. Provide an architecture and configuration plan including Integration Runtime sizing, parallelism throttles, retries/backoff, and monitoring.

Unlock Full Question Bank

Get access to all 36 Cloud Data Platforms and Managed Services interview questions and detailed answers.

Sign in to Continue

Join thousands of developers preparing for their dream job.