Replication, Partitioning, and Sharding Questions

Scaling and distributing data across nodes: primary-replica and multi-primary replication, read-replica scaling, horizontal partitioning, and sharding strategies with their key-selection and rebalancing challenges. Covers replication lag, failover and split-brain handling, cross-shard operations such as joins, distributed transactions, and global secondary indexes, and the operational cost of a partitioned topology. Key to designing databases that scale horizontally.

HardTechnical
101 practiced

Explain how database-level replication lag can affect read-after-write consistency in systems using read replicas. As a data engineer, describe patterns to ensure users see their own recent writes (for example, sticky sessions, session routing, read-your-writes guarantees) and trade-offs involved.

HardTechnical
95 practiced

How would you detect and mitigate silent data corruption or a split-brain scenario in a replicated database? Propose detection mechanisms, automated mitigation steps, and offline repair procedures that preserve data correctness.

HardSystem Design
76 practiced

Design alerting thresholds and anomaly detection for data divergence signals to minimize false positives. Describe statistical baselining, adaptive thresholds, scoring rules for multi-metric signals (e.g., checksum mismatch rate, replication lag, stale-read rate), and escalation policies for SRE teams.

HardSystem Design
103 practiced

Design an automated failover system for primary-replica database pairs that minimizes split-brain risk. Cover health checks, fencing mechanisms, promotion safety checks, and how you would validate and audit automatic promotions. Also explain what data can be lost during the failover and how RPO and RTO differ between manual and automatic failover.

MediumSystem Design
95 practiced

Design a monitoring and alerting strategy for a sharded database cluster. List per-shard and cluster-wide metrics to collect (latency percentiles, QPS, replication lag, disk utilization), alert thresholds and severity, and what actions/automation or runbooks should be triggered for common alerts.

Unlock Full Question Bank

Get access to all 14 Replication, Partitioning, and Sharding interview questions and detailed answers.

Sign in to Continue

Join thousands of developers preparing for their dream job.