Replication, Partitioning, and Sharding Questions

Scaling and distributing data across nodes: primary-replica and multi-primary replication, read-replica scaling, horizontal partitioning, and sharding strategies with their key-selection and rebalancing challenges. Covers replication lag, failover and split-brain handling, cross-shard operations such as joins, distributed transactions, and global secondary indexes, and the operational cost of a partitioned topology. Key to designing databases that scale horizontally.

HardTechnical
101 practiced

Write an operational playbook for recovering from a single-shard failure that causes read/write errors for a subset of users. Include detection, immediate mitigation steps, data integrity checks, and post-recovery validation steps relevant to a sharded SQL database.

HardTechnical
101 practiced

Explain how database-level replication lag can affect read-after-write consistency in systems using read replicas. As a data engineer, describe patterns to ensure users see their own recent writes (for example, sticky sessions, session routing, read-your-writes guarantees) and trade-offs involved.

MediumTechnical
72 practiced

You receive reports of stale reads for a small percentage of users. Outline a step-by-step troubleshooting plan to identify root cause: what logs, traces, metrics, and checks would you run; how to reproduce; and how to validate whether the issue stems from replication lag, caching, DNS, or client behavior.

MediumSystem Design
95 practiced

Design a monitoring and alerting strategy for a sharded database cluster. List per-shard and cluster-wide metrics to collect (latency percentiles, QPS, replication lag, disk utilization), alert thresholds and severity, and what actions/automation or runbooks should be triggered for common alerts.

HardTechnical
92 practiced

Tail latency spikes increased after increasing shard count. Outline how you would diagnose root causes (network issues, coordination overhead, compaction/GC, metadata lookups), which tools and traces you'd use, and what optimizations you would try to reduce tail latency in a sharded datastore.

Unlock Full Question Bank

Get access to all 16 Replication, Partitioning, and Sharding interview questions and detailed answers.

Sign in to Continue

Join thousands of developers preparing for their dream job.