Data Reliability and Fault Tolerance Questions

Designing pipelines that survive failures: retries, idempotency, checkpointing, exactly-once semantics, dead-letter handling, and recovery/replay. Covers reasoning about partial failures, poison messages, and consistency guarantees under faults. The resilience angle distinct from monitoring (detecting and alerting on a failure) and from Workflow Orchestration and Scheduling (the DAG/scheduler mechanics that decide whether and when a task runs again, including backfills and dependency management): this topic owns whether the data itself stays correct, not lost, not duplicated, not corrupted, when a process is retried or replayed.

HardTechnical
34 practiced

Provide pseudocode for the Chandy-Lamport distributed snapshot algorithm adapted to capture consistent operator state and in-flight messages in a streaming dataflow. Include handling for a snapshot attempt that partially fails and how to resume or abort safely. Then explain how you would adapt the same algorithm to implement consistent checkpoints across cooperating microservices that exchange events over Kafka or message queues, and discuss the practical challenges of capturing in-flight messages and integrating with persisted broker logs.

MediumTechnical
41 practiced

Your analytics database shows duplicate and occasionally missing user records created by your pipeline, and the job itself hasn't changed. Describe a step-by-step, time-boxed investigation plan to identify the root cause across producers, the message broker, stream processors, and sinks. Specify what logs, offsets, connector state, transactional metadata, and consumer checkpoints you'd inspect, which tests confirm at-least-once versus exactly-once behavior, and the short-term corrective action (deduplicate or roll back).

EasyTechnical
36 practiced

Differentiate between a dead-letter queue (DLQ), a poison message, and a retry policy, in both batch and streaming pipelines. Provide a rule set (an operational decision flow for when to retry, when to DLQ, and when to alert an engineer) that avoids infinite retry loops. Describe how you would design the DLQ message format for diagnostics (including failure reason, offsets, timestamps, schema version), retention/TTL considerations, and a small operational workflow for replaying messages after root-cause fixes.

HardTechnical
29 practiced

Compare coordinator-based two-phase commit (2PC), a write-ahead-log-plus-idempotent-sink (log-based/replay-and-compaction) approach, and eventual-consistency-with-compensating-transactions for writing to multiple heterogeneous sinks atomically. Discuss failure modes, blocking behavior, performance implications, recovery procedures, hybrid approaches that combine two of these, and cases where none of them is sufficient on its own.

MediumTechnical
38 practiced

Write Python-style pseudocode for a streaming operator that performs idempotent writes to an external datastore using Redis to track processed message IDs. Requirements: persist processed IDs in Redis atomically with write intent, use TTL or compaction to prevent unbounded growth, and handle crashes and replays safely. Explain the failure modes.

Unlock Full Question Bank

Get access to all 22 Data Reliability and Fault Tolerance interview questions and detailed answers.

Sign in to Continue

Join thousands of developers preparing for their dream job.