Database Monitoring, Troubleshooting, and Diagnostics Questions

Observing and fixing databases in production: health checks, metrics and alerting, and diagnosing common failures like slow queries, lock contention, replication lag, resource exhaustion, and data-integrity incidents such as duplicate keys or lost updates after a crash or migration. Covers a systematic troubleshooting method under incident pressure. Tests operational instincts distinct from design knowledge.

HardTechnical
35 practiced

Intermittent P99 write latency spikes are impacting global users, and the spikes seem to correlate with heavy analytics jobs that run hourly. Design a methodical approach to isolate the root cause: what instrumentation and sampling would you add, how would you validate that the analytics jobs are actually causing the spikes rather than just coinciding with them, and how would you decide which mitigation to reach for once causality is confirmed?

HardTechnical
35 practiced

A replicated MySQL cluster started reporting duplicate primary keys after a crash and restart. Describe diagnostic steps to determine whether binlog corruption, incomplete transactions, or a split-brain happened. Include commands to inspect binary logs, relay logs, GTIDs/positions, and server UUIDs, and describe how to reconcile the cluster and validate integrity after fixes.

MediumTechnical
37 practiced

Explain how you would use a managed database's built-in performance-insight tooling (for example AWS Performance Insights or Cloud SQL Insights) to find top wait events and high-impact SQL. Walk through a triage workflow that starts from an alert about increased database latency and ends with identifying and deploying a mitigation, and describe how you'd correlate your database-side findings with application traces and logs.

MediumTechnical
39 practiced

You suspect index bloat is causing performance regressions on a PostgreSQL cluster. How would you go about confirming that bloat is really the cause, and what's your plan to rebuild or defragment the affected indexes in production with minimal impact on live traffic?

MediumSystem Design
37 practiced

Design a monitoring dashboard and alert strategy for replication health for a cluster serving 100k QPS reads and 20k writes per minute with asynchronous replication. List the key metrics you would track, suggested alert thresholds, and the automated or operational responses when thresholds are crossed. Consider false positives and cost trade-offs.

Unlock Full Question Bank

Get access to all 29 Database Monitoring, Troubleshooting, and Diagnostics interview questions and detailed answers.

Sign in to Continue

Join thousands of developers preparing for their dream job.