Database Monitoring, Troubleshooting, and Diagnostics Questions

Observing and fixing databases in production: health checks, metrics and alerting, and diagnosing common failures like slow queries, lock contention, replication lag, resource exhaustion, and data-integrity incidents such as duplicate keys or lost updates after a crash or migration. Covers a systematic troubleshooting method under incident pressure. Tests operational instincts distinct from design knowledge.

HardTechnical
36 practiced

Production Postgres write latencies have spiked and fsync appears to be the bottleneck. Walk through the possible root causes you'd investigate, and propose a prioritized set of mitigations spanning database configuration, OS-level tuning, and hardware. Explain how you'd measure the impact of each change safely before rolling it out further.

MediumTechnical
54 practiced

A nightly ETL job reads from your production OLTP database, and the application has started slowing down noticeably during that window. Walk through your diagnostic checklist for finding the root cause, whether it's contending queries, lock waits, I/O pressure, or missing indexes, and describe the immediate mitigations you'd apply to protect production traffic while you investigate further.

HardTechnical
43 practiced

On a SQL Server instance, you're seeing heavy PAGELATCH_UP waits on tempdb under a parallel build-and-insert workload. Explain the root causes, how you'd diagnose whether this is allocation contention or latch contention, and what concrete mitigations you'd consider, spanning tempdb configuration, engine-level options, and schema or application changes.

HardTechnical
30 practiced

Write SQL or well-documented pseudo-SQL that helps detect overlapping updates or possible lost updates on an orders table defined as orders(order_id, user_id, status, updated_at). Propose a query that surfaces orders with very close successive updates (for example, multiple updates within 1 second) and explain your assumptions about timestamp granularity and available auditing or logging.

MediumTechnical
39 practiced

You suspect index bloat is causing performance regressions on a PostgreSQL cluster. How would you go about confirming that bloat is really the cause, and what's your plan to rebuild or defragment the affected indexes in production with minimal impact on live traffic?

Unlock Full Question Bank

Get access to all 37 Database Monitoring, Troubleshooting, and Diagnostics interview questions and detailed answers.

Sign in to Continue

Join thousands of developers preparing for their dream job.