Monitoring, Logging, and Observability Questions

Understanding running systems through their signals. Covers metrics, logs, and traces, instrumentation, dashboards, alerting design, and log analysis and correlation for debugging production. Emphasizes designing observability so problems are detectable and diagnosable before users are affected.

EasyTechnical
42 practiced

What would you monitor to know a customer-facing web service is healthy, and which of those signals would you prioritize if you could only page on a handful of them? Walk through how you'd decide what's essential versus nice-to-have.

HardTechnical
53 practiced

Product stakeholders are asking for 99.999% availability, but your historical data shows you've been running at 99.90%. How do you approach that conversation? Walk through how you'd figure out what's realistic, what closing the gap would actually cost, and how you'd present the trade-off.

EasyTechnical
53 practiced

What is a runbook, and what does a good one actually need to contain to be useful when someone's paged at 3am? Sketch what you'd want in one for a failed database migration.

HardTechnical
44 practiced

You're generating terabytes of logs per day and need a long-term retention strategy. How would you think about storing older logs cheaply while still being able to search them for forensic investigations and run batch analytics over them?

MediumTechnical
44 practiced

You inherit a dashboard with 40 panels that the on-call team has basically stopped looking at because it's too noisy to be useful during an incident. How would you go about fixing it?

Unlock Full Question Bank

Get access to all Monitoring, Logging, and Observability interview questions and detailed answers.

Sign in to Continue

Join thousands of developers preparing for their dream job.