Monitoring, Logging, and Observability Questions

Understanding running systems through their signals. Covers metrics, logs, and traces, instrumentation, dashboards, alerting design, and log analysis and correlation for debugging production. Emphasizes designing observability so problems are detectable and diagnosable before users are affected.

HardTechnical
53 practiced

You've got a request path that goes through three services in sequence, each with its own availability target. Users only care whether the whole request succeeded. How would you think about the end-to-end SLO, and how would you split the error budget across the teams that own those three services?

EasyTechnical
58 practiced

What is alert fatigue, and how would you go about preventing it on a team you're leading?

MediumTechnical
50 practiced

In a long-lived system, how do you evolve a structured logging or metrics schema over time, for example adding a new field or changing what a field means, without breaking dashboards, alerts, and tooling that depend on the old schema?

HardTechnical
44 practiced

You're generating terabytes of logs per day and need a long-term retention strategy. How would you think about storing older logs cheaply while still being able to search them for forensic investigations and run batch analytics over them?

MediumTechnical
48 practiced

What is metric cardinality, and why can high-cardinality labels be dangerous for a metrics backend like Prometheus? What are a few concrete strategies you would use to keep cardinality under control while still preserving useful signal?

Unlock Full Question Bank

Get access to all Monitoring, Logging, and Observability interview questions and detailed answers.

Sign in to Continue

Join thousands of developers preparing for their dream job.