Observability and Monitoring Architecture Questions

Building visibility into infrastructure and services: metrics, logs, and traces, dashboards and alerting, SLIs/SLOs, and the design of an observability stack. Covers instrumenting systems for actionable signal, reducing alert noise, and diagnosing production issues from telemetry. Infrastructure-wide observability, distinct from network-specific monitoring.

HardTechnical
36 practiced

Define SLOs for the observability pipeline itself: ingestion availability, storage durability, query freshness, and end-to-end latency for a dashboard query. For each one, propose a concrete target, how you'd measure it, and what action fires when it's breached.

HardSystem Design
27 practiced

Architect an observability pipeline that has to ingest and store genuinely high-cardinality metrics, millions of unique series, under a tight cost ceiling. Describe your strategy across label reduction, aggregation and rollups, sampling, storage tiering, and retention, and how you'd let someone reconstruct finer-grained detail at query time when they actually need it.

MediumTechnical
29 practiced

Set a concrete retention and downsampling policy for metrics and traces that balances cost against query fidelity, for example raw metrics for 14 days, downsampled metrics for a year, full traces for 30 days then sampled. Walk through your rationale and what it means for the kinds of queries you can still answer after each window closes.

HardSystem Design
34 practiced

A regulated customer requires that PII never lands in raw telemetry. Design an end-to-end pipeline that detects and redacts PII at ingestion while preserving enough context to debug production issues. Cover detection techniques (regex versus ML classifiers), whether masking is deterministic or tokenized, which enforcement point you'd use (agent, collector, or storage), the performance cost, and how you'd prove it's working to an auditor.

HardSystem Design
54 practiced

Design the mechanics of tail-based sampling at real scale, say 100,000 traces per second: spans have to be buffered somewhere until the sampling decision can be made, slow or erroneous traces need their full span set captured, and everything else gets thinned. How do you coordinate that buffering and decision-making across many collector instances without unbounded memory growth?

Unlock Full Question Bank

Get access to all 41 Observability and Monitoring Architecture interview questions and detailed answers.

Sign in to Continue

Join thousands of developers preparing for their dream job.