Monitoring, Logging, and Observability Questions
Understanding running systems through their signals. Covers metrics, logs, and traces, instrumentation, dashboards, alerting design, and log analysis and correlation for debugging production. Emphasizes designing observability so problems are detectable and diagnosable before users are affected.
Tell me about a time you led an effort to fix inadequate monitoring or alerting after it let a real incident go undetected too long. What did you change, how did you get buy-in from other teams, and how did you know it worked?
Sample Answer
Direct answer
A strong answer here is less about the specific tooling you shipped and more about the reasoning chain: how you diagnosed why the existing monitoring or alerting failed to surface a real incident in time, how you decided what to fix first, how you got other teams to actually adopt the change instead of just approving it, and how you confirmed afterward that it worked rather than just assuming it did. The story should show you treating a missed page as a debuggable system problem, not a one-off oversight.
Structured elaboration
Two common entry points, same underlying competency
This kind of story usually starts one of two ways, and either is a fine way to frame it: monitoring that had quietly degraded into noise over months (alerts nobody trusted anymore, so a real signal got lost in the pile), or monitoring that was specifically inadequate during one identifiable major incident (a gap that only became obvious in the post-incident review). Both point at the same skill: turning a detection failure into a durable fix.
Diagnosing the gap
Before proposing a fix, audit what actually happened: which alerts fired, which didn't, how long it took someone to notice and acknowledge, and who owned each alert. A useful lens is classifying existing alerts by signal versus noise (did this alert, historically, correlate with something someone needed to act on) rather than jumping straight to needing more alerts.
Prioritizing the fix
Not every gap is worth closing immediately. A reasonable framework is impact times likelihood times effort: fix the alerts that correlate strongly with real user-facing failures, have a low false-positive rate once tuned, and can ship in a small, reversible change. Anything that needs a bigger architectural shift, like a new tracing system or a new SLO framework, goes on a slower track.
Getting buy-in
Data-driven pilots beat mandates. Running the change on one canary service first, tuning thresholds against real traffic, and showing before-and-after alert quality to the teams affected is far more persuasive than presenting a plan in a meeting. It also surfaces objections early, from the people who will actually be paged, before a wider rollout.
Proving it worked
Define leading and lagging indicators up front rather than declaring victory from a gut feeling: time-to-detect, time-to-acknowledge, page volume per week, and the fraction of pages that turned out to require action. Track them for a meaningful window, not just the week after launch when everyone is paying extra attention, and revisit at a fixed cadence, folded into postmortems.
Worked example
Picture the story like this: a payments-adjacent service degraded slowly over a few hours due to a connection-pool leak, and the only alert that existed was a hard latency threshold that didn't trip until users were already seeing failures, well after the degradation started. In the postmortem you pull the alert history and find the team had actually been paged on latency dozens of times that quarter for transient blips, so the useful signal had been drowned out and largely muted by the time the real incident hit. You propose replacing the single hard threshold with an early-warning signal on the leading indicator itself, connection-pool saturation trending up, separated from the page-worthy signal, actual error rate breaching an SLO. You pilot it on that one service for a couple of weeks, and use the pilot's own page volume and detection timing, not a borrowed number from a different team, as the evidence you bring to the rollout conversation. The result you would actually claim afterward is directional and something you can defend if pressed: meaningfully fewer noisy pages, and the early-warning signal firing before user-facing errors rather than after, because those are the two things the redesign specifically targeted.
Trade-offs and pitfalls
The common wrong turn is treating inadequate alerting as not enough alerts and adding more, which usually makes the noise problem worse and erodes trust further. A related pitfall is fixing the alert without fixing ownership: if nobody is accountable for keeping a given alert tuned, it degrades back into noise within a few months, so a durable fix usually includes a clear owner and a recurring review, not just a one-time threshold change. Finally, be honest in the interview about what you would measure rather than asserting a suspiciously precise before-and-after number you cannot actually stand behind. It is far more credible to say what you tracked and why than to hand the interviewer a number that sounds fabricated.
What's the difference between an SLI, an SLO, and an SLA? Walk through how you'd define each one concretely for a service you've worked on, including how you'd measure the indicator and what time window you'd use.
Sample Answer
Direct answer
An SLI is the measured signal (e.g. the percentage of requests that were fast and successful), an SLO is the internal target for that signal over a time window (e.g. 99.9% of requests succeed under 300ms over a rolling 30 days), and an SLA is the external, often contractual promise built on top of an SLO, usually with margin, and real consequences (service credits, penalties) if it's breached. In short: the SLI measures, the SLO targets, the SLA promises with a business wrapper attached.
Defining each one concretely, for an e-commerce checkout service
| Term | Definition here | Concrete value |
|---|---|---|
| SLI | Percentage of checkout requests that return successfully (2xx) within 300ms | Measured every minute from load-balancer access logs |
| SLO | Internal reliability target for that SLI | 99.9% of checkout requests succeed under 300ms, measured over a rolling 30-day window |
| SLA | External commitment to customers, usually looser than the SLO to leave margin | 99.5% monthly checkout availability, with service credits below that threshold |
Explained for a non-technical stakeholder: the SLO is the bar the engineering team holds itself to internally so problems get caught and fixed before they become customer-visible; the SLA is the, usually more forgiving, bar the business is willing to be held to externally, on paper, with money attached if it's missed. The gap between the two is deliberate headroom, not sloppiness.
Worked example: computing the error budget
The error budget is the amount of allowed failure baked into the SLO. It's what lets a team ship changes at all instead of freezing forever.
error budget=(1−SLO target)For the 99.9% SLO above:
error budget=1−0.999=0.001=0.1%Over the 30-day measurement window (30 days equals 2,592,000 seconds):
0.001×2,592,000s=2,592s≈43.2 minutesSo this service is allowed about 43 minutes of budget-consuming failure (however that failure is defined: downtime, over-300ms responses, and so on) across a 30-day window before the SLO itself is breached. If checkout handles, say, 1,000,000 requests over that same window, the equivalent request-based budget is:
0.001×1,000,000=1,000 requests allowed to violate the SLIBoth framings, time-based and request-based, describe the same budget; which one is more useful depends on whether the SLI is availability-style (was the service up) or ratio-style (what fraction of requests succeeded).
Trade-offs and pitfalls
- Confusing SLO and SLA in conversation causes real problems: teams sometimes design their alerting and release-gating around the SLA (the looser, contractual number) instead of the SLO (the tighter, internal number), which means by the time anyone notices, the team is already close to breaching the external promise with no margin left to react.
- An SLO with no error-budget policy attached is just a number on a dashboard; the value comes from what happens when the budget is nearly exhausted (freeze risky launches, redirect engineering time to reliability work), not from the target itself.
- Picking an SLI that doesn't reflect real user experience (e.g. "server process is running" instead of "requests succeed within an acceptable time") gives a green dashboard while users are still unhappy. The SLI has to be as close to what the user actually experiences as the team can measure.
Your team is choosing between a hosted observability platform and a self-hosted open-source stack for a growing company with a small operations team. Walk through the trade-offs you'd weigh, things like cost, operational overhead, feature completeness, and vendor lock-in, and what would tip your recommendation one way or the other.
Sample Answer
Direct answer
For a growing company with a small operations team, default to a hosted platform unless one team already has strong operational muscle for a specific piece of the stack (most often metrics, via Prometheus). The variable that actually decides this is engineer-hours available for care and feeding, not sticker price: a self-hosted stack usually undercuts hosted pricing on paper, but that gap closes or reverses once you count the SRE time spent on cluster capacity, upgrades, and retention tuning.
Decision framework
| Dimension | Hosted (Datadog, New Relic, Grafana Cloud style) | Self-hosted OSS (Prometheus + Loki + Grafana + Tempo) |
|---|---|---|
| Upfront cost | Low, pay-as-you-ingest | Low licensing, but infra plus engineer time is a real cost |
| Ongoing cost at scale | Grows fast with hosts/ingestion, can dominate the infra bill | Grows with storage/compute you already control, more linear |
| Operational overhead | Near zero: vendor handles scaling, upgrades, HA | Real: cluster sizing, upgrades, backup/restore, on-call for the observability stack itself |
| Feature completeness | Turnkey APM, anomaly detection, log parsing UIs, SLO tooling out of the box | Comparable core signal collection, but polish (auto root-cause, ML anomaly detection) usually lags or needs extra tooling |
| Vendor lock-in | Real: proprietary query language, dashboards don't port cleanly | Low: OpenTelemetry, PromQL, and LogQL are portable across backends |
| Multi-tenant RBAC (role-based access control: who can see/query which logs) / PII redaction (log aggregation) | Usually built in (SSO, field-level masking) as a paid-tier feature | You build and maintain it yourself (access policies at the query layer, redaction at the log shipper) |
When to tip toward each:
- Hosted: team is small, time-to-value matters more than unit cost, nobody owns the observability stack as their primary job, or you need APM/anomaly detection you don't want to build yourself.
- Self-hosted: you already run Kubernetes at scale with a platform team that can treat observability as just another workload, data residency or compliance forces on-prem storage, or ingestion volume is high enough that hosted per-GB pricing becomes the single largest line item in the infra budget.
- Hybrid: a common middle path is hosted APM and log search (where turnkey correlation and UI matter most) paired with a self-hosted Prometheus for metrics you already understand and want fast, cheap, high-resolution queries on. This works because metrics are the cheapest and most mechanical piece to self-host, while tracing and log search UX is where vendors differentiate most. Re-evaluate the split as the team and ingestion volume grow.
Worked example
The crossover point between the two options can be derived, not guessed, once you write cost as a function of ingestion volume.
Let G be GB/day ingested. Model hosted cost as a per-GB ingestion rate, and self-hosted cost as a fixed infra floor plus a much cheaper per-GB storage rate plus ongoing engineer-hour labor:
Gbreak−even=30×(rhosted−rstorage)Cinfra+H⋅RlaborUsing illustrative, explicitly-assumed rates (not any specific vendor's current published price, since list prices change): hosted ingestion rate rhosted=$0.10/GB, self-hosted object-storage rate rstorage=$0.02/GB, fixed self-hosted compute floor Cinfra=$300/month, and H=4 hours/month of engineer time at a loaded rate Rlabor=$150/hour:
Gbreak−even=30×(0.10−0.02)300+4×150=2.4900=375 GB/dayUnder these assumptions, below roughly 375 GB/day of ingestion, hosted comes out cheaper on pure dollar terms even before counting the small-team operational risk. Above it, self-hosting's lower per-GB rate starts to outweigh the fixed compute and labor floor. The point isn't the exact number, it's that a candidate should reason in terms of where the cost curves cross, not assert a universal winner.
Trade-offs and pitfalls
- Sunk-cost fallacy: don't keep self-hosting because it was already built, once engineer time becomes scarce elsewhere the calculus can flip.
- Hidden costs of hosted: egress and API costs, per-seat pricing for dashboard users, and price increases once you're locked in and migration feels expensive.
- Hidden costs of self-hosted: the observability stack becomes a second production system that itself needs monitoring, on-call, and capacity planning. If it goes down during an incident, you're debugging blind.
- Migration cost is asymmetric: moving off a hosted platform later means rebuilding dashboards and alerts in a new query language; moving off self-hosted OSS is comparatively easier since OpenTelemetry-based data is portable to almost any backend.
- Common wrong turn: picking self-hosted purely because a spreadsheet says it's cheaper while ignoring engineer-hours, then rediscovering the true cost six months later at the first major version upgrade.
How would you set a naming and labeling convention for metrics across a multi-team organization, so ownership is discoverable and cardinality stays under control?
Sample Answer
Direct answer
A good metric naming and labeling convention is really an ownership and discoverability contract enforced through structure: names encode what's being measured and its unit so anyone can guess the meaning without documentation, and a small, mandatory set of labels makes every metric traceable back to an owning team, while everything else about label design is governed by an explicit cardinality budget rather than left to individual judgment.
The convention
Naming
- Pattern:
<namespace>_<subsystem>_<metric>_<unit>, snake_case, with a standard suffix that encodes type:_totalfor counters,_secondsor_bytesfor a base unit,_bucket/_sum/_countfor histogram components. - Names are nouns describing what's measured, not actions:
http_requests_total, nottrack_http_requests. - Treat a shipped metric name as a stable interface. If the semantics change, version it (
http_request_duration_seconds_v2) rather than silently redefining what an existing name means underneath dashboards and alerts that already depend on it.
Labeling for ownership and discoverability
- Mandatory labels on every metric:
service,team,environment. This is what makes "who owns this metric" answerable by a query, not a wiki page that goes stale. - A small, curated set of dimension labels beyond that (
region,status_code,method) that are explicitly allow-listed, not left open for anyone to add whatever seems useful in the moment. - No identifier-shaped labels, anything meant to be unique per request or per user, ever. That class of label belongs to logs and traces, not metrics.
Scaling the convention across teams without a central bottleneck
- Publish the convention as a small schema or lint rule, a CI check that validates new metric names and labels against the pattern and the mandatory-label list, rather than a document people are expected to remember, so it's enforced automatically at the point where it's cheapest to fix.
- Give teams a self-service allow-list process for adding a new dimension label to an existing metric family, with a lightweight review, so the convention doesn't become a bottleneck people route around.
- Use the
teamlabel to build an automatic ownership directory, which team owns which metrics, surfaced in the metrics catalog, so "who do I ask about this" is a query, not tribal knowledge.
Worked example
A recording rule that only works because the convention was followed consistently, and what breaks otherwise:
groups:
- name: service_latency
rules:
- record: service:request_latency_seconds:p95
expr: |
histogram_quantile(0.95,
sum(rate(http_request_duration_seconds_bucket[5m])) by (service, team, le)
)
This rule only produces a sane per-team latency rollup because every team's http_request_duration_seconds histogram uses the same metric name, the same bucket boundaries, and includes the mandatory service and team labels. If one team named theirs http_req_duration_ms instead, a different name and a different unit with no convention applied, this single rule can't include their data at all. They'd need a one-off query, which is exactly the discoverability cost the convention exists to prevent.
Trade-offs and pitfalls
- A convention that's too strict, a rigid required schema for every possible dimension, slows teams down and invites workarounds, like encoding extra information into the metric name itself to route around a label restriction, which is worse than the problem it was meant to solve. Keep the mandatory set small and make the allow-list process fast.
- A convention that's purely documented and not enforced by CI or lint decays within a quarter as new services get built by people who never read the doc. Enforcement has to be mechanical, not aspirational.
- Common wrong turn: retrofitting a naming convention onto an existing large fleet all at once. A big-bang rename breaks every dashboard and alert built on the old names simultaneously. Version and dual-ship, old and new names in parallel for a deprecation window, instead.
You're asked to lead the migration of dozens of services from legacy free-form text logs to a centralized, structured logging standard. Walk through your rollout plan: how you'd approach schema design and discovery across teams, how you'd instrument and validate services in phases rather than a single cutover, and how you'd handle an organization spanning multiple accounts or teams without breaking existing dashboards.
Sample Answer
Direct answer
Treat this as a backward-compatible API migration, not a cutover: discover the current logging landscape and define a minimal common schema, ship instrumentation as a shared library with CI schema validation, roll out per-service in risk-ordered phases using dual-write (old text logs and new structured logs side by side) rather than a single flip, and only retire the old format for a service once its dashboards and alerts have been remapped and verified against the new fields.
Structured elaboration
Discovery & schema design
Inventory every log producer, what format it emits today, and every consumer that depends on it (dashboards, alert rules, saved searches, downstream ETL). Design a minimal common schema (timestamp, severity, service name, environment, trace/span id, a small set of required business keys) plus a per-service extension namespace for fields that don't generalize. Review the schema with the teams who'll produce and consume it before writing any code; a schema designed in isolation gets rejected or worked around later.
Instrumentation tooling
Ship one shared library per language (not per team) that emits the schema correctly by construction, so individual engineers don't hand-write JSON log lines. Add a CI check that validates a sample of a service's emitted logs against the schema and blocks merges that regress it. This is what makes "structured logging" durable instead of a one-time cleanup that decays as new code gets added.
Phased rollout with dual-write
Group services by blast radius (customer-facing/critical, internal/important, low-risk/batch) and roll out lowest-risk first. For each service: instrument in dual-write mode (both the legacy text format and the new structured format emit in parallel), run for a defined validation window, confirm the structured stream reproduces what the legacy dashboards showed, then cut dashboards and alerts over to the structured fields, and only then stop emitting the legacy format. Dual-write is what prevents "the migration broke the on-call dashboard" incidents; it costs extra log volume for the overlap window, which is a deliberate, bounded trade.
Cross-account / cross-team handling
When the org spans multiple accounts, put the durable structured store in a central logging account rather than per-account, with each source account's service shipping into it via a scoped, least-privilege role (write-only into its own prefix, no read access to other accounts' data). This means an account's migration schedule doesn't block or get blocked by another account's, and a central index/search layer can serve org-wide dashboards without waiting for every account to finish.
Verification before cutover
Before retiring the legacy format for a service, run both streams in parallel and diff their derived metrics (error counts, request volumes by endpoint) for a validation window; only cut over once they agree within an expected tolerance. Skipping this step is how migrations quietly blind an alert without anyone noticing until an incident happens with no signal.
Worked example
Say the fleet is 60 services, grouped into 3 risk tiers of 20 each (low, medium, high blast radius). Assume (a stated planning assumption, not a timing claim about system performance) each tier needs a minimum 3-day dual-write validation window per service before cutover, and tiers are rolled out sequentially rather than in parallel to limit concurrent risk:
total validation calendar time≥3 tiers×3 days=9 daysThat's the floor if everything in a tier could validate in parallel; in practice services within a tier don't all reach "verified" on the same day, so the real plan should budget more like 2-3 weeks per tier once staggered starts and remediation cycles are accounted for. The point of doing this arithmetic explicitly in the plan is to give stakeholders an honest, assumption-labeled timeline rather than a single unqualified date.
flowchart LR
A[Service - Account A] -->|dual-write raw+structured| S1[Local agent]
B[Service - Account B] -->|dual-write raw+structured| S2[Local agent]
S1 --> X[Cross-account log shipper]
S2 --> X
X --> C[Central logging account]
C --> R[Raw archive - object storage]
C --> IDX[Search / index layer]
IDX --> D[Dashboards remapped to structured fields]
Trade-offs & pitfalls
- Dual-write roughly doubles log volume and ingest cost for every service during its validation window; keep that window as short as reliably possible rather than leaving it open indefinitely.
- Per-account IAM roles for cross-account shipping can sprawl into an unmanageable mess if each account hand-rolls its own; define one reusable role/policy template and apply it uniformly.
- The biggest hidden risk isn't the instrumentation, it's forgetting to remap a dashboard or alert rule before retiring the legacy stream; the phase order (validate, remap, verify, then retire) exists specifically to prevent silently blind alerting.
- A shared library that isn't versioned and centrally maintained will drift across services just like hand-written logging did; treat it as a product with an owner, not a one-time handoff.
Unlock Full Question Bank
Get access to all 11 Monitoring, Logging, and Observability interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.