Observability and Monitoring Architecture Questions
Building visibility into infrastructure and services: metrics, logs, and traces, dashboards and alerting, SLIs/SLOs, and the design of an observability stack. Covers instrumenting systems for actionable signal, reducing alert noise, and diagnosing production issues from telemetry. Infrastructure-wide observability, distinct from network-specific monitoring.
Define SLOs for the observability pipeline itself: ingestion availability, storage durability, query freshness, and end-to-end latency for a dashboard query. For each one, propose a concrete target, how you'd measure it, and what action fires when it's breached.
Sample Answer
Direct answer
Define one SLI (service level indicator, the measured signal used to judge whether a promise is being kept) per user-facing promise the observability pipeline makes: ingestion availability (did the write succeed), storage durability (did an acknowledged write survive), query freshness (how long until an ingested event is queryable), and end-to-end query latency (how fast a dashboard panel renders). Set a concrete target and measurement method for each, and wire alerting to error-budget burn rate rather than a raw threshold, so a brief severe outage and a sustained mild degradation both page appropriately instead of only one of them tripping an alert.
Structured elaboration
For each SLI: target, how it's measured, and what fires on breach.
| SLI | Target (30-day SLO) | Measurement | Breach action |
|---|---|---|---|
| Ingestion availability | 99.9% of write requests accepted | Synthetic canary producers every 10s + real gateway accept/reject counters | Auto-scale ingestion frontends; page if burn rate crosses fast-window threshold |
| Storage durability | 99.999% of acknowledged writes recoverable | Canary writes read back on a schedule; replication-lag monitoring | Fail over reads to a healthy replica; escalate if canary reads fail |
| Query freshness | p99 within 5 minutes of event time | A freshness consumer that tracks time-to-visible per event ID | Scale indexing workers; apply ingest backpressure to protect freshness for higher-priority signals |
| End-to-end query latency | p95 ≤ 500ms for canonical dashboard queries | Synthetic dashboard runs every minute against the canonical query set + real user latency traces | Enable/expand query result caching; shed non-critical query load |
Why error-budget burn rate, not a flat threshold: a flat "alert if availability drops below 99.9% this hour" either pages on noise (one bad minute in an otherwise fine hour) or misses a slow, sustained degradation that never dips low enough in any single short window to trip the threshold, while still burning through the monthly budget. A burn-rate alert asks a different question: "at this rate, how much of the month's error budget will this consume, and how fast?"
flowchart LR
G[Ingestion gateway] -- accept/reject --> M1[Availability SLI]
ST[(Storage)] -- canary read-back --> M2[Durability SLI]
IDX[Indexer] -- time-to-visible --> M3[Freshness SLI]
QE[Query engine] -- synthetic + real latency --> M4[Latency SLI]
M1 & M2 & M3 & M4 --> BR{Burn-rate calc}
BR -- fast window, high burn --> PAGE[Page on-call]
BR -- slow window, moderate burn --> TICKET[Ticket, review]
Worked example
Error budget for the ingestion availability SLO. At 99.9% over a 30-day period:
allowedErrorRate=1−0.999=0.001 errorBudgetMinutes=0.001×(30×24×60)=0.001×43,200=43.2 minutes/monthDeriving the multi-window burn-rate thresholds from first principles. A burn rate of b means the pipeline is failing b times faster than the allowed rate. If we want to alert fast when a burn rate would consume some fraction of the full 30-day (720-hour) budget within a short window:
burnRateThreshold=fractionOfBudget×windowHoursperiodHoursFor a fast window (page if 2% of the monthly budget would burn in 1 hour):
bfast=0.02×1720=14.4For a slower window (ticket if 5% would burn in 6 hours):
bslow=0.05×6720=6.0Applying it to an actual incident. Suppose the ingestion gateway starts rejecting 5% of writes (an actual observed error rate of 0.05 against an allowed rate of 0.001):
observedBurnRate=0.0010.05=50×Since 50>14.4, this immediately trips the fast-window page. It also tells you exactly how much runway remains before the entire month's error budget is gone at this rate:
hoursToExhaustBudget=50720=14.4 hoursThat's the number that should drive urgency in the room: at a 5% failure rate, this SLO's entire 30-day error budget is gone in 14.4 hours if nothing changes, which is why it pages immediately rather than waiting for a longer observation window.
Trade-offs & pitfalls
Common wrong turns: setting every SLI's target based on what current performance already achieves rather than what downstream consumers actually need (an SLO that's already being met with zero margin gives you no error budget to spend on legitimate risk, like a risky deploy); alerting on a single-window threshold instead of the fast/slow burn-rate pair, which either misses slow-burn incidents or pages too eagerly on noise; and treating "storage durability" and "ingestion availability" as the same SLI, when a write can be accepted (available) and still be lost before it's durable, so they need independent measurement even though they're both about "did the write succeed" in casual conversation. Also worth stating plainly in the room: these SLOs are for the observability platform's own reliability, which is a different (and often tighter) bar than the SLOs of the services it's monitoring, since a degraded observability platform blinds you during exactly the incidents where you need it most.
Architect an observability pipeline that has to ingest and store genuinely high-cardinality metrics, millions of unique series, under a tight cost ceiling. Describe your strategy across label reduction, aggregation and rollups, sampling, storage tiering, and retention, and how you'd let someone reconstruct finer-grained detail at query time when they actually need it.
Sample Answer
The architecture is a pipeline of five stages, each trading some fidelity for cost, plus a reconstruction path at query time for the cases where someone genuinely needs the detail back. The stages: label reduction, aggregation and rollups, sampling, storage tiering, and retention.
Pipeline
flowchart LR
A[Raw High-Cardinality Metrics] --> B[Label Reduction]
B --> C[Aggregation and Rollups]
C --> D[Sampling]
D --> E[Storage Tiering]
E --> F[Retention Enforcement]
F --> G[Query Layer]
G --> H[On-demand Rehydration: top-K exemplars]
- Label reduction: the single highest-leverage step. Cardinality is combinatorial across independent label dimensions: if
endpointhas 80 distinct values,statushas 6,podhas 800, andregionhas 4, the bounded series count is their product.
Adding one more label with effectively unbounded cardinality (like a raw request_id, close to one unique value per request) multiplies this by request volume instead of by a fixed factor, turning a bounded 1.5M-series metric into an unbounded one. Label reduction means identifying and removing or bucketing exactly those unbounded dimensions before ingestion, keeping the combinatorially-bounded ones.
- Aggregation and rollups: pre-compute coarser-grained series (per-service instead of per-pod, for example) at ingestion time so most dashboard and alert queries never touch the raw high-cardinality series at all.
- Sampling: for label values that are useful individually but too numerous to keep in full (e.g., per-customer-ID series for a B2B product with thousands of customers), keep exact series for the top-K by volume or spend, and sample or bucket the long tail.
- Storage tiering: recent raw data on fast storage, older data downsampled and moved to cheaper storage, as in a standard hot/warm/cold retention policy.
- Retention enforcement: hard expiry so the pipeline's cost stays bounded over time regardless of how ingestion volume trends.
Reconstructing detail on demand
The trick that makes this acceptable to users is that "finer detail" doesn't mean "keep everything forever," it means "keep enough breadcrumbs to go get the detail from a cheaper source when someone actually asks." Two mechanisms:
- Exemplars: attach a sampled trace ID or raw event reference to an aggregated metric bucket, so a spike in the aggregate can be drilled into by fetching the handful of exemplar traces/logs that were kept in full, even though the metric itself was aggregated.
- On-demand reprocessing: if raw pre-aggregation data still exists in a cheap cold tier (e.g., unindexed compressed blobs), a rare deep-dive query can trigger an offline reprocessing job rather than requiring the hot path to keep everything queryable in real time.
Sizing against a cost ceiling
Given an illustrative monthly storage budget of $5,000 and an illustrative unit cost of $0.02/GB-month (a labeled assumption for this worked example, not a live vendor quote):
max steady-state storage=0.025,000=250,000 GB=250 TBUsing the 1,536,000 bounded series from the label-reduction step above, at a 5-minute downsampled resolution with 8 bytes/point (consistent with the multi-aggregate downsampling estimate used for retention-tier design):
series = 1_536_000
bytes_per_point = 8
interval_s = 300
points_per_day = 86400 / interval_s
bytes_per_day = series * points_per_day * bytes_per_point # 3.539 GB/day
max_days = (250_000 * 1e9) / bytes_per_day
That affords roughly 70,600 days of 5-minute-resolution history under the budget, which is obviously far beyond any real retention need, so the budget is not actually the constraint at this series count and resolution. Repeating the same calculation for raw 15-second resolution at 2 bytes/sample instead gives 17.69 GB/day and about 14,128 affordable days, roughly 5x shorter than the downsampled case (the 20x fewer points at 5-minute resolution is partly offset by needing 4x more bytes per point to store min/max/sum/count instead of a single value, netting exactly 20/4 = 5x). The concrete lesson: at this series count, the budget comfortably covers years of retention either way, so cost pressure at 1.5M series is not what forces sampling or tiering, it forces label reduction to happen so the series count never gets to the point where the arithmetic above breaks down (e.g., adding the unbounded request_id label would blow past 250 TB in a matter of hours).
Trade-offs and pitfalls
- Treating sampling as the first line of defense (instead of label reduction) is the most common mistake: sampling a metric whose cardinality is unbounded because of a labeling error still leaves an unbounded number of series, each just sampled less; the series count itself, not just the sample rate, is what needs to be bounded first.
- Aggregation destroys the ability to answer "which specific instance caused this" without exemplars; a design that aggregates without keeping any drill-down path trades away debuggability that's expensive to get back later.
- A fixed cost ceiling naturally reframes the problem: it's not "how do we store everything cheaper," it's "what data can we afford to keep at what fidelity," and that framing should drive which stage of the pipeline (label reduction vs. sampling vs. tiering) absorbs the cost pressure, since as shown above they don't interact linearly.
- Query-time reconstruction only works if the cheap cold tier is actually queryable, even slowly; if cold data is written in a format nothing can read without a bespoke recovery process, "reconstruct on demand" is really "data is gone" with extra steps.
Set a concrete retention and downsampling policy for metrics and traces that balances cost against query fidelity, for example raw metrics for 14 days, downsampled metrics for a year, full traces for 30 days then sampled. Walk through your rationale and what it means for the kinds of queries you can still answer after each window closes.
Sample Answer
Set the policy by working backward from what each query pattern actually needs, then verify the storage savings with the arithmetic rather than picking round numbers and hoping. A reasonable concrete policy: raw metrics at native resolution for 14 days, 5-minute rollups for 1 year, hourly rollups for years 2 through 5; full traces for 30 days, then 1% sampled for the following 11 months.
Rationale by window
flowchart LR
A[Raw ingest: native res] -->|14 days| B[Raw tier: hot]
B -->|downsample| C[5-min tier: 1 year]
C -->|downsample| D[Hourly tier: years 2-5]
D -->|expire| E[Deleted]
F[Trace ingest] -->|30 days full| G[Full trace tier]
G -->|sample 1%| H[Sampled trace tier: 11 months]
- 14 days raw: covers essentially all incident debugging, since almost every retro or root-cause investigation happens within two weeks of the event, and alerting needs full resolution on recent data to avoid missing short spikes.
- 1 year at 5-minute rollups: supports capacity planning and seasonal comparisons (week-over-week, month-over-month) without needing per-second precision; 5 minutes is short enough to still show diurnal patterns clearly.
- Years 2-5 at hourly rollups: supports long-term trend and year-over-year growth analysis; anything finer than hourly at this age is rarely queried and expensive to keep.
- 30 days full traces: matches the raw-metrics window for the same reason, full-fidelity root cause work happens fast, and traces are the most expensive telemetry type per unit.
- 1% sampled for 11 more months: preserves enough statistical signal for "did this class of error exist a few months ago" investigations without paying for full trace volume; always retain 100% of traces tied to errors or SLO breaches regardless of the sampling rate (a fixed-percentage sample can otherwise miss the rare traces investigators actually want).
Verifying the storage savings
For 1,000,000 active series, using the same 2-bytes/compressed-raw-sample and 8-bytes/downsampled-point (4 aggregates: min, max, sum, count, at roughly 2 bytes each) assumptions used elsewhere in TSDB capacity planning:
series = 1_000_000
compressed_bytes_per_raw_sample = 2
agg_bytes_per_downsampled_point = 8
def samples(days, interval_s):
return series * (days * 86400 / interval_s)
raw_bytes = samples(14, 15) * compressed_bytes_per_raw_sample
ds_bytes = samples(365, 300) * agg_bytes_per_downsampled_point
hourly_bytes = samples(1460, 3600) * agg_bytes_per_downsampled_point
total_tiered_bytes = raw_bytes + ds_bytes + hourly_bytes
allraw_bytes = samples(14 + 365 + 1460, 15) * compressed_bytes_per_raw_sample
Result: raw tier = 161.28 GB, 5-min tier = 840.96 GB, hourly tier = 280.32 GB, total tiered storage over the full 5-year window ≈ 1.283 TB, versus an all-raw-forever equivalent of ≈ 21.19 TB for the same window, a 16.5x reduction. The 5-minute tier dominates total storage (840 GB of the 1.28 TB) precisely because it covers the most time (1 year) at the finest surviving resolution; that's useful to know when deciding whether to push the raw window shorter or the 5-minute window's resolution coarser if the budget gets tighter.
For traces, at 10,000 traces/sec with an assumed 4 KB compressed size per trace: full 30-day retention stores about 103.68 TB, while the following 11 months at 1% sampling adds roughly 11.40 TB, so the sampled tail costs about 11% as much as the initial 30-day full window despite covering over 10x the time span.
What you can and can't still answer after each window closes
| Window | Still answerable | No longer answerable |
|---|---|---|
| After 14 days (raw metrics gone) | Was there a sustained regression this week vs. last month, at 5-minute granularity | Exact second-level spike shape of an incident 3 weeks ago |
| After 1 year (5-min rollups gone) | Year-over-year seasonal comparison at hourly granularity | Any sub-hour pattern from 13 months ago |
| After 30 days (full traces gone) | Error-tagged and SLO-breach traces remain at full fidelity indefinitely (by policy) | A specific successful request's full span tree from 6 weeks ago, unless it happened to fall in the 1% sample |
| After 11 months (sampled traces gone) | Aggregate error-rate and latency-percentile trends from metrics, which persist far longer than traces | Any trace-level detail at all from over a year ago |
Trade-offs and pitfalls
- Percentile aggregates (p95, p99) do not survive naive downsampling: averaging five 1-minute p99 values is not the same number as the true p99 across that 5-minute window. If percentile fidelity matters at the rollup tier, you need to store enough of a histogram or sketch (not just min/max/sum/count) to recompute percentiles, which raises the per-point byte cost above the 8-byte assumption used here.
- A fixed sampling percentage for traces (1% flat) will statistically under-represent rare-but-important request types unless it's stratified or combined with the always-keep-errors/SLO-breach rule; a pure random sample optimizes for "typical" traffic, which is exactly what you don't need for debugging.
- Retention policy that isn't enforced automatically (a manual cleanup job, or "we'll get to it") tends to silently become the accidental real retention policy; tie expiry to the storage engine's native TTL/compaction mechanism rather than a side script.
- Communicating the policy to teams matters as much as setting it: if engineers don't know that a 3-week-old spike is only visible at 5-minute resolution, they'll draw wrong conclusions from a smoothed-out graph without realizing detail was lost.
A regulated customer requires that PII never lands in raw telemetry. Design an end-to-end pipeline that detects and redacts PII at ingestion while preserving enough context to debug production issues. Cover detection techniques (regex versus ML classifiers), whether masking is deterministic or tokenized, which enforcement point you'd use (agent, collector, or storage), the performance cost, and how you'd prove it's working to an auditor.
Sample Answer
Direct answer
Detect PII with a fast deterministic layer (regex/pattern rules for structured PII like emails, SSNs, phone numbers) combined with an ML classifier for unstructured or contextual PII (names, addresses in free text), and treat neither as sufficient alone. Enforce redaction at the collector, not the agent or the storage layer alone: agent-side redaction is best for privacy but hardest to update and audit across a fleet, storage-side is the last line of defense but means raw PII already crossed the network. Use deterministic, salted tokenization (not plain masking) so debugging can still correlate the same underlying value across events without ever storing or logging the raw value, and produce an immutable audit record for every redaction decision so you have concrete evidence for an auditor, not just a policy document claiming it happens.
Structured elaboration
Detection techniques:
- Regex/pattern rules: fast, deterministic, good for structured PII with a known shape (email, SSN, credit card, IP). Cheap enough to run on every event inline.
- ML classifiers: needed for unstructured, contextual PII (a name or address embedded in free-text log messages) that no fixed pattern reliably catches. Higher recall on the hard cases, but adds latency and needs periodic retraining/evaluation as it can drift.
- Hybrid: run regex first (catches the bulk of structured PII cheaply); route ambiguous or high-risk fields to the ML classifier as a second pass.
Masking versus tokenization:
- Deterministic tokenization (HMAC, a one-way keyed hash function, of the value with a per-tenant secret salt): the same raw value always maps to the same token, which preserves the ability to correlate "this is the same user across these events" for debugging, without the token being reversible back to the raw value.
- Reversible tokenization (format-preserving encryption, encryption whose output keeps the same shape and length as the input, via a KMS-backed vault): needed only when an authorized party may legitimately need the original value back, gated behind strict access control and its own audit trail, separate from the deterministic path used for ordinary debugging.
- Partial masking (
j***e,a****@domain.com): preserves a debugging hint about shape/length without preserving correlatable identity; use where correlation isn't needed, just "this field had a name-shaped value."
Enforcement point, and why it's a layered decision, not a single choice:
| Enforcement point | Privacy strength | Update/audit cost | Role |
|---|---|---|---|
| Agent | Strongest: raw PII never leaves the host | Hardest: rules ship per-agent, slow to patch a false negative | First line, not sole line |
| Collector | Strong, centralized | Easiest: one place to update rules/models | Primary enforcement point |
| Storage-layer guard | Weakest on its own, raw PII already transited the network | Cheap to add as a backstop | Last line of defense, catches what leaked past the collector |
Auditability: every redaction decision writes an immutable audit record: input hash (never the raw value), detection method, confidence score, rule/model version, timestamp, event ID. This is what turns "we redact PII" into something an auditor can actually verify happened, on a specific event, using a specific rule version, rather than a claim about the system's design.
flowchart LR
A[Agent] -- light regex mask --> C[Collector]
C --> RX[Regex pass]
RX -- ambiguous field --> ML[ML classifier pass]
RX -- clear match --> TOK[Deterministic tokenization]
ML --> TOK
TOK --> ST[(Redacted storage)]
TOK --> AUD[(Immutable audit log)]
ST -- storage-level guard --> ST
Worked example
Combined detection recall. Assume the regex layer alone catches 92% of true PII instances (illustrative assumption), and the ML secondary pass catches 70% of what the regex layer missed:
regexMissRate=1−0.92=0.08 combinedMissRate=0.08×(1−0.70)=0.08×0.30=0.024 combinedRecall=1−0.024=0.976(97.6%)At an assumed 10,000,000 PII-bearing fields/day across the fleet:
residualLeakPerDay=10,000,000×0.024=240,000 fields/day still undetectedThis is the number that justifies the storage-level guard as a real requirement rather than a nice-to-have: even a well-tuned two-stage detector with 97.6% combined recall still lets 240,000 fields/day through in this example, which is why enforcement can't stop at the collector alone for a regulated customer.
Audit log storage overhead. Using the pipeline's ingestion baseline of 200,000 events/sec, assume 15% of events trip at least one detector:
auditEventsPerSec=200,000×0.15=30,000At 120 bytes/audit record (hash + method + confidence + rule/model version + event ID):
auditBytesPerDay=30,000×86,400×120 bytes=311.04 GB/dayThree hundred gigabytes a day of audit trail alone is a real storage line item, not a footnote, worth surfacing explicitly when scoping this design rather than discovering it after the audit log has been running unmanaged for a month.
Trade-offs & pitfalls
| Approach | Latency cost | Coverage |
|---|---|---|
| Regex only | Sub-millisecond, negligible | Misses contextual/unstructured PII entirely |
| Regex + ML hybrid (this design) | ML pass adds real but bounded per-field cost, offloadable to an async path for non-blocking fields | 97.6% combined recall in the worked example above |
| Storage-level guard only | Cheapest to add | Raw PII already transited agent-to-collector network in the clear until this point |
Common wrong turns: treating any single detection layer as sufficient and skipping the residual-leakage math above, which is what a regulated auditor will actually ask for ("what's your false-negative rate, and what's the backstop"); using reversible tokenization everywhere by default because it's more flexible, when most debugging only needs deterministic correlation and every reversible token is a standing risk that needs its own access control and audit trail; and logging the audit record with the raw value "just in case," which defeats the entire point of redaction by creating a second place PII lives, this time inside the system meant to prove PII doesn't land anywhere.
Design the mechanics of tail-based sampling at real scale, say 100,000 traces per second: spans have to be buffered somewhere until the sampling decision can be made, slow or erroneous traces need their full span set captured, and everything else gets thinned. How do you coordinate that buffering and decision-making across many collector instances without unbounded memory growth?
Sample Answer
The key mechanic is consistent-hash routing by trace_id at the load balancer, so every span belonging to a given trace lands on the same collector instance and no cross-collector coordination is needed to assemble a trace before deciding whether to keep it. Memory is then bounded with a fixed decision window plus per-shard and per-trace caps, not by trying to hold every trace indefinitely.
Coordination architecture
flowchart LR
A[Application Spans] --> B[Load Balancer: hash by trace_id]
B --> C[Collector Shard 1: buffer]
B --> D[Collector Shard 2: buffer]
B --> E[Collector Shard N: buffer]
C --> F[Sampling Decision Engine]
D --> F
E --> F
F -->|keep| G[Export Full Trace]
F -->|drop| H[Discard]
F -->|timeout| I[Partial-Trace Fallback]
Because routing is consistent-hash on trace_id, adding or removing collectors only reshuffles a small fraction of trace-to-collector assignments (standard consistent-hashing property), so scaling the fleet doesn't require a coordinated rebalance of in-flight traces.
Sizing the buffer
Take the stated 100,000 traces/sec, an average of 20 spans/trace (a typical microservice call depth), an average compressed span size of 500 bytes (a labeled assumption), and a 10-second decision window (wait up to 10 seconds after a trace's apparent last span before deciding, which covers the large majority of trace completion times):
traces_per_sec = 100_000
avg_spans_per_trace = 20
avg_span_bytes = 500
decision_window_s = 10
num_collectors = 50
span_rate = traces_per_sec * avg_spans_per_trace # 2,000,000 spans/sec
spans_buffered_systemwide = span_rate * decision_window_s # 20,000,000 spans
bytes_buffered_systemwide = spans_buffered_systemwide * avg_span_bytes # 10 GB
spans_per_collector = spans_buffered_systemwide / num_collectors # 400,000 spans
bytes_per_collector = spans_per_collector * avg_span_bytes # 200 MB
At 50 collectors, each instance buffers about 400,000 spans (200 MB), a footprint that fits comfortably in a modest container (2-4 GB), while the system-wide live buffer is about 10 GB spread across the fleet. This is the concrete argument for sharding by trace_id: a single collector holding the full 10 GB buffer would need a memory profile most container platforms would flag as oversized, while 50 shards each holding 200 MB is unremarkable.
Bounding memory beyond the happy path
The 10-second window handles typical traces, but slow or stuck traces need explicit handling so they don't grow the buffer without bound:
- Per-trace TTL: destroy a trace's buffer if no new span arrives within some multiple of the decision window (e.g., 2x), forcing a decision (keep as partial, or drop) rather than waiting indefinitely.
- Per-shard memory cap with eviction: each collector enforces a hard memory ceiling; if exceeded, evict the lowest-priority buffered traces first (e.g., traces with no error/latency signal yet) rather than failing open.
- Global admission control: a lightweight control-plane process aggregates each collector's buffer occupancy and kept-rate on a slow control loop (seconds, not per-request) and adjusts the decision window or sampling probability fleet-wide if the system is trending toward the memory ceiling, rather than each collector reacting in isolation and potentially over-correcting.
Decision logic
- Cheap deterministic rules first (status code indicates an error, latency exceeds a fixed threshold): mark "must keep" immediately without waiting for the full window, since these are unambiguous.
- For everything else, wait out the decision window, then apply either a lightweight scoring model (feature-based, comparing this trace's shape against a recent rolling baseline) or a straightforward probabilistic sample at a rate tuned to the fleet-wide keep-rate budget.
- On TTL expiry before a full decision, export whatever spans were captured as a partial trace rather than silently dropping everything; a partial error trace is still more useful for incident response than nothing.
Trade-offs and pitfalls
- The most common design mistake is trying to coordinate the sampling decision across collectors (e.g., a central service that all collectors ask before deciding); at 2,000,000 spans/sec that coordination service becomes the bottleneck. Consistent-hash-by-
trace_idavoids this entirely by guaranteeing the decision can be made locally. - A fixed decision window is a trade-off, not a free parameter: too short and slow-but-successful traces (a legitimately slow but non-erroring downstream call) get truncated into partial traces; too long and the buffer grows for no benefit on traces that were always going to be dropped.
- Eviction policy under memory pressure needs to bias toward keeping traces that already show error/latency signal; a naive LRU eviction can evict exactly the traces you most want to keep just because they arrived earlier.
- Rebalancing collector count changes which shard owns which traces going forward, but in-flight traces already buffered on their original shard need to either finish there or be explicitly drained; a hash-ring change that silently orphans in-flight buffers loses those traces' decisions.
Unlock Full Question Bank
Get access to all 41 Observability and Monitoring Architecture interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.