Event-Driven Architecture and Asynchronous Messaging Questions
Designing systems around events and message passing: publish/subscribe, message queues, event streaming, choreography versus orchestration, and decoupling producers from consumers. Covers delivery semantics (at-least-once, at-most-once), ordering, backpressure, dead-letter handling, and the operational tradeoffs of asynchronous flows. Includes async processing patterns for offloading slow work.
Compare synchronous REST and asynchronous messaging for inter-service communication. For each approach explain: failure semantics, coupling, observability, latency, consistency guarantees, and operational complexity. Provide examples of when you'd prefer one over the other.
Sample Answer
Direct answer
Synchronous REST and asynchronous messaging differ on every axis that matters for inter-service communication: failure semantics, coupling, observability, latency, consistency guarantees, and operational complexity. REST gives an immediate, explicit success/failure answer at the cost of temporal coupling; asynchronous messaging decouples caller and callee at the cost of an implicit, eventually-resolved answer. Prefer REST when the caller needs a definitive result to proceed; prefer messaging when the interaction can tolerate delay and the two sides should be able to fail, scale, and deploy independently.
Structured elaboration
Take each axis for both approaches directly:
| Axis | Synchronous REST | Asynchronous messaging |
|---|---|---|
| Failure semantics | Caller gets an explicit HTTP status (2xx/4xx/5xx) or a timeout; failure is immediate and attributable to one call. | Failure is implicit and delayed: a message can fail after N retries and land in a dead-letter queue (DLQ), discovered later by whoever monitors the DLQ, not by the original producer. |
| Coupling | Temporal coupling (both services must be up simultaneously) plus contract coupling to the exact response shape and latency. | Only schema coupling to the message contract; producer and consumer need not be online at the same time, and consumers can be added without producer changes. |
| Observability | A single request/response pair is easy to trace with standard distributed tracing (a request ID follows one call stack). | Requires correlating a message across an asynchronous boundary (a shared correlation ID propagated through message headers) and tracking consumer lag, queue depth, and DLQ growth as first-class signals, since "is this working" is no longer visible in a single trace. |
| Latency | Bounded by the caller's timeout; the caller experiences the callee's latency directly, and a slow callee makes the caller slow. | The producer's latency is just "message accepted by the broker," typically single-digit milliseconds; end-to-end processing latency is decoupled from the producer and instead depends on consumer throughput and backlog. |
| Consistency guarantees | Strong consistency is achievable at the call boundary: the caller knows the callee's result before proceeding. | Eventual consistency: the caller only knows the message was durably accepted, not that the consumer has processed it; there is a window, bounded by consumer lag, during which downstream state has not caught up. |
| Operational complexity | Lower: no broker to run, fewer moving parts, well-understood tooling (load balancers, HTTP status codes, standard retries). | Higher: a broker (or managed equivalent) to operate or pay for, plus delivery-semantics decisions (at-least-once handling), consumer-side idempotency, ordering/partition-key design, and DLQ/retry policy to build and monitor. |
Worked example
A ride-hailing platform's "request ride" endpoint needs to synchronously call a pricing service, because the rider must see a confirmed price before confirming the ride; if pricing is slow or down, the request should fail fast rather than silently succeed with an unknown price, so REST with a tight timeout (e.g., 500ms) and a clear 503 on failure is the right fit here: failure semantics are explicit, and the caller cannot proceed without the answer. In contrast, once the ride completes, updating the driver's lifetime earnings dashboard and running fraud-pattern analysis on the trip do not gate anything the rider or driver is waiting on; publishing a "ride.completed" event lets those two consumers process independently, at their own pace, and a backlog in the fraud-analysis consumer (say, several minutes of lag during a traffic spike) has zero effect on ride completion, which is exactly the isolation asynchronous messaging is bought for.
Trade-offs and pitfalls
The common wrong turn is treating this as an architecture-wide choice ("we are a REST shop" or "we are event-driven") rather than a per-interaction decision on these six axes; most real systems need both, often for different steps of the same business flow. A subtler pitfall is picking asynchronous messaging for its scalability story while underestimating the observability tax: without correlation IDs threaded through message headers and consumer-lag/DLQ-depth dashboards from day one, an asynchronous flow that silently stalls is far harder to detect than a synchronous call that returns an explicit error, because nothing "fails" in a way that pages anyone; the failure just accumulates quietly as growing lag until a downstream SLA (service-level agreement) is missed.
In a shopping-cart checkout flow, decide which sub-steps should be synchronous (e.g., payment authorization) and which can be asynchronous (e.g., sending confirmation email, analytics). Explain how your choices affect user experience, system correctness, error-handling, and eventual consistency guarantees.
Sample Answer
Direct answer
In a checkout flow, keep synchronous only the sub-steps whose outcome the user or the next step must know before the transaction can be considered complete: inventory availability check and payment authorization. Everything that does not gate the "did this order succeed" answer, such as sending the confirmation email and recording analytics events, should be asynchronous, published as events once the order is durably created.
Structured elaboration
Apply a single test to each sub-step: "if this step fails or is slow, must the user's checkout fail or wait?" Payment authorization fails the test: if the card is declined, the order must not be placed, so it has to complete (or definitively fail) before the checkout response returns. Confirmation email and analytics both pass the test in the other direction: if the email provider is down, the order is still valid and the customer should not see an error; if the analytics pipeline is backed up, that has zero bearing on whether the customer got their item.
Effects of this split:
- User experience. The customer sees a fast, honest response: "order confirmed" as soon as payment clears, not "order confirmed, email pending" or a spinner while a marketing analytics call finishes. Moving email/analytics off the critical path directly lowers perceived latency, since the response no longer waits on the slowest of several unrelated systems.
- System correctness. Correctness now hinges only on the synchronous steps actually being atomic or safely retryable: payment authorization must be idempotent (a network retry must not double-charge) and the order record must not be marked "placed" unless payment is confirmed. The asynchronous steps cannot violate correctness of the order itself, because they consume from an event that is only published after the order is already valid; at worst, a failed email send means a customer does not get a receipt email, not that they have a wrong order.
- Error handling. Synchronous steps need explicit user-facing error handling (declined card, out-of-stock) because the user is waiting on the answer. Asynchronous steps need consumer-side error handling instead: retries with backoff, and a dead-letter queue (DLQ, a queue that holds messages a consumer could not process after exhausting retries) for a persistently failing email send, which an on-call engineer or automated remediation job can drain later without ever bothering the customer.
- Eventual consistency guarantees. The order itself is strongly consistent the moment checkout returns (it either happened or it did not). The order's "surrounding" state, such as "has the customer been emailed" or "has this purchase been counted in today's revenue dashboard," becomes eventually consistent: it will be true within some bounded window (seconds to low minutes, driven by consumer lag) but is not guaranteed true at the instant checkout returns. That gap needs to be a conscious guarantee you can state, not an accident: e.g., "confirmation email delivered within 5 minutes of order placement, monitored via consumer lag on the notifications topic."
Worked example
Sequence for a 75order:(1)synchronouslyreserveinventoryfortheSKUandauthorizepaymentfor75; if either fails, return an error to the user immediately and nothing else happens. (2) On success, write the order row and, in the same database transaction (or via the transactional outbox pattern, where an "OrderPlaced" row is written to an outbox table in the same commit and a separate relay publishes it), emit an "OrderPlaced" event. (3) Return "order confirmed" to the user at this point, without waiting on anything downstream. (4) A notifications consumer subscribed to "OrderPlaced" sends the confirmation email, retrying up to 3 times with backoff on transient send failures before landing the message in a DLQ. (5) An independent analytics consumer subscribed to the same event increments the day's revenue counter. Steps 4 and 5 run in parallel, are unaware of each other, and neither can block or fail step 1 through 3.
Trade-offs and pitfalls
The main pitfall is drawing the line by "which steps feel slow" rather than "which steps the user's success/failure outcome depends on"; a fast email send is still a UX and correctness bug if it is on the synchronous path, because it adds a dependency the checkout does not need. The opposite pitfall is making payment authorization asynchronous "to be consistent" with the rest of the flow: that forces the UI into an awkward "we'll email you when your payment clears" pattern for something users expect an immediate answer to, and it reopens the question of what state the order is in while payment is pending. A senior answer also flags that moving a step asynchronous introduces a durability requirement: if the order write and the event publish are not atomic, a crash between them can silently drop confirmation emails and analytics events for orders that did place successfully, which is exactly the failure mode the outbox pattern exists to close.
Define idempotency in the context of event-driven architectures. As a Solutions Architect, design a simple idempotency/deduplication strategy for an email-sending consumer that reads messages with payload {email_id, recipient, template}. Describe storage choices for dedup keys, TTL policies, memory vs disk trade-offs, and how to handle retries and long outage recovery.
Sample Answer
Direct answer
Idempotency in an event-driven system means processing the same message any number of times produces exactly the same effect as processing it once, which matters because message brokers commonly redeliver: a crash between processing and acknowledging, a network blip, or a producer retry can all cause the same logical message to arrive twice. For the email-sending consumer, that means checking a durable, atomic "have I already sent this email_id" record before every send, not trusting that the broker will only ever deliver a message once.
Structured elaboration
Dedup key
- Use
email_idas the deduplication key if oneemail_idalways corresponds to exactly one intended send; if a singleemail_idcould legitimately target multiple recipients, use(email_id, recipient)as a composite key instead. - Store a small record per key: status (in progress or sent) and when it was written.
Storage choices for dedup keys
- A fast in-memory store (a Redis-style key-value cache) for the hot path: low latency, high throughput, supports an atomic check-and-set operation needed to avoid a race between two consumer instances both trying to claim the same message.
- A durable, disk-backed store (a relational or key-value database) as the system of record: survives a full cache restart or a long outage, at higher latency than the in-memory tier.
- Hybrid: check the fast in-memory tier first; on a miss, fall back to the durable store before concluding the message is genuinely new, so a cache restart cannot cause a previously-sent email to be resent.
Time-to-live (TTL) policies
- The in-memory tier's entries expire after a TTL sized to the realistic redelivery window (long enough that any plausible retry or redelivery is still caught, short enough not to grow unbounded).
- The durable tier does not need a TTL for correctness (it is the long-term source of truth for "was this ever sent"), though it may have its own separate retention policy for storage cost or compliance reasons, independent of the dedup TTL.
Memory vs disk trade-offs
- Memory: fast, ideal for the common case (a redelivery arriving within seconds to minutes of the original), but limited capacity and lost on a restart unless backed by persistence.
- Disk: slower per lookup, but durable and cheap to retain for long windows, which is exactly what is needed for the "long outage recovery" requirement below.
- The two-tier design exists specifically to get the low latency of memory for the common case without losing correctness when memory alone is not enough.
Handling retries and long outage recovery
- Use an atomic check-and-set (an operation like Redis's
SETNX, or an equivalent conditional write) to claim a key before sending, so two concurrent consumer instances processing a redelivered pair of the same message cannot both decide they are the one to send it. - After a long outage (the in-memory cache is empty on restart, or entries aged out past their TTL during the outage), a redelivered message must not be trusted as new just because the fast cache has no record of it; falling back to the durable store catches exactly this case.
Worked example
The following script implements the two-tier dedup store (an in-memory "hot" tier with a TTL, backed by a durable tier with no TTL) and an idempotent consume function for the email-sending consumer, and was executed exactly as shown. A logical tick counter stands in for wall-clock time so TTL expiry is deterministic and reproducible, rather than asserting any real elapsed-time claim.
"""
Pinned, deterministic demo of a consumer-side dedup store for an email-sending
consumer, A logical tick counter
stands in for wall-clock time so TTL expiry is reproducible (no sleep()).
Run with: python3 idempotent_email_consumer.py
"""
from dataclasses import dataclass
from typing import Dict, Optional
@dataclass
class DedupRecord:
status: str # "in_progress" | "sent"
written_at_tick: int
class DedupStore:
"""Emulates a Redis-style store with SETNX (atomic check-and-set) and a
TTL measured in logical ticks, backed by a durable table for anything
that survives past the TTL window (the 'long outage recovery' path)."""
def __init__(self, ttl_ticks: int):
self.ttl_ticks = ttl_ticks
self._hot: Dict[str, DedupRecord] = {} # Redis-like cache
self._durable: Dict[str, DedupRecord] = {} # DB-backed, no TTL
def _expire_if_needed(self, key: str, now_tick: int):
rec = self._hot.get(key)
if rec and now_tick - rec.written_at_tick > self.ttl_ticks:
del self._hot[key]
def try_claim(self, key: str, now_tick: int) -> bool:
"""Atomic check-and-set: returns True if this call claimed the key
(i.e., no prior attempt is in flight or completed), False if a
duplicate delivery should be skipped."""
self._expire_if_needed(key, now_tick)
if key in self._hot:
return False # in_progress or sent, already claimed
if key in self._durable:
# survived a hot-cache eviction or full outage; still dedup
self._hot[key] = self._durable[key]
return False
self._hot[key] = DedupRecord(status="in_progress", written_at_tick=now_tick)
return True
def mark_sent(self, key: str, now_tick: int):
rec = DedupRecord(status="sent", written_at_tick=now_tick)
self._hot[key] = rec
self._durable[key] = rec # durable write-through, no TTL
def status(self, key: str) -> Optional[str]:
if key in self._hot:
return self._hot[key].status
if key in self._durable:
return self._durable[key].status
return None
class FakeEmailProvider:
def __init__(self):
self.sent_log = []
def send(self, email_id: str, recipient: str, template: str):
self.sent_log.append((email_id, recipient, template))
def consume(store: DedupStore, provider: FakeEmailProvider, message: dict, now_tick: int):
key = message["email_id"]
claimed = store.try_claim(key, now_tick)
if not claimed:
return "skipped_duplicate"
provider.send(message["email_id"], message["recipient"], message["template"])
store.mark_sent(key, now_tick)
return "sent"
if __name__ == "__main__":
TTL = 5 # ticks
store = DedupStore(ttl_ticks=TTL)
provider = FakeEmailProvider()
msg = {"email_id": "welcome-e100", "recipient": "a@example.com", "template": "welcome_v2"}
print("=== First delivery ===")
r1 = consume(store, provider, msg, now_tick=0)
print(f"result={r1}, provider.sent_log={provider.sent_log}")
assert r1 == "sent"
assert provider.sent_log == [("welcome-e100", "a@example.com", "welcome_v2")]
print("\n=== At-least-once redelivery of the SAME message, 1 tick later (well within TTL) ===")
r2 = consume(store, provider, msg, now_tick=1)
print(f"result={r2}, provider.sent_log={provider.sent_log}")
assert r2 == "skipped_duplicate"
assert len(provider.sent_log) == 1, "must not send twice"
print("\n=== Two consumers race on the SAME message at the SAME tick (concurrent redelivery) ===")
store2 = DedupStore(ttl_ticks=TTL)
provider2 = FakeEmailProvider()
claim_a = store2.try_claim("race-1", now_tick=0)
claim_b = store2.try_claim("race-1", now_tick=0)
print(f"consumer A claimed={claim_a}, consumer B claimed={claim_b}")
assert claim_a is True and claim_b is False, "only one consumer may claim the key"
print("\n=== Hot cache entry expires after TTL, but durable store still dedups (outage-recovery path) ===")
# Simulate a long outage: the in-memory/Redis-style cache is empty on
# restart (process restarted, or the key aged out past its TTL), but the
# durable table still has the record from before the outage.
after_outage_tick = 0 + TTL + 1 # past the TTL window
store._expire_if_needed("welcome-e100", after_outage_tick)
print(f"hot cache has key? {'welcome-e100' in store._hot}")
assert "welcome-e100" not in store._hot, "TTL should have evicted the hot entry"
r3 = consume(store, provider, msg, now_tick=after_outage_tick)
print(f"redelivery after TTL expiry + simulated outage -> result={r3}, provider.sent_log={provider.sent_log}")
assert r3 == "skipped_duplicate", "durable store must still catch the duplicate"
assert len(provider.sent_log) == 1, "still must not have sent twice"
print("\n=== A genuinely NEW message with a different email_id sends normally ===")
msg2 = {"email_id": "welcome-e101", "recipient": "b@example.com", "template": "welcome_v2"}
r4 = consume(store, provider, msg2, now_tick=after_outage_tick)
print(f"result={r4}, provider.sent_log={provider.sent_log}")
assert r4 == "sent"
assert len(provider.sent_log) == 2
print("\nALL ASSERTIONS PASSED")
Actual output from running python3 idempotent_email_consumer.py:
=== First delivery ===
result=sent, provider.sent_log=[('welcome-e100', 'a@example.com', 'welcome_v2')]
=== At-least-once redelivery of the SAME message, 1 tick later (well within TTL) ===
result=skipped_duplicate, provider.sent_log=[('welcome-e100', 'a@example.com', 'welcome_v2')]
=== Two consumers race on the SAME message at the SAME tick (concurrent redelivery) ===
consumer A claimed=True, consumer B claimed=False
=== Hot cache entry expires after TTL, but durable store still dedups (outage-recovery path) ===
hot cache has key? False
redelivery after TTL expiry + simulated outage -> result=skipped_duplicate, provider.sent_log=[('welcome-e100', 'a@example.com', 'welcome_v2')]
=== A genuinely NEW message with a different email_id sends normally ===
result=sent, provider.sent_log=[('welcome-e100', 'a@example.com', 'welcome_v2'), ('welcome-e101', 'b@example.com', 'welcome_v2')]
ALL ASSERTIONS PASSED
Walking the scenarios: the first delivery of welcome-e100 sends normally. A redelivery of the identical message one tick later (well within the 5-tick TTL) is skipped, and the send log still shows only one entry, proving no duplicate email went out. Two consumer instances racing on the same key at the same tick show only one successfully claims it (consumer A claimed=True, consumer B claimed=False), which is what the atomic check-and-set is for. After simulating the hot cache's TTL expiring (past tick 5) to represent a long outage, a redelivery of welcome-e100 is still correctly skipped, because the durable tier (no TTL) still has the record even though the fast tier does not, exactly the outage-recovery path the design requires. A genuinely new message (welcome-e101) sends normally and independently, confirming the dedup logic does not over-match on unrelated messages.
Trade-offs and pitfalls
- Common wrong turn: relying on the in-memory cache alone with no durable backing. It is fast for the common case, but a cache restart or an outage that outlasts the TTL causes exactly the duplicate-send failure idempotency was supposed to prevent.
- Common wrong turn: using a non-atomic "check, then set" (two separate operations) instead of a single atomic check-and-set. Between the check and the set, a second consumer instance can slip through and both end up believing they are the one to send.
- Common wrong turn: generating a fresh key on every delivery attempt (for example, from the broker's own delivery identifier) instead of using the message's own stable
email_id. A key tied to delivery mechanics changes on redelivery and stops deduplicating the thing that actually matters. - Senior signal: treating the TTL'd fast tier and the untimed durable tier as two different concerns (latency versus correctness-across-outages) rather than a single cache with an arbitrary expiry. The same two-tier idempotent-consumer pattern applies unchanged to other at-least-once consumers with a different natural key, for example a payment-notification consumer keyed by
payment_idinstead ofemail_id.
Explain Command Query Responsibility Segregation (CQRS). As a data engineer, when is CQRS valuable for analytics or operational workloads? Discuss trade-offs including complexity, eventual consistency of read-models, and strategies to make reads 'fresh' when required.
Sample Answer
Direct answer
Command Query Responsibility Segregation (CQRS) is the pattern of using a different model for writes (commands that change state) than for reads (queries that return state), instead of forcing one schema to serve both well. As a data engineer, it earns its added machinery when the write side's transactional shape and the read side's analytical or lookup shape diverge enough that one schema serves neither well, for example narrow row-level operational writes versus wide, pre-aggregated reporting reads. Treat it as a deliberate trade: schema simplicity for the ability to scale, model, and store reads and writes independently.
Structured elaboration
What CQRS actually separates
- Command side: an authoritative model (a normalized transactional store, or an event log) that enforces write-time invariants and produces state changes.
- Query side: one or more purpose-built read models (denormalized tables, search indexes, in-memory caches), each shaped for a specific access pattern rather than for correctness enforcement.
- A projector connects the two asynchronously: it consumes the write side's changes (domain events, or change-data-capture (CDC) records) and updates the read model(s).
When CQRS is valuable for analytics workloads
- Reporting or business intelligence (BI) queries need aggregation shapes (rollups by category, hour, region) that would otherwise require expensive joins or full scans against the transactional schema.
- Several independent consumers need different projections of the same data (a finance rollup, a fraud-detection view, a customer dashboard); three denormalized read models are cheaper to operate than three sets of ad-hoc joins against the online transaction processing (OLTP) store.
- Analytical queries would otherwise contend for locks and I/O with operational writes on the same tables.
When CQRS is valuable for operational workloads
- Write throughput and read throughput need to scale independently and at different rates (write-heavy ingestion feeding a low-cardinality operational dashboard).
- Write-side invariants are complex enough (state machines, multi-step validation) that mixing them with read-optimization concerns would make the write model harder to reason about.
- Not valuable: a small application with one read pattern that already matches the write schema. There CQRS adds a projector, extra storage, and extra failure modes with no offsetting benefit.
Trade-off: complexity
You now operate an additional pipeline (the projector), additional storage (one or more read stores), and additional failure modes: projector lag, projector crashes mid-batch, and schema drift between the write shape and the read shape.
Trade-off: eventual consistency of read models
Because the read model updates asynchronously, a query issued immediately after a write can observe stale data. The size of that staleness window is a direct function of projector throughput and batching, not something a team can design around by ignoring it.
Strategies to make reads "fresh" when required
- Read-your-writes for the writer: return enough state in the command response (or a version/sequence number) that the client who just wrote never needs to trust the read model for its own write.
- Tighten the pipeline: smaller batches and event-driven push instead of periodic batch pull shrinks the staleness window, at the cost of more frequent projector invocations.
- Expose staleness explicitly: attach a last-updated version or timestamp to read-model responses so callers can judge whether the data is fresh enough, instead of the system silently presenting stale data as current.
- Selective synchronous update: for a small, well-identified set of critical fields, update the read model synchronously in the write path (accepting some coupling) while everything else stays asynchronous.
Worked example
An order system accepts writes at 500 orders per minute (about 8 to 9 orders per second) into a transactional order table. A "revenue by category, per hour" read model is built by a projector that drains the order-events stream every 60 seconds and applies that batch of updates.
- Worst-case staleness for that read model equals the batch interval: 60 seconds. An order committed just after a batch run will not appear until the next run.
- Average staleness is roughly half the batch interval, about 30 seconds, if orders arrive close to uniformly across the minute.
If the product requirement is "the dashboard must reflect a new order within 10 seconds," this projector cadence fails outright: 60 seconds worst case exceeds the 10-second bound. The fix is either to drop the batch interval below 10 seconds, or to read the specific "orders placed today" counter synchronously from the write side while the rest of the dashboard stays on the 60-second cadence.
Trade-offs and pitfalls
- Common wrong turn: adopting CQRS because it sounds architecturally sophisticated for a workload that has a single read pattern already matching the write schema. That is pure overhead with no payoff.
- Common wrong turn: treating "eventually consistent" as a detail to sort out later. Staleness needs an explicit, stated bound (or an explicit "no bound" with a user-facing affordance for it) decided at design time, not discovered in production when a user cannot see the order they just placed.
- Senior signal: naming a concrete staleness budget and matching the pipeline's cadence to it, rather than discussing CQRS only in the abstract.
A producer team wants to remove a field that several downstream consumer teams currently read from an event. Two consumer teams say removing it will break their service; a third says they don't use it. Walk through how you would let the producer team make this change (and future ones like it) without breaking consumers who depend on the old shape, and what you would need in place beforehand for that to be possible.
Sample Answer
Direct answer
Do not trust the third team's "we don't use it" self-report as the basis for removing the field: verify usage empirically (via consumer-driven contract tests or a lineage/usage audit against the registry) before touching anything, then make the change additive and reversible rather than destructive: deprecate the field with a stated window while the two dependent teams migrate to whatever replaces it, and only physically remove it after the registry and the audit both agree nobody depends on it. What has to be in place beforehand for this (and every future change like it) to be low-risk is a schema registry with enforced compatibility, consumer-driven contracts wired into the producer's continuous integration (CI) pipeline, and field-level usage visibility, not policy written down but unchecked by tooling.
Structured elaboration
Step 1: verify the claims, don't just count votes
- Two teams say removal breaks them; one says they don't use it. Before acting on any of these statements, check field-level usage against real traffic (a lineage/audit view keyed to the registry, or the third team's own consumer-driven contract, if one exists) rather than relying on memory or a quick grep of code that might miss dynamic field access. A team that "doesn't use" a field today may still have a downstream job or dashboard reading it that the team itself is not aware of.
- If no usage-tracking tooling exists yet, this is itself the finding: the producer team cannot safely make this change (or any future one) without it, and building that visibility becomes the first deliverable, not an afterthought.
Step 2: never remove directly, deprecate first
- Mark the field deprecated in the schema registry with a stated time-to-live (a concrete window, for example 60 to 90 days depending on how disruptive the change is), rather than removing it in the next release. Deprecation is reversible; removal is not.
- If the field is being replaced rather than dropped outright (a common case: a flatter
carrier_codestring replaced by a richercarrierobject), ship the replacement as an additive, backward-compatible field alongside the old one first, so the two dependent teams can migrate to the new field on their own schedule while the old one still works.
Step 3: give the two dependent teams a real migration path, not a deadline
- Communicate the deprecation and its window through the same channel every consumer team already watches (the registry's deprecation surface, not a one-off message that is easy to miss).
- Track each dependent team's migration status against the deprecation window on a shared dashboard, so "day 89 and two teams still on the old field" is visible before it becomes an incident, not discovered at the deadline.
- If a team cannot realistically migrate within the standard window, that is a scheduling negotiation with a visible tracked extension, not silent, indefinite delay and not a forced break either.
Step 4: remove only once verified safe
- Remove the field only after both signals agree: the deprecation window has closed, and the usage audit independently confirms zero real traffic reads it, including the third team's prior "we don't use it" claim, now verified rather than assumed.
What must be in place beforehand for this, and future changes like it, to be possible
- A schema registry that enforces compatibility on every change automatically, so an accidental breaking change cannot ship silently regardless of what any individual engineer intends.
- Consumer-driven contracts (or equivalent field-level usage tracking / lineage) wired into the producer's continuous integration (CI) pipeline, so "who actually depends on this field" is a query, not a Slack thread across three teams.
- A documented, tooled deprecation workflow (a way to mark a field deprecated with a TTL, and tooling that surfaces that to consumers and eventually blocks removal until the window and the usage audit both clear), so this becomes a repeatable, low-drama process rather than a one-off negotiation every time a producer wants to evolve its data.
- Without these three things in place, the producer team's only honest options for a genuinely breaking change are: negotiate with every consumer team by hand each time (slow, does not scale past a handful of consumers), or ship the breaking change and hope (which is what created this exact situation).
Worked example
Say the field in question is legacy_shipping_zone on a shipment-events topic with 8 consumer teams subscribed. Team A and team B say they read it (matches the scenario's "will break" teams); team C says they don't.
- A usage audit against 30 days of production traffic (using consumer group lag/read metrics keyed to which fields a consumer's deserializer actually accesses, or the registry's usage tracking if wired up) confirms: team A genuinely reads it in a nightly job, team B reads it in a real-time dashboard, and team C's claim holds, zero reads from team C's consumer group in the 30-day window.
- The field is marked deprecated with a 60-day TTL. A replacement
shipping_zone_v2field ships additively alongside it in the same event, so team A and team B can migrate independently: team A's nightly job is updated in week 2 (low urgency, batch job, easy to redeploy); team B's real-time dashboard, wired into a customer-facing screen, is updated in week 7 after more careful testing. - At day 60, the audit is re-run: both team A and team B now read
shipping_zone_v2exclusively; nobody readslegacy_shipping_zone. The field is removed. Team C, whose original claim was correct, was never blocked or delayed by any of this.
Trade-offs and pitfalls
- Common wrong turn: trusting a consumer team's self-reported "we don't use it" without verification. Self-reports are honest but incomplete; a dynamic field access, a downstream job the reporting team forgot about, or stale documentation can all make a well-intentioned "we don't use it" wrong.
- Common wrong turn: removing the field once the two dependent teams say they've migrated, without an independent audit confirming zero remaining traffic. "We think we're done migrating" and "the traffic confirms nobody reads the old field" are different claims, and only the second one is safe to act on.
- Common wrong turn: treating this as a one-time negotiation to get through, rather than as evidence that the team needs standing tooling (registry, contracts, usage visibility) so the next field removal does not require the same manual, three-team back-and-forth.
- Senior signal: naming the prerequisite tooling explicitly as the actual answer to "what would you need in place beforehand," rather than only describing the sequence of steps for this one field.
Unlock Full Question Bank
Get access to all 25 Event-Driven Architecture and Asynchronous Messaging interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.