Event-Driven Architecture and Asynchronous Messaging Questions
Designing systems around events and message passing: publish/subscribe, message queues, event streaming, choreography versus orchestration, and decoupling producers from consumers. Covers delivery semantics (at-least-once, at-most-once), ordering, backpressure, dead-letter handling, and the operational tradeoffs of asynchronous flows. Includes async processing patterns for offloading slow work.
A customer needs an immutable audit trail and the ability to rebuild multiple read models quickly. Compare event sourcing + CQRS against a traditional relational database augmented with Change Data Capture (CDC). Discuss complexity, operational cost, replayability, schema evolution, developer ergonomics, and scenarios where event sourcing is or is not justified.
Sample Answer
Direct answer
For an immutable audit trail with fast rebuild of multiple read models, both event sourcing plus Command Query Responsibility Segregation (CQRS) and a relational database augmented with change-data-capture (CDC) can work, but they solve it differently: event sourcing stores domain intent as the source of truth and replays it deterministically, while CDC turns an existing relational database's row-level changes into a changelog that downstream consumers materialize into views. Pick event sourcing when the business needs the "why," not just the "what," and needs a canonical replayable log; pick relational-plus-CDC when the relational model already fits the domain and the team wants a lighter operational footprint.
Structured elaboration
Complexity
- Event sourcing + CQRS: higher architectural complexity. Domain events are the source of truth; the team must implement an append-only event store, event versioning, snapshotting, and one or more projections, and must handle at-least-once delivery into those projections.
- Relational + CDC: lower incremental complexity if the team already runs a relational database. CDC (typically reading the database's write-ahead log, for example via Debezium) streams committed transactional changes to downstream systems without changing the core domain model.
Operational cost
- Event sourcing needs operational maturity for the event store itself: scaling, backups, compaction, and snapshotting are all new operational surfaces, plus projection workers and a message bus.
- CDC leans on existing database tooling and log-shipping infrastructure; fewer bespoke components, generally lower operational overhead, but the pipeline is now coupled to the source database's internal log format and retention.
Replayability (folding the append-only-log / materialized-view nuance)
- Event sourcing's replay is native and deterministic: any read model or a brand-new projection can be rebuilt by replaying the event log from the start (or from a snapshot). The event log is intentionally an append-only log of domain intent, and every read model is a materialized view derived from it.
- CDC is structurally similar (the database's write-ahead log is also an append-only log, and CDC consumers building denormalized tables are also building materialized views), but the two differ in what the log records: CDC's changelog captures row-level state deltas ("column X became Y"), not business intent ("customer upgraded their plan"). Reconstructing intent from a stream of row deltas is lossy and sometimes ambiguous, and a full CDC-based rebuild requires retaining the source database's change history for as long as you might need to replay it, which is a weaker retention guarantee than an event store's log is designed to give.
Immutable audit trail and intent
- Event sourcing stores intent explicitly as first-class domain events; the audit trail directly answers "what business action happened and why."
- CDC provides a factual log of persisted state changes, useful for audit of what changed, but not necessarily why, unless the application already wrote that intent into the row (e.g., an explicit
reasoncolumn).
Schema evolution
- Event sourcing requires an explicit event-versioning strategy (adding fields with defaults, or introducing a new event type for a breaking change, plus upcasters that translate old event versions when replayed).
- CDC schema evolution tracks the source table's schema; adding a nullable column is generally safe, but column renames, type changes, or table restructuring can break the CDC pipeline's mapping and every downstream consumer of it.
Developer ergonomics
- Event sourcing has a steeper learning curve: developers must think in events, eventual consistency, and projection design, but gain very clear auditability and strong support for temporal questions ("what did this account look like on March 3rd?").
- CDC lets developers keep writing familiar create/read/update/delete (CRUD) code against the relational schema; projections consume the CDC stream with comparatively little domain-model rework.
When event sourcing + CQRS is justified
- A complex domain with rich auditability or regulatory replay requirements, where the business needs to reconstruct exact past system state.
- Many read models that change frequently and need fast, correct rebuilds.
- The business explicitly wants to capture intent, not just state, for analytics or downstream machine learning use.
When it is not justified
- A straightforward create/read/update/delete domain where the relational schema already models the business well and the database's own transaction log already satisfies audit and retention requirements.
- Limited engineering bandwidth or a need for fast delivery where the extra event-sourcing machinery would slow the team down for no corresponding benefit.
Worked example
A payments team is choosing between the two approaches for a ledger requiring 7 years of audit retention and 3 read models (customer statement, fraud-review queue, regulatory export).
- Event sourcing: the event store holds roughly 40 million domain events over 7 years (about 15,600 events/day on average for a mid-size ledger). A new read model, say a fourth "tax reporting" view added in year 5, is built by replaying those 40 million events once, deterministically, against the new projection logic; correctness is verifiable because the same event log produces the same output every time it is replayed.
- Relational + CDC: the same 7-year retention means either keeping 7 years of database write-ahead log history available to the CDC pipeline (expensive and often beyond what most databases retain by default) or accepting that a full historical rebuild of a new view is not actually possible from CDC alone, only from that point forward. This is the concrete cost of CDC's weaker replay guarantee versus a purpose-built event store: it shows up exactly when the business asks for a new view of old data.
Trade-offs and pitfalls
- Do not choose event sourcing for its audit-trail marketing value alone; a relational database with proper CDC and immutable audit columns can satisfy many audit requirements at a fraction of the operational cost.
- Do not underestimate CDC's replay ceiling: if "rebuild any read model from any point in history" is a hard requirement, verify the source database's log retention actually supports it before committing to CDC as the long-term answer.
- Senior signal: naming the retention and rebuild requirement in concrete terms (how far back, how many read models, how often they change) before picking a side, rather than treating this as a purely stylistic architecture preference.
Explain backpressure and flow-control techniques across networked services and message brokers: TCP flow control, HTTP/2 flow control, reactive streams (request-n), credit-based broker flow control, and producer-side throttling. For each technique explain when it's most appropriate and how you'd design end-to-end controls to prevent OOMs and cascading failures.
Sample Answer
Direct answer
These five techniques operate at different layers, from the raw byte stream up to an individual message, and they compose: none of them alone protects an entire pipeline from an out-of-memory (OOM) crash or a cascading failure, the end-to-end design comes from layering them so each protects the specific thing the layer below it cannot.
Structured elaboration
| Technique | Mechanism | Layer | Most appropriate when |
|---|---|---|---|
| Transmission Control Protocol (TCP) flow control | Receiver advertises a byte window in each acknowledgment; sender never has more unacknowledged bytes outstanding than that window | Transport, automatic for any TCP connection | Baseline protection for any point-to-point byte stream, especially when you do not control the application protocol running on top of it |
| HTTP/2 flow control | A credit-like byte window, but per multiplexed stream as well as per connection | Application-transport boundary | Many logical exchanges (for example gRPC calls) multiplexed over one TCP connection, where one slow stream should not starve the others sharing that connection |
Reactive streams (request(n)) | The subscriber explicitly tells the publisher how many items it is ready to receive next; the publisher cannot push more than requested | Application, in-process or library-level | An async pipeline or stream-processing library where backpressure needs to be an explicit, item-oriented part of the consumer's own control flow, not inferred from bytes |
| Credit-based broker flow control | A consumer grants the broker a bounded number of unacknowledged in-flight messages (a prefetch or credit limit); each acknowledgment replenishes one unit | Message-broker to consumer relationship | Bounding one consumer's own in-flight work without needing the producer to coordinate with it directly, the broker mediates |
| Producer-side throttling | The producer itself is rate-limited before sending, based on a fixed budget or on live downstream health signals (queue depth, consumer lag, error rate) | Outermost, protects the whole system including the broker | Protecting the system as a whole, including the broker's own storage, not just one consumer |
Why layering matters, not picking one. TCP and HTTP/2 flow control protect the wire, an application can still accumulate an unbounded in-memory backlog even while the underlying connection is perfectly healthy, because those layers only bound bytes in flight on the network, not messages queued in application memory waiting to be processed. Credit-based broker flow control is what actually bounds an individual consumer's own memory footprint, by capping how many unacknowledged messages the broker will hand it. Producer-side throttling is the only one of the five that protects the broker's own memory or disk, without it, a healthy consumer relationship does nothing to stop the broker itself from accumulating an unbounded queue if producers keep publishing faster than the system can drain.
Preventing cascading failure specifically. Without an upstream signal, a slow consumer causes the broker's queue depth to grow unbounded, risking the broker's own out-of-memory (OOM) crash or disk exhaustion. If the broker then starts shedding load (rejecting connections or writes) to protect itself, producers that do not back off in response to that rejection can retry-storm the broker, making the incident worse instead of better. This is why producer-side throttling needs to be dynamic, responsive to live broker or consumer health signals, rather than a fixed rate calibrated only for normal conditions; a fixed rate does nothing extra during the exact moment it is needed most.
Worked example
A consumer configured with a prefetch (credit) limit of 50 unacknowledged messages, average payload size 2 kilobytes (KB).
With credit-based flow control: the broker never has more than 50 unacknowledged messages outstanding to this consumer, so its in-flight memory footprint for this consumer's queue is bounded at roughly 50 x 2 KB = 100 KB, regardless of how large the total backlog waiting behind those 50 becomes.
Without it (an unbounded prefetch): the broker keeps pushing every available message, so the consumer's in-flight count, and its memory footprint, grows with the size of the backlog itself, with no structural bound, until the backlog stops growing or the consumer runs out of memory.
The structural difference (bounded versus unbounded growth) is what the mechanism buys you; the 100 KB figure is a direct, pinned-input calculation (50 messages times 2 KB each), not a measurement.
Trade-offs and pitfalls
- Setting broker credit or prefetch too low sacrifices throughput, many small round trips to replenish a small credit budget.
- Setting it too high defeats the purpose, it looks bounded on paper but the bound is too large to meaningfully protect memory.
- Assuming TCP or HTTP/2 flow control alone is "enough" backpressure is the single most common pitfall named implicitly by this question: those protect the wire, not an application's own queues or memory.
- Static, fixed-rate producer throttling calibrated for normal conditions does not actually prevent cascading failure during a real incident, since the fixed rate was never designed to respond to the incident happening; dynamic throttling that reacts to consumer lag or broker health is what closes that gap.
Propose a backpressure/flow-control design when a fast producer floods a slow consumer connected via a queue system. Include mechanisms on both producer and broker sides (bounded queues, rate-limiting, pause/resume, token buckets), and describe how to implement graceful degradation while preserving important messages.
Sample Answer
Direct answer
When a fast producer floods a slow consumer behind a queue, the fix is layered: constrain the producer's effective send rate so it can never overwhelm the broker in the first place, cap what the broker will hold so an unconstrained burst fails fast instead of growing without limit, and give the consumer an explicit signal to slow the producer down. On top of that, degrade by shedding low-value messages first, never by silently dropping everything.
Structured elaboration
Producer-side mechanisms
- Token bucket rate limiter in front of the publish call, sized to the consumer's sustained throughput plus a small burst allowance, not to the producer's natural output rate.
- Pause/resume (credit-based flow control): the broker or consumer publishes a credit or high/low watermark signal; the producer pauses publishing when credits are exhausted and resumes when the consumer signals capacity again. This is the mechanism that actually closes the loop. A token bucket alone just smooths a rate, it does not react to real backlog.
- Client-side buffering with its own bounded size, so a paused producer does not itself become an unbounded memory sink. Once that local buffer is full, callers see backpressure (a blocking call, or a rejected/deferred write) instead of the process growing without limit.
Broker-side mechanisms
- A bounded queue with an explicit capacity ceiling. Once full, the broker rejects new publishes (or blocks the producer, depending on the client contract) rather than growing memory without bound. This turns "the consumer is a little slow" into a visible, actionable signal instead of a silent memory leak.
- Priority lanes or multiple queues by message class, so that once the bound is reached, the broker can shed low-priority traffic first while still admitting high-priority messages.
- Consumer-side autoscaling triggered off queue depth or lag, so backpressure is a symptom you fix by adding capacity over time, not just a state you tolerate forever.
Multi-stage and hierarchical topologies. In a pipeline with more than one hop (producer, broker, an intermediate aggregator, final consumer), apply backpressure independently at each hop rather than only at the outermost edge. A hierarchical fan-out (broker to regional relays to final consumers) needs the same bounded-queue-plus-token-bucket pattern at each stage, otherwise one stage's overflow just moves the flood one hop downstream instead of resolving it. On the signaling side, a producer that receives an explicit "too busy" response (an HTTP 429 status code from an API-fronted queue, or a broker-specific backpressure response) should treat repeated instances as its own local circuit breaker: stop sending for a cooldown window rather than retrying immediately into a system that just told you it is full.
Graceful degradation, preserving important messages
- Classify messages into priority tiers at publish time (a header or routing key), not after the fact.
- When the bounded queue approaches capacity, drop or defer the lowest tier first (an analytics ping before a checkout confirmation), and reject new low-priority publishes at the producer's client library so the rejection happens as close to the source as possible.
- Keep the shed decision observable: emit a metric and a structured log entry for every dropped or throttled message, so shedding is a visible design choice, not silent data loss.
Worked example
Say the producer bursts at 5,000 messages/sec, and the consumer side is a pool of 8 workers, each sustaining about 100 messages/sec, for a total of 800 messages/sec.
net fill rate=5000−800=4200 msg/sIf the broker's bounded queue is capped at 50,000 messages, an unmitigated burst fills it in:
time to fill=420050000≈11.9 sUnder 12 seconds from full-speed burst to a broker that starts rejecting or blocking. That is the argument for a producer-side token bucket sized to consumer capacity (800/sec) with a small burst allowance (say 960/sec, 20% headroom) rather than relying only on the broker's cap: throttling at the source means the queue never gets anywhere near 50,000 in the first place, and the 12-second figure above becomes the worst case only if the token bucket is missing or misconfigured.
Trade-offs and pitfalls
- A token bucket sized to the producer's natural rate instead of the consumer's sustained rate just moves the flood downstream. Size it to what the consumer can actually drain.
- Bounding the queue without pause/resume just converts "slow consumer" into "producer errors," which is safer than unbounded memory growth but is still an outage if nothing tells the producer to slow down instead of retry-storming.
- Dropping messages under pressure without a priority scheme treats a checkout confirmation the same as a page-view ping. Senior answers always tier the shed.
- Consumer autoscaling has ramp-up lag (new workers take time to become ready), so it complements but does not replace bounded queues and rate limiting. It is the medium-term fix, not the instantaneous one.
Event-sourcing stores all state changes as events. Discuss the trade-offs between storing only events versus introducing periodic snapshots. As a data engineer, explain snapshotting frequency, snapshot storage, snapshot validation, rehydration cost, and strategies for compaction or archival to control event-store growth.
Sample Answer
Direct answer
Storing only events gives perfect auditability and deterministic rebuilds, but rehydration cost (the work to replay events into current state) grows with the event count, so as a stream ages, reads and rebuilds get slower and more expensive. Periodic snapshots cap that cost by giving rehydration a recent starting point instead of the beginning of time, at the price of extra storage and a validation problem: a snapshot must be provably consistent with the events it claims to summarize. As a data engineer, pick a snapshot cadence and compaction policy driven by measured rehydration cost, not by a fixed rule of thumb.
Structured elaboration
Snapshotting frequency
- Event-count-based: snapshot every N events (commonly in the low thousands, tuned to the aggregate). Predictable worst-case replay cost regardless of how much wall-clock time has passed.
- Time-based: snapshot daily or hourly, better for aggregates with low or bursty event velocity where event count alone is not a reliable trigger.
- Hybrid: aggressive event-count triggers for hot, frequently-updated aggregates; time-based triggers for long-lived but rarely-updated ones.
- Whichever trigger is chosen, tune it against measured rehydration cost (replay time and CPU per rehydration), not a value picked without data.
Snapshot storage
- Store snapshots as immutable objects in cost-efficient storage, tagged with the aggregate identifier, the schema version, and the sequence number of the last event folded into the snapshot.
- Keep hot, frequently-accessed snapshots in fast storage or a cache; move cold snapshots to cheaper storage tiers with lifecycle rules.
Snapshot validation
- A snapshot must carry: the last-applied event sequence number, a schema version, and a checksum of its contents.
- On load, verify the checksum and confirm sequence continuity (no gap between the snapshot's last-applied sequence and the next event to replay). On mismatch, fall back to a full replay from the last known-good snapshot or from the beginning.
- Run periodic background verification jobs that independently rebuild a sample of aggregates from events alone and diff against the stored snapshot, to catch silent snapshot corruption before it is relied upon.
Rehydration cost
- Rehydration cost is: load the snapshot, then replay only the events recorded after that snapshot's sequence number. Cost scales with events-since-snapshot, not with the aggregate's total lifetime event count.
- The snapshot cadence directly bounds the worst case: with a snapshot taken every N events, no rehydration ever replays more than N-1 events.
Compaction and archival to control event-store growth (folding the replay/backfill and derived-dataset-reprocessing nuance)
- Compact by taking a full snapshot and, only where retention rules permit it, archiving or truncating events older than that snapshot's sequence to cold, cheaper storage rather than deleting them outright; legal or audit requirements often forbid true deletion.
- The archival tier still has to remain replayable: any derived dataset (an analytics table, a machine learning feature store, a rebuilt read model) that was originally built by consuming the full event history needs that same history available if it must ever be reprocessed, for example after a bug fix in the transformation logic. A retention policy that only optimizes for "rehydrate a live aggregate quickly" and quietly discards old events breaks backfill for any derived dataset that depended on that full history, even though the live aggregates themselves are unaffected. Decide retention and compaction against both use cases explicitly, not just the aggregate-rehydration one.
- Use a tiered retention: hot recent events in the primary event store, older events moved to compressed, partitioned cold storage with an index that supports selective replay by aggregate and time range.
- Mark truncation points explicitly (a compaction marker event, or metadata) so any consumer replaying the stream can detect where the live tier's history starts and knows to fetch older ranges from the archive if it needs them.
Worked example
An account aggregate receives 50 events per day on average.
- Without snapshots, rehydrating that account after 5 years of history means replaying 5 * 365 * 50 = 91,250 events.
- With a snapshot taken every 1,000 events, the worst-case replay after any snapshot is 999 events (999 / 91,250 = 1.09%, so rehydration cost is bounded to roughly 1% of the no-snapshot case at year 5, not strictly under 1%: it is a hair over).
- At 50 events/day, a snapshot every 1,000 events fires roughly every 20 days (1,000 / 50 = 20), so the account accumulates at most 20 days' worth of unsnapshotted events at any time.
- If the business later needs to reprocess the full 5-year history to backfill a new derived dataset (say, a new fraud-scoring feature that needs every historical event, not just the latest snapshot), the archived event range covering all 91,250 events must still be retrievable even though live rehydration never touches most of them.
Trade-offs and pitfalls
- Common wrong turn: choosing a snapshot cadence without measuring actual rehydration cost, then discovering it under- or over-snapshots (too frequent wastes storage and write bandwidth on snapshotting; too infrequent leaves rehydration slow).
- Common wrong turn: treating snapshot validation as optional. An unvalidated, silently corrupt snapshot is worse than no snapshot: it produces confidently wrong current state instead of forcing a (correct, if slow) full replay.
- Common wrong turn: designing retention purely around live-aggregate rehydration speed and discovering, only when a derived dataset needs to be rebuilt, that the events required for that rebuild were already archived out of reach or deleted.
- Senior signal: naming the specific numeric relationship between snapshot cadence and worst-case replay cost, and treating archival policy as a decision that serves more than one consumer (live rehydration and derived-dataset reprocessing), not just the first one that comes to mind.
Explain the difference between publish-subscribe and point-to-point (producer-consumer) messaging patterns. Provide concrete scenarios where pub/sub is a better fit (e.g., notifications, analytics) and where queues are preferable (e.g., work queues, task processing), particularly in multi-tenant SaaS and event-driven microservice architectures.
Sample Answer
Direct answer
Publish-subscribe (pub/sub) delivers each message to every interested subscriber, so it fits situations where multiple independent parties need to know the same fact happened, such as notifications or analytics. Point-to-point (producer-consumer, work-queue style) delivers each message to exactly one consumer among a pool, so it fits situations where a unit of work must be done exactly once by whichever worker picks it up, such as background task processing. The distinction is about fan-out (one-to-many awareness) versus load distribution (one-of-many execution), and both patterns are commonly implemented on the same underlying broker.
Structured elaboration
Publish-subscribe. A publisher emits an event to a topic; every subscriber with an active subscription receives its own copy. Subscribers are typically unaware of each other, can be added or removed without changing the publisher, and each one processes the event for its own purpose. This is the right model whenever "N different systems need to react to the same fact" is the actual requirement: a "UserSignedUp" event might be consumed by an email-welcome service, an analytics pipeline, and a fraud-scoring service simultaneously, with none of them competing for the message.
Point-to-point (work queues). A producer places a task on a queue; a pool of competing consumers pulls from the same queue, and each task is handled by exactly one consumer. This is the right model for "this unit of work needs to happen once, by whichever worker is free," such as resizing an uploaded image or sending a single transactional email: you do not want three workers all resizing the same image.
Where pub/sub is the better fit. Notifications is the clearest case: a single "OrderShipped" event needs to reach a push-notification service, an SMS service, and an in-app activity feed, each independently, and adding a fourth channel later should not require touching the producer. Analytics is the same shape: every business event (page view, purchase, signup) typically needs to reach an analytics pipeline in addition to whatever else consumes it, without competing with those other consumers for the message.
Where queues are preferable. Work queues and task processing are the clear case: a video-transcoding job, a report-generation job, or an outbound-email send should be picked up and completed by exactly one worker, with the queue's competing-consumers model providing natural load balancing and horizontal scaling (add more workers, they compete for the same backlog) without any risk of duplicate execution beyond what at-least-once delivery already requires the consumer to handle idempotently.
Multi-tenant SaaS (software as a service) and event-driven microservices. In a multi-tenant SaaS system, pub/sub is what lets independently-owned services (billing, usage-metering, audit logging) all react to the same tenant-level event, such as "SubscriptionUpgraded," without the team that owns the upgrade flow needing to know or coordinate with every downstream consumer; new consumers subscribe without any change to the publisher. Point-to-point queues, by contrast, are what those same microservices use internally for their own background work, such as a billing service's queue of pending invoice-generation tasks, where exactly-once-effective execution by one worker in the pool is the requirement, not fan-out to observers.
Worked example
A multi-tenant SaaS platform publishes a "TenantUpgraded" event when a customer moves from a free to a paid plan. Three independent subscribers exist on this topic: a billing service that starts metered invoicing, a feature-flag service that unlocks paid features, and a customer-success service that triggers an onboarding email sequence. All three receive their own copy of the same event; the team that owns the upgrade flow never had to know these three consumers existed. Separately, the feature-flag service's own onboarding-email trigger enqueues an actual "send welcome email" task onto a point-to-point work queue consumed by a pool of 5 worker processes; only one of those 5 workers ends up sending that specific email, because the queue hands each task to a single competing consumer, not to all 5.
Trade-offs and pitfalls
The common mistake is using a work queue where pub/sub was needed: if a "TenantUpgraded" task were placed on a single point-to-point queue instead of published to a topic, only one of billing, feature-flags, or customer-success would ever see it, and the other two would silently never fire, which is a subtle and easy-to-miss integration bug. The opposite mistake is using pub/sub where a work queue was needed for a task that must be done exactly once: if "resize this uploaded image" were published to a topic with multiple subscribed workers, every worker would independently resize the same image, wasting resources and, if the workers write to the same output path, potentially racing each other. A senior answer names this fan-out-versus-load-distribution distinction explicitly, rather than treating "pub/sub" and "queue" as interchangeable synonyms for "asynchronous messaging."
Unlock Full Question Bank
Get access to all 35 Event-Driven Architecture and Asynchronous Messaging interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.