Architectural Patterns and Anti-Patterns Questions
Architecture-level patterns and the anti-patterns that signal a wrong turn. Patterns: layered and n-tier architecture, including where cross-cutting concerns like authentication, rate-limiting and tracing belong, dependency injection trade-offs, and thin-versus-fat controller design; hexagonal (ports and adapters) and clean architecture; CQRS and event sourcing; backend-for-frontend; plugin (microkernel) extension models; and the coupling, cohesion, encapsulation and separation-of-concerns principles behind them, including when each applies and what it costs. Anti-patterns: distributed monolith, chatty services, shared-database coupling, cyclic service dependencies, leaky abstractions that expose internal schemas, and golden-hammer pattern adoption. Covers the detection signals (deploy coupling, call-graph fan-out, change amplification, trace evidence), incremental remediation, and architecture governance that keeps smells from recurring. This is about diagnosing and fixing the smell in an existing design, not the monolith-versus-microservices decision itself.
A team proposes an event-sourced architecture for a business-critical domain (for example, bookings or orders). Evaluate it: describe the main components (event store, projections, snapshots) and the operational challenges (replay cost, migration, storage growth). As the engineering manager, what organizational prerequisites (testing strategy, developer tooling, team readiness) would you require before approving this architecture?
Sample Answer
Direct answer
Event sourcing stores every change to a business entity as an immutable event ("BookingCreated", "BookingDateChanged", "BookingCancelled") and derives current state by replaying those events, instead of overwriting a row. For a business-critical domain like bookings it can be the right call when the business genuinely needs a complete, trustworthy history (audit, disputes, "what did we know at the time?") or needs many different read models of the same facts. It is expensive to operate: events are permanent, so schema mistakes live forever, rebuilding read models can take hours, and the team needs skills most CRUD (create, read, update, delete) teams do not yet have (event modelling, projection design, and replay/rebuild tooling among them; see the prerequisites below). As the engineering manager, I would approve it only if the team can name the business need that a normal database plus an audit table cannot meet, and only once specific testing, tooling and readiness prerequisites are in place. Otherwise I would push back towards a conventional design, possibly with event publishing for integration (publishing a small, deliberately stable set of events other teams can subscribe to, kept separate from whatever internal event shapes the service uses for its own logic).
The main components
flowchart LR
Cmd["Command: change booking date (a request to make a change)"] --> Agg[Booking aggregate: load events, check rules]
Agg -->|append new events| ES[(Event store: append-only log per booking)]
ES --> P1[Projection: booking detail view]
ES --> P2[Projection: daily occupancy report]
ES --> Snap[(Snapshots every N events)]
Snap --> Agg
P1 --> Q[Queries from UI and APIs]
- Event store: an append-only log, one ordered stream per entity (per booking). It must support "append these events only if the stream is still at version N" (this is called optimistic concurrency: instead of locking the stream first, you attempt the append and let the version check reject it if another write already happened), so two concurrent changes cannot both win. This can be a dedicated product or a relational table with a unique (stream id, version) constraint.
- Aggregate: the code that loads a booking's events, rebuilds its current state in memory, checks business rules ("cannot cancel after check-in"), and emits new events.
- Projections: read models built by consuming events, such as a booking-detail table for the UI or an occupancy report. Each is disposable: drop it and replay the events to rebuild it. This is closely related to CQRS (command query responsibility segregation: separate models for writes and for reads), which event sourcing almost always implies.
- Snapshots: a saved copy of an aggregate's state at event number N, so loading a long-lived entity replays only the events after the snapshot instead of all of them.
Operational challenges, with numbers
Assumptions for a mid-size booking platform: 200,000 bookings per day, an average of 8 events per booking over its life, 1 KB (1,000 bytes) per stored event including metadata.
Storage growth
- Per day: 200,000 × 8 × 1 KB = 1,600,000 KB = 1.6 GB.
- Per year: 1.6 GB × 365 = 584 GB.
- After 3 years: 1.752 TB, plus indexes and projections.
Events are never updated, so this only grows. The plan needs an archival policy (move closed bookings' streams older than the retention period to cheaper storage) and a clear answer to data-deletion requests: personal data inside immutable events is a real problem under privacy laws such as the GDPR (General Data Protection Regulation). The usual approach is to keep personal data out of events or encrypt it with a per-customer key that can be destroyed ("crypto-shredding").
Replay cost
After 3 years there are 200,000 × 8 × 365 × 3 = 1,752,000,000 events. If a new projection processes 20,000 events per second (an assumption; measure your own), a full rebuild takes 1,752,000,000 / 20,000 = 87,600 seconds, about 24.3 hours. That has consequences:
- A bug in a projection means up to a day of rebuilding while the old projection keeps serving.
- New read models need a blue-green approach: build the new projection alongside the old one, switch reads when it catches up.
- Partitioning the rebuild across workers (by stream) is needed to bring this down.
Aggregate load time and snapshots
Most bookings have about 8 events and need no snapshot. A long-lived entity such as a corporate account with 50,000 events would replay all of them on every command; snapshotting every 500 events caps a load at the snapshot plus up to 499 events.
Migration and schema evolution
You cannot run ALTER TABLE on history. When the meaning of an event changes you need:
- Versioned event types (
BookingCreated.v1,v2). - Upcasters: code that converts old event versions to the current shape when they are read.
- Occasionally a copy-and-transform migration of the whole store, which is a major, risky operation.
Every one of these must be tested against real historical events, not just new ones.
Eventual consistency
Projections lag behind the event store by milliseconds to seconds. A user who changes a booking and immediately reloads may see the old date unless the UI reads from the aggregate or waits for the projection to catch up. The product team must accept that behaviour or it must be designed around.
Organizational prerequisites I would require before approving
Testing strategy
- Given-when-then tests on aggregates: given these past events, when this command arrives, then these new events (or this rejection) result. These are the core of correctness and are fast.
- Projection tests that replay a fixed event sequence and assert the resulting read model.
- Upcaster tests against a sample of real production events from every historical version.
- A full-replay rehearsal in a staging environment on a production-sized copy, with a measured rebuild time, before go-live.
Developer tooling
- An event catalog: every event type, its versions, its schema, its owner.
- Tools to inspect a single stream ("show me every event for booking 123") for support and debugging.
- A projection rebuild tool with progress reporting and the ability to run a new projection alongside the old one.
- Monitoring of projection lag, with an alert when a projection falls behind.
Team readiness
- At least one or two engineers who have run an event-sourced system in production, or budgeted time for a proof of concept on a non-critical slice first.
- Agreement from support and product on eventual consistency and on how corrections are made (you append a compensating event (a new event that records the correction, rather than editing the old one); you never edit history).
- On-call runbooks for a stuck projection, a poison event (one that crashes a projection on every attempt), and a failed upcast.
The approval question itself
I would ask the team to answer, in writing: which requirement fails if we use a normal relational model plus an audit log table and publish integration events? If the answer is "none, but event sourcing is elegant", that is golden-hammer adoption (reaching for a favourite pattern regardless of fit) and I would decline. If the answer is "disputes require reconstructing exactly what the booking looked like at any moment, and we need five read models that disagree on shape", that is a real case.
Recommendation shape
- Approve for the booking aggregate only, not the whole platform; surrounding domains (customer profiles, content) stay conventional.
- Gate go-live on the staged replay rehearsal and the tooling above.
- Review after one quarter in production: projection lag, rebuild time, incident count, and how long a new engineer takes to ship their first change.
Pitfalls
- Event sourcing everything. Most domains do not need it; applying it platform-wide multiplies the operational cost.
- Using events as the integration contract. Internal events change as the model evolves; publish separate, stable integration events for other teams.
- Ignoring deletion and privacy until later. Immutable history and data-deletion obligations collide, and retrofitting encryption onto three years of events is painful.
- Treating the event store as a message queue. It is the system of record (the one authoritative place a fact is permanently and correctly stored, not a transient message that can be dropped or replayed); durability, backups and ordering guarantees matter more than throughput.
Your organization's microservice landscape has become chatty and tightly coupled: excessive cross-service calls, a few cyclic dependencies, and some services with very high coupling to others. As the engineering manager, produce a prioritized multi-phase refactor plan, quick wins plus risk-managed bigger changes, the metrics you'd track for stability and delivery velocity, and how you'd communicate and land the plan with your teams without stalling product delivery.
Sample Answer
Direct answer
I would run this as a funded, measured program in four phases, not a big-bang rewrite: measure first so we fix the couplings that actually hurt, then quick wins that cut call counts without changing ownership, then break the cycles, then the risk-managed structural changes (moving data ownership, merging or splitting badly drawn services), each behind a feature flag (a toggle that turns a new behavior on for a subset of traffic without a new deploy, so it can be switched back off instantly) and with a rollback path (a pre-agreed way to revert to the previous behavior if something goes wrong). It is funded as a fixed share of capacity (about 20% of each team's sprint) rather than a freeze, so product delivery continues. Success is judged on two sets of metrics tracked from the start: stability (incident rate, p99 latency, availability of the key user flows) and delivery velocity (deployment frequency, lead time (the time from a change merging to it running in production), share of releases needing coordination).
First, the vocabulary
- Chatty services: services that make many small calls to each other to serve one request, so latency and failure risk add up across calls.
- Cyclic dependency: A calls B and B (directly or through C) calls A. There is no safe order to deploy or restart them, and a slowdown in one feeds back into itself.
- High coupling: a service that many others depend on for internal details, or that depends on many others, so it cannot change safely.
Phase 0 (weeks 1 to 3): baseline and prioritise
Build a shared, objective picture before changing anything:
- From distributed traces (records that follow one user request as it moves across services, showing every downstream call it made and how long each took), list the top 20 user-facing endpoints by traffic and, for each, the number of downstream calls per request and the longest synchronous chain.
- From the service call graph, list every cycle.
- From the deploy log, list service pairs that deploy together most often.
- From incident reviews of the last two quarters, tag which incidents involved cascading calls or cycles.
Then rank each problem on pain (traffic affected, incidents caused, teams slowed) against effort and risk. The output is a one-page ranked backlog that product and engineering leads review together.
Worked example: why the baseline changes priorities
Suppose the order-history page makes 1 call to Orders plus 1 call to Catalog per order line to fetch product names, and a typical page shows 30 lines.
- Calls per page view: 1 + 30 = 31.
- If each Catalog call takes 8 ms and they are made one after another, that is 30 × 8 = 240 ms spent on Catalog alone.
- A batch endpoint (
getProducts(ids[])) turns 30 calls into 1. At, say, 15 ms for the batch call, the page drops about 225 ms and Catalog's request volume from this page drops 30-fold.
That is a two-sprint change for one team, and it would not have been obvious from an architecture diagram. The baseline is what surfaces it.
Phase 1 (weeks 3 to 8): quick wins
Low risk, no ownership changes, visible results that build trust in the program:
- Batch endpoints for loops of per-item calls (as above).
- Parallelise independent calls that are currently made one after another.
- Cache slow-changing reference data (product names, currency lists, configuration) locally with a short TTL (time-to-live, how long a cached value is trusted), instead of calling the owner on every request.
- Delete dead calls: traces often show calls whose results are never used.
- Aggregate for the client: where a web or mobile client makes 10 calls to render one screen, a backend-for-frontend (a thin service shaped for one client that composes the calls server-side) reduces client round trips. Keep it thin, or it becomes a new god service (a single service that has absorbed so many unrelated responsibilities that no one can change it safely).
Phase 2 (weeks 6 to 16): break the cycles
For each cycle, decide which direction of the dependency is legitimate and invert the other:
- Replace the back-call with an event. If Orders calls Notifications, and Notifications calls back into Orders to fetch order details, have Orders publish
OrderPlacedwith the fields Notifications needs. The back edge disappears. - Extract the shared piece. If A and B call each other because both need one capability, move that capability into its own module or service that both depend on.
- Merge. If two services are in a cycle because they are one capability split in two, merge them. This is often the cheapest fix and is not a failure.
Add a CI check that fails any change introducing a new cycle, so the count can only go down.
Phase 3 (quarter 2 onward): the risk-managed bigger changes
These change data ownership or service boundaries, so each gets an architecture decision record (a short written decision with alternatives and consequences) and a safe rollout pattern:
- Strangler approach: build the new path next to the old one, route a growing share of traffic to it behind a feature flag, and remove the old path only when the new one has run clean.
- Data ownership moves: when a highly coupled service owns data that several others read directly, give the data one owner, have others read through its API or keep local copies fed by events, and run old and new reads in parallel, comparing results, before cutting over.
- One structural change per team at a time, each with an explicit rollback step.
Metrics I would track
| Category | Metric | Source | Direction |
|---|---|---|---|
| Stability | Incidents involving cascading failures per month | Incident reviews | Down |
| Stability | p99 (99th percentile) latency and availability of the top 5 user flows | Monitoring | Latency down, availability up |
| Coupling | Downstream calls per request on top endpoints | Traces | Down |
| Coupling | Number of dependency cycles | Call graph | To zero |
| Velocity | Deployment frequency per service | Deploy log | Up |
| Velocity | Lead time from merge to production | Pipeline data | Down |
| Velocity | Share of releases that needed another service to ship at the same time | Deploy log | Down |
| Program health | Share of capacity actually spent on the program vs the planned 20% | Sprint data | Stable |
Deployment frequency and lead time are two of the DORA metrics (from the DevOps Research and Assessment program), which gives a common language with leadership. Report them monthly, with the baseline from Phase 0 on every chart.
Communicating and landing it without stalling delivery
- Frame it in product terms. Not "we are removing cycles" but "checkout availability and the time to ship a pricing change". Show the baseline numbers from Phase 0 to product leadership and agree the 20% allocation as an explicit trade-off.
- Capacity, not a freeze. Each team keeps 80% on roadmap work. Refactor items live in the same backlog as features, so trade-offs are visible, not hidden.
- Attach refactors to features. When a roadmap feature touches a coupled area, do the decoupling as part of that feature. This pays for itself and avoids "architecture sprints" nobody wants.
- Give teams ownership of their items. Each team owns the fixes in its services; a small working group (one engineer per team, plus a staff engineer: a senior individual-contributor engineer whose scope spans many teams) owns the shared plan and the cross-team changes.
- Show progress early. Phase 1 exists partly to produce visible latency and incident improvements within two months, which buys patience for Phase 3.
- Revisit monthly. If a phase is not moving its metric, stop and re-rank rather than finishing it out of momentum.
Pitfalls
- The rewrite. A "v2 platform" that pauses features for a year usually never lands and leaves two systems to maintain.
- Fixing what is visible rather than what hurts. The ugliest part of the diagram may carry little traffic; the baseline decides.
- Replacing sync calls with events that carry the same internal data model. Call count drops, but deploy coupling stays.
- Measuring only activity (tickets closed, services refactored) instead of outcomes (incidents, lead time).
Design an architectural governance process to detect and prevent anti-patterns (like distributed monoliths, chatty services, or shared databases) before they take hold at scale. What review checkpoints, health metrics, and enforcement mechanisms would keep this from becoming a bottleneck for teams?
Sample Answer
Direct answer
I would build governance as automation first, humans for the few decisions that are expensive to reverse. Most anti-patterns leave machine-detectable evidence (a service reading another service's tables, a call chain that keeps getting deeper, services that always deploy together), so those checks run continuously in CI (continuous integration: the automated pipeline that builds and tests every change before it is allowed to merge) and against production telemetry, the way tests do, and they give teams a fast, objective answer. Human review is reserved for a small set of triggers, such as a new service, a new cross-team synchronous dependency or a new shared datastore, and it runs with a published turnaround time so it never becomes a queue. Every rule has an owner, a metric, and an exception process with an expiry date.
Quick definitions of the three smells named in the question: a distributed monolith is a set of services that cannot be changed or deployed independently, so you pay the cost of a network without getting independence; chatty services make many small calls to each other to serve one request; a shared database is two or more services reading and writing the same tables, so neither can change its schema alone.
Why governance usually fails
Two failure modes, and the design has to avoid both:
- The architecture review board as a gate. Every change waits for a weekly meeting. Teams learn to route around it (calling a new service a "module", skipping the form), and the smells arrive anyway, just undocumented.
- Principles on a wiki. "Services must own their data" is written down and never checked. Six months later four services share an orders table.
The answer is to turn principles into fitness functions: automated checks that measure whether the architecture still has a property you care about, run as often as tests.
The process
flowchart TB
PR[Pull request or infra change] --> F[Fitness functions in CI]
F -->|pass| Merge[Merge]
F -->|violation| X{Approved exception on file?}
X -->|yes, not expired| Merge
X -->|no| Block[Build fails with rule link]
PR --> T{Hits a review trigger?}
T -->|no| Merge
T -->|yes| ADR[Architecture decision record plus async review, 3-day turnaround]
ADR --> Merge
Prod[Production traces and deploy log] --> H[Weekly health report per team]
H --> Q[Quarterly review of worst trends]
1. Review checkpoints (few, triggered, time-boxed)
Review is triggered by the kind of change, not by every change:
| Trigger | Why it matters | Output |
|---|---|---|
| New service proposed | Wrong boundaries are the root of chatty and god services (a god service is one service that has absorbed far more responsibility, and far more of the codebase's inbound calls, than any one team can safely own) | ADR with the capability owned and the data it owns |
| New synchronous dependency between teams | Each one adds a runtime and deploy coupling edge | ADR stating why an async event or local data copy will not do |
| Any service granted access to another's datastore | The shared-database anti-pattern starts here | Default answer is no; exception needs an end date |
| New shared library carrying domain types | Lockstep upgrades across consumers | ADR, owner, versioning policy |
An ADR (architecture decision record) is a one-page document: context, decision, alternatives considered, consequences. It is written by the team, reviewed asynchronously by one architect from a rotating pool (plus a peer from an affected team), and approved or pushed back within 3 working days. If the reviewer misses the deadline, the decision stands as written. That default is what stops review from becoming a bottleneck.
2. Automated checks (the enforcement layer)
Each anti-pattern maps to a check with a concrete data source:
| Anti-pattern | Check | Data source |
|---|---|---|
| Shared database | No service's database credentials can reach another service's schema | Database grants, infrastructure-as-code (defining and provisioning infrastructure, here database permissions, through versioned config files rather than manual changes) definitions, scanned in CI |
| Chatty services | Per endpoint, count of downstream calls per inbound request; alert above a threshold (say 10, chosen because this system's own traces show a typical endpoint makes 3 to 6 downstream calls, so 10 sits clearly above the busiest normal case, about 1.7 times the top of that range, not the full double it can read as) or a 25% rise over the previous release (which catches a regression even on an endpoint whose steady-state fan-out is already high) | Distributed traces (for example OpenTelemetry spans) |
| Cyclic dependencies | The service call graph has no cycles; a new edge that closes a cycle fails | Call graph built from traces or declared dependencies |
| Distributed monolith | Co-deploy rate: share of releases in which a service had to ship together with another | Deploy log |
| Leaky internal schema | Public API schemas may not reference internal table or entity types | Schema linting (an automated check that a schema definition follows a set of rules, run here against the public API's own schema file) on API definitions |
| God service | Number of distinct teams committing to one service per quarter; count of inbound dependencies (how many other services call it) | Version control history, call graph |
Concretely, one of these compiles to an actual query: the shared-database check runs something like SELECT grantee, table_schema FROM information_schema.role_table_grants WHERE table_schema NOT IN (SELECT owned_schema FROM service_registry WHERE service = grantee) against the database's own permission tables, expecting zero rows back; any row it returns names one real illegal grant, and CI fails the build listing it. The other rows in the table above compile the same way, into a database query, a static-analysis pass over the declared call graph, or a schema-linter rule, each with a pass condition as concrete as this one.
3. Health metrics (trends, not gates)
Some signals are too noisy to block a build but are exactly what leadership should watch. A weekly per-team report shows:
- Co-deploy rate (target below 10% of releases).
- Median and p95 (95th percentile) synchronous call depth (how many services deep a chain of blocking calls goes before the original request can complete) per user-facing request.
- Count of cross-service database grants (target zero, trending down).
- Count of open exceptions and how many are past their expiry.
- Delivery outcomes, so architecture is tied to results: deployment frequency and change lead time (two of the DORA metrics, from the DevOps Research and Assessment programme).
4. Enforcement with a ratchet
A new rule never starts as a hard block on a codebase that already violates it, or every team is instantly red:
- Measure: run the rule in report-only mode, publish the baseline (for example, 14 cross-service database grants today).
- Ratchet: the build fails if the number increases. Existing violations are grandfathered (excused from the new rule for now, because they predate it, but only as a dated, tracked exception, not a permanent pass) as dated exceptions.
- Burn down: each exception has an owner and an expiry; expired exceptions show in the weekly report and in quarterly planning.
- Harden: once the baseline hits zero, the rule becomes an absolute block.
Worked example: rolling out the "no cycles" rule
Suppose the trace-derived call graph for 40 services has 3 cycles, involving 7 services.
- Week 1: the check runs report-only; the 3 cycles are listed with the teams that own each edge.
- Week 2: the ratchet goes live. A pull request adding a call from Notifications to Orders would close a 4th cycle (Orders already calls Notifications), so it fails with a link to the rule and two suggested alternatives: Notifications subscribes to an
OrderPlacedevent, or Orders includes the needed fields in the call it already makes. - The 3 existing cycles become exceptions expiring in two quarters, placed on the owning teams' roadmaps.
- Two quarters later: if 2 of the 3 are gone and 1 remains, the remaining one is escalated in the quarterly review with a choice: fix, merge the two services (a cycle often means they were one capability), or extend with a written reason.
Keeping it from becoming a bottleneck
- Paved road over permission: publish templates (service skeleton with tracing, contract tests, its own database) so the compliant path is also the fastest path.
- Federated reviewers: one architect per group of teams, on rotation, not a central board.
- Time-boxed defaults: silence after the deadline means approved.
- Measure the process itself: median ADR turnaround and share of builds blocked by governance rules. If blocked builds exceed a few percent, the rules are too strict or too noisy and need tuning.
Pitfalls
- Metrics as targets without context. A team can cut "calls per request" by merging two endpoints into one god endpoint (the same god-service problem named above, now at the endpoint level: one endpoint absorbing so many responsibilities that it becomes the new coupling point). Review trends together with the ADRs that explain them.
- Rules without owners. A fitness function nobody maintains starts flapping (passing and failing intermittently on the same code, for reasons unrelated to whether the rule is actually being violated), teams add it to an ignore list, and the governance erodes silently.
- Governing everything. If every schema change needs review, the important reviews drown. Scope review to decisions that are expensive to reverse.
That is every published Architectural Patterns and Anti-Patterns question for Engineering Manager so far. Browse the other topics in this category, or practice this one interactively.