Architectural Patterns and Anti-Patterns Questions
Architecture-level patterns and the anti-patterns that signal a wrong turn. Patterns: layered and n-tier architecture, including where cross-cutting concerns like authentication, rate-limiting and tracing belong, dependency injection trade-offs, and thin-versus-fat controller design; hexagonal (ports and adapters) and clean architecture; CQRS and event sourcing; backend-for-frontend; plugin (microkernel) extension models; and the coupling, cohesion, encapsulation and separation-of-concerns principles behind them, including when each applies and what it costs. Anti-patterns: distributed monolith, chatty services, shared-database coupling, cyclic service dependencies, leaky abstractions that expose internal schemas, and golden-hammer pattern adoption. Covers the detection signals (deploy coupling, call-graph fan-out, change amplification, trace evidence), incremental remediation, and architecture governance that keeps smells from recurring. This is about diagnosing and fixing the smell in an existing design, not the monolith-versus-microservices decision itself.
Your team moved to microservices, but engineering velocity has stalled: releases need cross-team coordination, on-call keeps getting paged for cascading issues, and a schema change in one service quietly breaks two others. Which architectural anti-patterns do you suspect are at play? For each one, give a concrete detection signal you'd look for in monitoring or traces, and one pragmatic, incremental remediation step.
Sample Answer
Direct answer
Those three symptoms point to six suspects, including one that is worth checking even though it does not map cleanly to a single symptom below. Lockstep releases suggest a distributed monolith: services that are packaged and deployed separately, so they look like microservices, but are coupled tightly enough that they still have to change and ship together, the coordination cost of a monolith without the deployment independence a microservice split is supposed to buy. Cascading pages (one service's trouble tripping alerts through several other services it calls, or that call it, in a chain) suggest chatty services and long synchronous call chains, possibly with cyclic dependencies. A god service, one service that has absorbed so many responsibilities that it shows up in almost every trace and almost every incident, is that additional suspect: it can force lockstep releases because every dependent team waits on the team that owns it, and it can make pages cascade because it sits on the path of most requests. A schema change breaking two other services points to shared-database coupling or a leaky abstraction (an interface meant to hide a service's internal details, that lets them show through anyway) that exposes one service's internal schema to others. Confirm each with evidence from traces, deploy history and database access logs before acting, then fix incrementally, starting with the one that causes the most pages. Do not propose a rewrite.
Symptom to anti-pattern, detection signal and first remediation step
| Symptom | Suspected anti-pattern | Detection signal | One incremental remediation |
|---|---|---|---|
| Releases need cross-team coordination | Distributed monolith: independently deployed services that must still ship together | Deploy coupling: share of releases where 2 or more services had to deploy in the same window for one change; co-change in version control (the same ticket touching several repositories) | Make API changes backward compatible (add fields, never rename or remove in one step) and add consumer-driven contract tests (each consumer publishes the requests it relies on; the provider's CI fails if it breaks them), so each service can ship alone |
| On-call paged for cascading issues | Chatty services: many fine-grained calls to render one result | Trace span count per user request and fan-out (how many downstream calls one request triggers); many sequential calls to the same service inside one trace (the N+1 pattern: one call to fetch a list, then one extra call per item in that list, instead of a single batched call) | Replace the loop of per-item calls with one batch endpoint (GET /prices?ids=...), or keep a local read-only copy of slowly changing reference data |
| Same cascade, pages across many teams | Cyclic service dependencies: A calls B, B calls C, C calls A | Build the service call graph from trace data and look for cycles; a single incident paging several teams' services in a loop | Break one edge of the cycle: move the back-call to an event the upstream service publishes, so the dependency points one way |
| One service appears in almost every trace | God service: one service that owns too many responsibilities | High in-degree (a graph term for many incoming edges, here meaning many other services call into it) in the call graph; one repository with commits from many teams; its deploys correlate with others' incidents | Carve out the single most frequently changed responsibility behind its own interface first |
| Schema change in one service breaks two others | Shared-database coupling | Database logs or per-role query statistics (usage broken down by database login/credential, not by job role) showing several services' credentials reading or writing the same tables; incidents in other services following a migration | Give each table exactly one owning service; as a first step, publish a read-only database view or an API as the stable contract and move readers onto it |
| Same, via the API or events | Leaky abstraction exposing internal schema | API responses or events whose field names mirror table columns; change-data-capture streams (a feed of every row-level insert, update and delete, read straight off the database) of raw tables consumed directly by other teams | Introduce an explicit, versioned public model and map the internal schema to it, so internal columns can change freely |
Span: one timed operation in a distributed trace, such as one HTTP call. Trace: the tree of spans produced by one user request.
Worked example: putting numbers on the cascade
A checkout trace shows one user request producing 23 spans (a span is one timed operation in the trace, such as one HTTP call; 23 means some services are called more than once) across 7 services, with 6 of the calls made one after another, which is what makes the request slow. Separately from that latency picture, checkout's correctness depends on all 7 services succeeding, whether a given call sits on the sequential chain or not: if any one of them fails, the checkout fails. Suppose each of the 7 services is available 99.9% of the time and checkout needs all of them to succeed.
- Chained availability: 0.999 raised to the 7th power (one factor per required service) = 0.99302, about 99.30%. That looks like a small drop from 99.9%, because both numbers are close to 100%.
- Over a 30-day month (43,200 minutes), that is (1 - 0.99302) x 43,200, about 301 minutes of failed checkouts, versus 0.001 x 43,200 = 43.2 minutes for a single 99.9% service. 301 / 43.2 is about 7: it is the expected downtime, not the availability percentage, that scales roughly with the number of required services.
So a design where every request synchronously depends on seven services has about seven times the expected downtime of any one of them (301 minutes a month versus 43.2), before even counting the extra latency the 6 sequential calls add on top, even though the availability percentages themselves look close (99.30% versus 99.9%). The 23-span, 6-sequential-call shape is why the request is also slow; the 7-services-required shape is why it is fragile. That is the architectural reason on-call keeps getting paged: the dependency structure multiplies failure, and no single team's service looks broken.
For deploy coupling, pull the last quarter's release records. If, for example, 31 of 50 releases needed a coordinated deploy, that 62% is the number to drive down, and the co-change data shows which service pairs to fix first.
Sequencing the remediation
The aim is to reduce risk and restore delivery speed without a big-bang rewrite:
- Measure first. Record the baseline: deploy-coupling rate, pages per week by root-cause service, median spans per key request, number of services with write access to each shared table.
- Stop the bleeding on pages. Fix the worst chatty path or break the cycle on the most-paged flow; these are local changes with fast payback.
- Stop new coupling. Add contract tests in CI, and a rule that no new service gets credentials to another service's tables. This keeps the smell from recurring while you pay down the old debt.
- Pay down the shared database table by table, starting with the tables whose migrations caused incidents.
- Report progress in delivery terms: lead time for changes (the time from a commit to it running in production) and change failure rate (the share of production changes that cause an incident), two of the DORA metrics (the software delivery measures from the DevOps Research and Assessment programme), plus the deploy-coupling rate.
Pitfalls
- Armouring instead of fixing. Timeouts and circuit breakers (a guard that stops calling a failing dependency for a while, instead of piling retries onto a call that keeps failing) limit the blast radius (how much of the system one failure can drag down) of a bad dependency but leave the dependency in place; they are not a remediation of the architecture.
- Merging everything back. Sometimes two services that always change together should become one, and that is a legitimate fix, but decide it per pair from co-change evidence, not as a reaction to pain.
- Guessing from symptoms alone. The same "cascading pages" symptom can come from chatty calls or from a cycle; the trace graph tells you which.
- Fixing everything at once. Coordinated remediation across all teams recreates the coordination problem you are trying to remove.
Design an architectural governance process to detect and prevent anti-patterns (like distributed monoliths, chatty services, or shared databases) before they take hold at scale. What review checkpoints, health metrics, and enforcement mechanisms would keep this from becoming a bottleneck for teams?
Sample Answer
Direct answer
I would build governance as automation first, humans for the few decisions that are expensive to reverse. Most anti-patterns leave machine-detectable evidence (a service reading another service's tables, a call chain that keeps getting deeper, services that always deploy together), so those checks run continuously in CI (continuous integration: the automated pipeline that builds and tests every change before it is allowed to merge) and against production telemetry, the way tests do, and they give teams a fast, objective answer. Human review is reserved for a small set of triggers, such as a new service, a new cross-team synchronous dependency or a new shared datastore, and it runs with a published turnaround time so it never becomes a queue. Every rule has an owner, a metric, and an exception process with an expiry date.
Quick definitions of the three smells named in the question: a distributed monolith is a set of services that cannot be changed or deployed independently, so you pay the cost of a network without getting independence; chatty services make many small calls to each other to serve one request; a shared database is two or more services reading and writing the same tables, so neither can change its schema alone.
Why governance usually fails
Two failure modes, and the design has to avoid both:
- The architecture review board as a gate. Every change waits for a weekly meeting. Teams learn to route around it (calling a new service a "module", skipping the form), and the smells arrive anyway, just undocumented.
- Principles on a wiki. "Services must own their data" is written down and never checked. Six months later four services share an orders table.
The answer is to turn principles into fitness functions: automated checks that measure whether the architecture still has a property you care about, run as often as tests.
The process
flowchart TB
PR[Pull request or infra change] --> F[Fitness functions in CI]
F -->|pass| Merge[Merge]
F -->|violation| X{Approved exception on file?}
X -->|yes, not expired| Merge
X -->|no| Block[Build fails with rule link]
PR --> T{Hits a review trigger?}
T -->|no| Merge
T -->|yes| ADR[Architecture decision record plus async review, 3-day turnaround]
ADR --> Merge
Prod[Production traces and deploy log] --> H[Weekly health report per team]
H --> Q[Quarterly review of worst trends]
1. Review checkpoints (few, triggered, time-boxed)
Review is triggered by the kind of change, not by every change:
| Trigger | Why it matters | Output |
|---|---|---|
| New service proposed | Wrong boundaries are the root of chatty and god services (a god service is one service that has absorbed far more responsibility, and far more of the codebase's inbound calls, than any one team can safely own) | ADR with the capability owned and the data it owns |
| New synchronous dependency between teams | Each one adds a runtime and deploy coupling edge | ADR stating why an async event or local data copy will not do |
| Any service granted access to another's datastore | The shared-database anti-pattern starts here | Default answer is no; exception needs an end date |
| New shared library carrying domain types | Lockstep upgrades across consumers | ADR, owner, versioning policy |
An ADR (architecture decision record) is a one-page document: context, decision, alternatives considered, consequences. It is written by the team, reviewed asynchronously by one architect from a rotating pool (plus a peer from an affected team), and approved or pushed back within 3 working days. If the reviewer misses the deadline, the decision stands as written. That default is what stops review from becoming a bottleneck.
2. Automated checks (the enforcement layer)
Each anti-pattern maps to a check with a concrete data source:
| Anti-pattern | Check | Data source |
|---|---|---|
| Shared database | No service's database credentials can reach another service's schema | Database grants, infrastructure-as-code (defining and provisioning infrastructure, here database permissions, through versioned config files rather than manual changes) definitions, scanned in CI |
| Chatty services | Per endpoint, count of downstream calls per inbound request; alert above a threshold (say 10, chosen because this system's own traces show a typical endpoint makes 3 to 6 downstream calls, so 10 sits clearly above the busiest normal case, about 1.7 times the top of that range, not the full double it can read as) or a 25% rise over the previous release (which catches a regression even on an endpoint whose steady-state fan-out is already high) | Distributed traces (for example OpenTelemetry spans) |
| Cyclic dependencies | The service call graph has no cycles; a new edge that closes a cycle fails | Call graph built from traces or declared dependencies |
| Distributed monolith | Co-deploy rate: share of releases in which a service had to ship together with another | Deploy log |
| Leaky internal schema | Public API schemas may not reference internal table or entity types | Schema linting (an automated check that a schema definition follows a set of rules, run here against the public API's own schema file) on API definitions |
| God service | Number of distinct teams committing to one service per quarter; count of inbound dependencies (how many other services call it) | Version control history, call graph |
Concretely, one of these compiles to an actual query: the shared-database check runs something like SELECT grantee, table_schema FROM information_schema.role_table_grants WHERE table_schema NOT IN (SELECT owned_schema FROM service_registry WHERE service = grantee) against the database's own permission tables, expecting zero rows back; any row it returns names one real illegal grant, and CI fails the build listing it. The other rows in the table above compile the same way, into a database query, a static-analysis pass over the declared call graph, or a schema-linter rule, each with a pass condition as concrete as this one.
3. Health metrics (trends, not gates)
Some signals are too noisy to block a build but are exactly what leadership should watch. A weekly per-team report shows:
- Co-deploy rate (target below 10% of releases).
- Median and p95 (95th percentile) synchronous call depth (how many services deep a chain of blocking calls goes before the original request can complete) per user-facing request.
- Count of cross-service database grants (target zero, trending down).
- Count of open exceptions and how many are past their expiry.
- Delivery outcomes, so architecture is tied to results: deployment frequency and change lead time (two of the DORA metrics, from the DevOps Research and Assessment programme).
4. Enforcement with a ratchet
A new rule never starts as a hard block on a codebase that already violates it, or every team is instantly red:
- Measure: run the rule in report-only mode, publish the baseline (for example, 14 cross-service database grants today).
- Ratchet: the build fails if the number increases. Existing violations are grandfathered (excused from the new rule for now, because they predate it, but only as a dated, tracked exception, not a permanent pass) as dated exceptions.
- Burn down: each exception has an owner and an expiry; expired exceptions show in the weekly report and in quarterly planning.
- Harden: once the baseline hits zero, the rule becomes an absolute block.
Worked example: rolling out the "no cycles" rule
Suppose the trace-derived call graph for 40 services has 3 cycles, involving 7 services.
- Week 1: the check runs report-only; the 3 cycles are listed with the teams that own each edge.
- Week 2: the ratchet goes live. A pull request adding a call from Notifications to Orders would close a 4th cycle (Orders already calls Notifications), so it fails with a link to the rule and two suggested alternatives: Notifications subscribes to an
OrderPlacedevent, or Orders includes the needed fields in the call it already makes. - The 3 existing cycles become exceptions expiring in two quarters, placed on the owning teams' roadmaps.
- Two quarters later: if 2 of the 3 are gone and 1 remains, the remaining one is escalated in the quarterly review with a choice: fix, merge the two services (a cycle often means they were one capability), or extend with a written reason.
Keeping it from becoming a bottleneck
- Paved road over permission: publish templates (service skeleton with tracing, contract tests, its own database) so the compliant path is also the fastest path.
- Federated reviewers: one architect per group of teams, on rotation, not a central board.
- Time-boxed defaults: silence after the deadline means approved.
- Measure the process itself: median ADR turnaround and share of builds blocked by governance rules. If blocked builds exceed a few percent, the rules are too strict or too noisy and need tuning.
Pitfalls
- Metrics as targets without context. A team can cut "calls per request" by merging two endpoints into one god endpoint (the same god-service problem named above, now at the endpoint level: one endpoint absorbing so many responsibilities that it becomes the new coupling point). Review trends together with the ADRs that explain them.
- Rules without owners. A fitness function nobody maintains starts flapping (passing and failing intermittently on the same code, for reasons unrelated to whether the rule is actually being violated), teams add it to an ignore list, and the governance erodes silently.
- Governing everything. If every schema change needs review, the important reviews drown. Scope review to decisions that are expensive to reverse.
What is a 'distributed monolith' anti-pattern? Describe two signs that a microservices deployment has effectively become a distributed monolith, and propose one concrete mitigation step.
Sample Answer
Direct answer
A distributed monolith is a system that is split into separately deployed services but still behaves like one program: the services cannot be changed, tested or released independently. You pay every cost of a distributed system (network latency, partial failures, harder debugging) and get none of the main benefit, which is independent teams shipping independently. Two strong signs are lockstep deployments (a feature routinely requires several services to release together, in a specific order) and long synchronous call chains (one user request cannot complete unless many services are all up at the same moment). A concrete first mitigation is to make deploys independent with versioned, backward-compatible APIs checked by consumer-driven contract tests, so each service can ship without waiting for the others.
What it looks like
Think of an online store split into Cart, Pricing, Inventory, Orders and Payments. On paper these are five microservices. In practice:
- Adding a "gift wrap" option means changing all five, and the release plan is a spreadsheet saying "deploy Pricing first, then Orders, then Cart, within 30 minutes, or checkout breaks".
- Placing an order calls Cart → Pricing → Inventory → Orders → Payments, one after another, synchronously.
- Orders and Inventory read and write the same
productstable.
Sign 1: lockstep deployment
What you observe: releases are coordinated across services; there is a shared release train (a fixed, recurring release schedule that several services are bundled onto and ship together, rather than each shipping on its own whenever its change is ready) or a "deploy order" document; rolling back one service requires rolling back others.
Why it happens: services share a data model or API that changes without versioning. If Pricing renames a field in its response, Cart breaks until Cart is updated, so they must deploy together.
How to detect it: from the deploy log, compute the fraction of releases in which a service was deployed within the same short window as another service for the same ticket. If Cart ships with Pricing in most of its releases, those two are one deployable in disguise.
Sign 2: synchronous call chains that multiply failure
What you observe: a distributed trace (a record, stitched together from the timestamped spans each service emits for the same request, showing every service that request touched and how long each one took) of one request shows a deep chain of blocking calls, and an outage in any one service takes down the whole user flow.
Why it matters, with numbers: if a request needs 5 services in series and each is independently available 99.9% of the time, the chain is available only when all 5 are up:
- 0.999 × 0.999 × 0.999 × 0.999 × 0.999 = 0.999^5 ≈ 0.99501, about 99.50%.
- Over a 30-day month (30 × 24 × 60 = 43,200 minutes), one service at 99.9% is down about 0.001 × 43,200 = 43.2 minutes.
- The 5-service chain is down about (1 − 0.99501) × 43,200 ≈ 215.6 minutes, roughly five times as much.
A monolith would have had one component's worth of downtime; the distributed monolith has five, with the same features. (This simple model assumes failures are independent; correlated failures change the number but not the lesson.)
Other signs worth naming
- Services sharing a database schema, so a column change needs coordination.
- A shared library of domain objects that every service must upgrade together.
- End-to-end tests that need every service running before anyone can merge.
The mitigation: break deploy coupling first
Why start here: lockstep deploys are the most expensive symptom (they slow every team, every release) and fixing them is incremental. Steps:
- Version the contract. API changes become additive: add a new field, never rename or remove one in place. Consumers use a tolerant reader (ignore fields you do not recognise, do not fail on missing optional ones).
- Add consumer-driven contract tests. Each consumer (Cart) publishes the exact requests and response fields it depends on; the provider (Pricing) runs those contracts in its own CI. Tools such as Pact implement this. Now Pricing learns in its own build, before deploying, whether it would break Cart.
- Deploy in expand-then-contract order. To rename
pricetounit_price: Pricing ships a version returning both; Cart switches tounit_pricewhenever it likes; once contracts show no consumer readsprice, Pricing removes it. Three independent deploys, no coordination window. - Measure the result. Track the lockstep share from the deploy log. If Cart and Pricing co-deployed in 8 of 10 releases before, the target is near zero within a couple of quarters.
If the contracts reveal that two services cannot evolve separately no matter what (they change together on every feature), the honest fix is to merge them back into one service. A boundary that always moves together was never a real boundary.
Pitfalls
- Fixing the symptom with tooling. A better release-orchestration tool makes lockstep deploys easier to run, which hides the coupling instead of removing it.
- Adding more services. Splitting further to "fix" coupling usually adds more edges.
- Replacing sync calls with async messages carrying the same shared data model. The deploy coupling survives; it just moved into the message schema.
That is every published Architectural Patterns and Anti-Patterns question for Cloud Architect so far. Browse the other topics in this category, or practice this one interactively.