Architectural Patterns and Anti-Patterns Questions
Architecture-level patterns and the anti-patterns that signal a wrong turn. Patterns: layered and n-tier architecture, including where cross-cutting concerns like authentication, rate-limiting and tracing belong, dependency injection trade-offs, and thin-versus-fat controller design; hexagonal (ports and adapters) and clean architecture; CQRS and event sourcing; backend-for-frontend; plugin (microkernel) extension models; and the coupling, cohesion, encapsulation and separation-of-concerns principles behind them, including when each applies and what it costs. Anti-patterns: distributed monolith, chatty services, shared-database coupling, cyclic service dependencies, leaky abstractions that expose internal schemas, and golden-hammer pattern adoption. Covers the detection signals (deploy coupling, call-graph fan-out, change amplification, trace evidence), incremental remediation, and architecture governance that keeps smells from recurring. This is about diagnosing and fixing the smell in an existing design, not the monolith-versus-microservices decision itself.
Describe a typical three-tier (multi-tier) layered web architecture: presentation/UI, application/business logic, and persistence/data. For each tier, name its responsibilities, then trace how a single request flows through the system, from the client through ingress, load balancing, and the application tier down to the data tier and back. What are the trade-offs of this architecture (scalability, deployment complexity, testability) compared to a simpler two-tier design or a flatter, event-driven one?
Sample Answer
Direct answer
A three-tier architecture splits a web application into three independently deployed parts: the presentation tier (what runs in or is sent to the user's browser or app), the application tier (the servers that apply business rules) and the data tier (the databases and stores that keep state). Each tier talks only to its neighbour, so the browser never touches the database directly. You pay for an extra network hop and more moving parts, and in exchange you get a stateless middle tier you can scale by adding servers, one place to enforce security and rules, and parts you can test and deploy separately.
A useful distinction up front: a tier is a physical deployment unit (separate machines or containers), while a layer is a logical grouping of code inside one program. Three-tier is about tiers.
The three tiers and what each owns
| Tier | Responsibilities | Typical technology | Should NOT do |
|---|---|---|---|
| Presentation | Render screens, capture input, client-side validation for fast feedback, call the API | React or server-rendered HTML, mobile app, static assets on a CDN (content delivery network: servers near users that cache files) | Hold business rules or database credentials |
| Application | Authenticate and authorize, validate input authoritatively, apply business rules, coordinate transactions, call other services, shape responses | Stateless API servers (Java, Go, Node, Python) behind a load balancer | Keep user session state in local memory (breaks horizontal scaling) |
| Data | Store and retrieve durable state, enforce integrity (keys, constraints), run indexed queries, replicate and back up | PostgreSQL or MySQL, plus a cache such as Redis and object storage for files | Encode business workflows in stored procedures that nobody can test |
Tracing one request: GET /orders/123
- Client. The browser already loaded the presentation code from the CDN. The user opens "My orders", and the JavaScript sends
GET /api/orders/123over HTTPS with a session token. - Ingress and load balancing. DNS resolves the API hostname to a load balancer (or, in Kubernetes (a container-orchestration system: software that runs and manages many server instances across a fleet of machines for you, restarting failed ones and scheduling new ones), an ingress controller: the component that admits outside traffic into the cluster). It terminates TLS (decrypts HTTPS), checks which application instances are passing health checks (a periodic ping each instance must answer, so the load balancer stops sending traffic to one that has stopped responding), and forwards the request to one of them.
- Application tier. The chosen instance validates the token, checks the rule "a customer may only read their own orders", and asks its data-access code for order 123. It may check a cache first.
- Data tier. A connection is borrowed from the instance's connection pool (a set of already-open database connections the instance keeps ready, so a request reuses one instead of paying the cost, tens of milliseconds, of opening a fresh connection to the database each time), an indexed
SELECTruns against the orders table, and rows come back. - Back up the chain. The application maps rows to a JSON response (dropping internal columns), returns 200, the load balancer relays it, and the browser renders the page.
sequenceDiagram
participant B as Browser
participant LB as Load balancer
participant App as App server
participant DB as Database
B->>LB: GET /api/orders/123 over HTTPS
LB->>App: forward to a healthy instance
App->>App: check token and ownership rule
App->>DB: SELECT order 123
DB-->>App: rows
App-->>LB: 200 JSON
LB-->>B: 200 JSON
Because step 3 keeps no per-user state in memory, any instance can serve the next request from the same user. That property is what lets the middle tier scale out.
Worked example: where the bottleneck moves
Suppose peak load is 1,200 requests per second (RPS) and a load test shows one application instance sustains about 150 RPS.
- Instances needed: 1,200 / 150 = 8. Run 10 so that losing one or two still leaves capacity (25% headroom).
- Each instance keeps a database connection pool of 20 connections: 10 x 20 = 200 connections.
- PostgreSQL's
max_connectionsdefault is typically 100 (the documentation notes it can be lower if kernel settings do not support 100).
So the design that scales the application tier by adding boxes will exhaust the data tier at peak, exactly when the autoscaler (the system that automatically adds or removes application instances based on load) adds instances. Fixes: shrink each pool (10 x 8 = 80 connections), or put a connection pooler such as PgBouncer between the tiers: a lightweight service that holds a smaller, fixed number of real connections to Postgres and multiplexes many application-side requests through them. PgBouncer's own default is session pooling, which hands one real connection to an application connection for its whole session and would not shrink the connection count here; the setting this design actually needs is transaction pooling mode, which hands a connection out per transaction and takes it back the instant that transaction commits or rolls back, then reuses it for the next request. PgBouncer also offers a statement pooling mode that returns the connection after every individual query, but that mode explicitly disallows transactions spanning multiple statements, which would break a multi-step operation like a funds transfer, so transaction mode, not the default and not statement mode, is the setting to reach for. That lets the same 200 application-side connections be served through, say, 50 real database connections instead of colliding with the 100-connection ceiling. The general lesson is that the application tier scales horizontally cheaply and the data tier does not, so in a three-tier system the bottleneck usually migrates to the database.
Trade-offs versus two-tier and event-driven designs
A two-tier design has the client talk straight to the database (a desktop app with a database driver, or a small server-rendered app where UI code and SQL live in one process). An event-driven design has components publish events ("OrderPlaced") to a broker such as Kafka (a system dedicated to durably queueing and delivering these events to every interested component, so the publisher does not need to know who is listening or whether they are online right now), and other components react asynchronously instead of being called directly.
| Dimension | Two-tier | Three-tier | Event-driven |
|---|---|---|---|
| Scalability | Every client holds database connections; the database takes all load | Stateless middle tier scales out; the database remains the limit | Consumers scale independently; the broker absorbs spikes |
| Deployment complexity | Lowest: one app plus a database, but rolling out a desktop client means updating every install | Moderate: load balancer, app fleet, database, each with its own deploy | Highest: broker, topics, consumers, schema versioning for events |
| Testability | Rules mixed with UI or SQL are hard to test in isolation | Business rules testable in the application tier without a UI | Each consumer testable alone, but end-to-end flows are hard to test and debug |
| Security | Database credentials sit on or near the client | Database reachable only from the app tier | Must also secure the broker and every topic |
| Time to market | Fastest for small internal tools | Reasonable default for most web products | Slowest to start; pays off when work can happen later |
| Failure behaviour | Database down means everything is down | A slow database slows every request synchronously | A slow consumer builds a backlog but the user-facing request still succeeds; data is eventually consistent (other parts catch up after a delay) |
Recommendation. Use three-tier as the default for a customer-facing web product. Use two-tier only for small internal tools with a handful of trusted users. Add event-driven pieces for work the user does not need to wait for (emails, search indexing, analytics), which gives a hybrid: synchronous three-tier for the request path, events for side effects. What would flip this: if most work is naturally asynchronous (ingesting sensor data, processing uploads), an event-driven core is the better starting point.
Pitfalls that separate strong answers
- Stateful application servers. Sessions kept in instance memory force "sticky" routing (the load balancer must keep sending one user to the same instance every time, because that instance is the only one holding their session in memory), and users get logged out when an instance dies. Keep sessions in a token or a shared store.
- A pass-through middle tier. If the application tier only forwards requests to SQL, you paid for a hop and got nothing; the rules have leaked into the client or into stored procedures (business logic written and run inside the database itself, rather than in the application tier).
- Confusing tiers with layers. Splitting every code layer onto its own network hop adds latency without adding isolation you need.
- Forgetting that tiers fail together. In a synchronous chain, a slow data tier makes the whole request slow; the architecture does not isolate that by itself.
Design patterns can become anti-patterns when they're reached for out of habit rather than fit. Describe a realistic case where adopting a popular pattern (at the service or architecture level) created more complexity or risk than it solved. What made it the wrong fit for that context, and what would have signaled the mismatch earlier?
Sample Answer
Direct answer
A realistic case is a small product team that adopted backend-for-frontend (BFF), one dedicated API service per client type, because it was "the recommended pattern when you have multiple frontends". BFF solves a real problem: very different clients, owned by separate teams, that need very different data shapes. This team had none of those forces. Its web and mobile apps showed nearly the same screens and were built by the same engineers, so the pattern mostly bought triplicated code, drift between clients and extra operational load. The mismatch was visible early in measurable signals: nearly identical endpoints across the BFFs, business rules leaking into them, and every feature needing several pull requests.
This is the golden-hammer anti-pattern: choosing a familiar or fashionable tool before naming the problem it is meant to solve.
The case (story skeleton)
Context. A seven-engineer team builds a B2B (business-to-business) ordering product: a web app for office staff and iOS and Android apps for field sales. One core API service owns orders, pricing and customers.
Decision. Early on, the team adds three BFF services (web, iOS, Android), each aggregating calls to the core API and shaping responses for its client. The design doc cites the pattern and large companies using it; it does not state which client need the core API fails to meet.
What happened over the following months.
- Change amplification: a new field on the order screen needed a change in the core API plus the same change in three BFFs: four pull requests, four deploys, four sets of tests for one user story.
- Logic drift: "just for now", a discount-eligibility check was added in the web BFF to unblock a release. The mobile BFFs never got it, so mobile users briefly saw prices that the web app correctly refused. The bug came from business logic living in an aggregation layer, which a BFF is not supposed to hold.
- Operational load: three extra services to patch, scale, monitor and put on call, for a team with no dedicated platform engineer.
- Latency: every call now crossed an extra network hop for responses the core API could have served directly.
Remediation. The team collapsed the three BFFs into the core API, which gained optional field selection (clients request the fields they need) and a couple of screen-oriented composite endpoints. The discount rule moved into the pricing domain code where every client gets it. A single BFF was kept later, but for a genuinely different consumer: a partner integration with its own authentication, rate limits and response format.
Why it was the wrong fit
Every pattern resolves specific forces. Checking them against this context:
| Force BFF resolves | Present here? |
|---|---|
| Clients need substantially different data shapes or call patterns | No: the same screens on every client |
| Separate teams own separate clients and need to ship independently | No: one team built everything |
| Client-specific concerns (device-specific auth flows, payload limits, protocol differences) | Minimal |
| A general-purpose API that is hard to change | No: the same team owned the core API |
With none of the forces present, the pattern contributed only its costs: more deployables, more duplicated code and a new place for business logic to hide. The value of a pattern is always relative to the problem; out of context, what remains is its overhead.
Signals that would have exposed the mismatch earlier
- The design doc named the pattern before naming the problem. A good architecture decision record (ADR, a short document capturing context, decision and consequences) states the forces first. "Because Netflix does it" is not a force.
- Endpoint similarity. Diffing the three BFFs' routes and response shapes would have shown them to be nearly identical. If most endpoints are copies, the per-client split has no purpose.
- Change amplification per story. Tracking how many repositories or services a typical feature touches: one story needing four pull requests is a structural smell, not a process one.
- Business logic outside the domain. Any
ifabout pricing, eligibility or permissions in an aggregation layer is a warning; a lint rule or review checklist can flag it. - Ownership mismatch. In Sam Newman's original framing, a BFF is ideally owned by the team that owns the client. One team owning all of them is a hint the split is not buying autonomy.
- Operational cost against team size. Services per engineer rising without a matching rise in independent release cadence.
A lightweight guard against golden hammers
Before adopting any architecture-level pattern, answer four questions in writing:
- What specific force or pain does it resolve, and is that pain happening now?
- What does it cost to build, run and staff?
- What is the simplest alternative, and why is it insufficient?
- How hard is it to reverse, and what signal would tell us to reverse it?
If question 1 has no concrete answer, defer the pattern. Adding a BFF later, when a genuinely divergent client appears, is cheap; removing three that have accumulated business logic is not.
Trade-offs and pitfalls
- The overcorrection: refusing all patterns leads to the opposite smell, one general API contorted to serve incompatible clients. The point is fit, not avoidance.
- Fixing the symptom: moving the discount rule into all three BFFs "for consistency" would have hidden the drift while cementing the duplication.
- Sunk cost: teams keep an ill-fitting pattern because removing it feels like admitting a mistake. Treat reversals as a normal outcome of an ADR with an explicit reversal signal.
Explain dependency injection (DI) and the differences between constructor injection, setter injection, and the service-locator approach. When designing a layered backend, what are the advantages and the potential pitfalls of leaning on a DI framework in your service and repository layers?
Sample Answer
Direct answer
Dependency injection (DI) means an object receives the collaborators it needs (a repository, a clock, an HTTP client) from outside, instead of creating or looking them up itself. Constructor injection passes required dependencies as constructor arguments, so the object is complete the moment it exists. Setter injection assigns dependencies after construction through setter methods, which suits optional ones. A service locator is the opposite direction: the object asks a global registry for what it needs at run time, which hides its dependencies. Use constructor injection by default, setters only for genuinely optional dependencies, and keep service locators out of business code.
The three approaches compared
| Constructor injection | Setter injection | Service locator | |
|---|---|---|---|
| How dependencies arrive | Constructor parameters | Setter methods after construction | Object calls Registry.get("x") itself |
| Are dependencies visible? | Yes, in the signature | Partly (must read the setters) | No, buried in method bodies |
| Can the object exist half-built? | No | Yes, until every setter runs | Looks complete, fails later |
| Immutability | Fields can be final or read-only | Fields must stay mutable | N/A |
| When a dependency is missing | Fails at construction (or at application startup with a container) | Null reference at first use | Lookup error at first call, possibly in production |
| Good for | Required collaborators | Optional collaborators with a sensible default | Framework glue at the edge of the application |
Some frameworks also allow field injection (the container writes directly into a private field). It hides dependencies like a locator does and makes plain unit testing awkward; Spring's own reference documentation recommends constructor injection for mandatory dependencies and setters for optional ones.
Worked example: why hidden dependencies hurt
This runnable Python program builds the same service two ways.
class Registry: # a service locator: a global lookup table
_services = {}
@classmethod
def register(cls, name, obj):
cls._services[name] = obj
@classmethod
def get(cls, name):
return cls._services[name]
class OrderServiceWithLocator:
def total(self, order_id):
repo = Registry.get("order_repo") # hidden dependency
return sum(repo.line_amounts(order_id))
class OrderServiceWithConstructor:
def __init__(self, repo): # dependency is visible in the signature
self.repo = repo
def total(self, order_id):
return sum(self.repo.line_amounts(order_id))
class FakeOrderRepo:
def line_amounts(self, order_id):
return [1200, 350]
print("constructor:", OrderServiceWithConstructor(FakeOrderRepo()).total("o1"))
svc = OrderServiceWithLocator() # constructs fine, nothing looks wrong yet
try:
print("locator:", svc.total("o1"))
except KeyError as e:
print("locator failed at call time, missing:", e)
Registry.register("order_repo", FakeOrderRepo())
print("locator after global setup:", svc.total("o1"))
Output:
constructor: 1550
locator failed at call time, missing: 'order_repo'
locator after global setup: 1550
The constructor version tells you from its signature that it needs a repository, and a test simply passes a fake (1,200 + 350 = 1,550). The locator version constructs without complaint and fails only when total runs; to test it you must mutate global state, which leaks between tests that run in the same process.
Advantages of a DI framework in the service and repository layers
- Wiring large object graphs. A container (Spring, .NET's built-in container, Guice, Dagger) builds hundreds of objects in the right order so you do not hand-write the assembly code.
- Lifecycles and scopes. It manages which objects are shared for the whole application (singleton), created per request, or created each time.
- Environment-specific configuration. Swap the real payment client for a sandbox client in staging through configuration, not code changes.
- Declarative cross-cutting behaviour. Transactions or metrics applied by annotation (a marker attached to code, such as
@Transactional, that tells the framework to wrap it with extra behaviour automatically) around service methods. - Testability. Because services receive their repositories, tests hand them fakes, and the business logic never needs a database to be exercised.
Pitfalls to watch for
- Framework leaking into the core. Annotations and container types in domain classes couple your business rules to the framework; hexagonal and clean architecture (architectural styles that keep business rules as plain code with no framework dependencies, pushing frameworks, databases and other details to the outside) both require that dependencies point inward, toward framework-free code. Keep the core as plain objects and do the wiring in one composition root (the single place at startup where the object graph is assembled).
- Scope mismatch. A singleton service that receives a request-scoped object (one the container creates fresh for each incoming request and discards afterward) at startup keeps that one instance forever, so data from one request can bleed into another. Containers offer proxies (a lightweight stand-in injected into the singleton that, on every call, fetches whichever request-scoped instance is current) or providers (a small factory object injected instead of the real dependency, which the singleton calls each time it needs a fresh, current instance) for this; you have to know to use them.
- Circular dependencies. With constructor injection, a cycle (A needs B, B needs A) cannot be constructed and the container fails at startup, which is good: it exposes a design problem. Setter or field injection can let the container paper over the cycle.
- Hidden god classes (classes that quietly accumulate far more responsibility than they should, because adding one more dependency is so easy). Adding a dependency costs one more constructor parameter, so bloat creeps in silently. A service with 10 or more injected collaborators is a cohesion smell (a warning sign that a class is doing too many unrelated things to be well designed); split it.
- Magic and startup cost. Classpath scanning (the container inspecting every class in the application at startup to find the ones it should manage) and reflection (code that inspects or calls into other code at run time, by name, rather than at compile time) make it hard to answer "where does this object come from?" and slow startup, which matters for serverless (functions that start fresh for each invocation and shut down between them, so every millisecond of startup cost is paid repeatedly) or short-lived processes. Compile-time DI (Dagger) (wiring code generated when the project is built, instead of discovered by scanning and reflection at startup) or plain manual wiring avoids both.
- Interfaces for everything. Creating an interface for every class "for DI" adds ceremony. Introduce an interface where there is a real second implementation or a test double worth having.
Recommendation: constructor injection everywhere in service and repository code; a container only at the outer layer; plain manual wiring for small services, where a 20-line composition root is clearer than a framework.
You are designing a new SaaS web application with a frontend, API gateway, application layer, business logic layer, and persistence layer. For each layer, describe its core responsibilities, what interfaces it should expose, and how you'd enforce separation of concerns to ease testing, deployment, and security boundaries across teams.
Sample Answer
Direct answer
I would treat the five layers as logical boundaries with a strict dependency direction, not as five separately deployed services. As a request flows through the system, each layer calls only the layer directly beneath it (frontend to gateway, gateway to application) or an interface owned by an inner layer (application to persistence, through a repository interface the domain layer defines), never sideways, and never back up to a layer closer to the user. All of those calls still obey one further rule about dependencies, which is a separate thing from calls: dependencies point inward toward the business rules, meaning the domain layer is the one thing every other layer's code is allowed to depend on (import, call, or implement an interface from), and the domain layer depends on nothing outside itself. The next section spells out what "inward" means and why that lets application call persistence directly without breaking the rule. The business logic layer knows nothing about HTTP or SQL. That gives each concern one home: the gateway handles cross-cutting edge concerns (things every request needs regardless of which use case it is, such as authentication and TLS, handled once at the boundary instead of being repeated inside every use case), the application layer orchestrates use cases, the domain layer holds the rules, and persistence hides the database. I enforce the boundaries mechanically (build-time architecture tests, separate modules, explicit data-transfer objects) rather than by convention, because conventions erode under deadline pressure.
flowchart TB
FE[Frontend: SPA on CDN] -->|HTTPS + JSON contract| GW[API gateway]
GW -->|authenticated request + tenant/user context| APP[Application layer: use cases]
APP -->|calls domain objects/services| DOM[Business logic: domain model]
APP -->|repository interfaces| PER[Persistence adapters]
PER --> DB[(Postgres)]
PER -.implements interfaces owned by.-> DOM
Two labels in that diagram are worth spelling out before going further. SPA means single-page application: a browser app that loads once and then updates itself by calling the API, rather than requesting a fresh HTML page for every click. Persistence adapters are the concrete, database-specific code (here, Postgres-specific) that implement an interface the domain layer defines; "adapter" is the standard name for that kind of outer-layer implementation, and the next paragraph explains the pattern it belongs to.
The dotted line is the important one: the domain layer defines the repository interface ("load this invoice", "save this invoice"), and persistence implements it. That is dependency inversion, and it is what lets the business rules be tested with no database.
This pattern has a name worth knowing: ports and adapters (also called hexagonal architecture). A port is an interface the domain defines for something it needs from the outside world (here, "a place to load and save invoices"); an adapter is the concrete, outer-layer code that plugs into that port (here, the Postgres-specific code in persistence). Picture the domain sitting at the centre, with every other layer forming a ring around it: "inward" means every layer's code is allowed to depend on the layer closer to the centre, never the reverse. That reconciles the two rules above: the runtime call from application to persistence looks like it reaches past the domain layer, but the interface it calls through is one the domain owns, so the dependency created by that call still points at the domain, not around it. Persistence, for its part, imports and implements that domain-owned interface, so its dependency also points inward, toward the centre, even though the request itself flows outward from application to persistence to the database.
Layer by layer
| Layer | Core responsibilities | Interface it exposes | Must NOT do |
|---|---|---|---|
| Frontend | Rendering, client-side validation for UX, navigation, local UI state | Nothing to other layers; consumes the public API contract | Enforce security or business rules (it runs on the user's machine and can be bypassed) |
| API gateway | TLS termination (decrypting incoming HTTPS at the gateway, so the connection arrives encrypted from the internet but the layers behind the gateway can talk plain HTTP internally), authentication (validating the token), rate limiting, request size limits, routing, request IDs for tracing | The public HTTP API (documented with OpenAPI, a machine-readable API description format) | Business decisions, data shaping per use case, database access |
| Application layer | One handler per use case ("CreateInvoice"): authorization in context (may this user act on this tenant's invoice?), transaction boundaries, orchestration of domain objects and repositories, mapping between API DTOs (data-transfer objects, plain request/response shapes) and domain objects | Use-case functions or command/query handlers that take DTOs and return DTOs | Contain business rules (it coordinates them), know SQL |
| Business logic (domain) | Entities and rules: an invoice cannot be issued with zero line items, tax is computed per jurisdiction, a paid invoice cannot be edited | Domain objects and domain services; repository and external-service interfaces (ports) it needs | Import any web framework, ORM (object-relational mapping library), or HTTP client |
| Persistence | Mapping domain objects to tables, queries, migrations, connection handling, tenant scoping as defence in depth (a second, independent enforcement of tenant isolation here, in case the application layer above ever forgets its own check; it backs up that check, it does not replace it) | Implementations of the repository interfaces | Leak ORM entities or table shapes upward; decide business outcomes |
Worked example: "issue invoice"
- The frontend sends
POST /invoices/881/issue. - The gateway validates the JSON Web Token (JWT, the signed login token), applies the tenant's rate limit, attaches a request ID, and forwards the call with
user_idandtenant_idas trusted context. - The application layer's
IssueInvoicehandler checks that this user has the billing role on this tenant, opens a transaction, and loads invoice 881 throughInvoiceRepository. - The domain object's
issue()method enforces the rules (not already issued, at least one line, totals recomputed) and records an "issued" state. - The handler saves through the repository, commits, and maps the result to an
InvoiceResponseDTO. - Persistence writes the rows, with the query scoped by
tenant_id, and Postgres row-level security (RLS: a database policy that hides rows belonging to other tenants) as a second guard.
If the tax rule changes, one domain class changes. If the database moves from Postgres to another store, only persistence adapters change. If the mobile app needs a new response shape, only the API DTO mapping changes.
How I enforce separation of concerns
For code boundaries
- One module (package, project or workspace) per layer, with the build graph forbidding inward-to-outward imports.
- An architecture test in CI that fails the build on a forbidden dependency: ArchUnit's
layeredArchitecture()rules in Java, import-linter "layers" contracts in Python, dependency-cruiser rules in TypeScript. This is the single most effective mechanism because it turns a design principle into a red build. - API DTOs are separate types from domain objects and from ORM entities, so a column rename cannot silently change the public response.
For testing
- Domain: fast unit tests with no mocks and no I/O.
- Application: tests using in-memory fakes of the repository interfaces.
- Persistence: integration tests against a real Postgres in a container, because a fake cannot catch a wrong SQL query.
- Frontend-to-API: contract tests against the OpenAPI document, so either side learns immediately when the other breaks the contract.
For deployment
- Frontend ships to the CDN (content delivery network) on its own schedule.
- Gateway configuration (routes, rate limits) is versioned and deployed separately.
- Application, domain and persistence ship as one deployable. Splitting them into separate network services would turn each in-process call into a network hop, which is how layered designs become chatty distributed monoliths.
For security boundaries
- Authentication at the gateway, authorization in the application layer (where the use case and the resource are both known), tenant isolation enforced again in the database.
- Only the persistence layer holds database credentials, with a least-privilege role (no DDL rights at runtime: DDL, or data-definition-language statements such as CREATE TABLE or ALTER TABLE, change the schema itself, as opposed to the DML statements that read and write rows; migrations run under a separate role that does hold DDL rights).
For teams
- Code ownership files map layers to owning teams, so a change to domain rules requires a domain owner's review, and platform engineers own the gateway configuration.
Trade-offs and pitfalls
- Anemic domain: all rules end up in application handlers and the "domain" is getters and setters. Watch for handlers that contain
ifstatements about business policy. - Layer skipping: a controller (an application-layer handler, the thing that receives the incoming call in step 3 of the worked example above) calling the repository directly for "just a quick read". Allow it only for explicitly read-only query paths, and make that an explicit rule rather than an accident.
- God gateway: teams start adding per-endpoint transformations and business checks in gateway plugins because it is the easiest place to deploy. The gateway should stay boring.
- Ceremony cost: three mapping layers for a simple CRUD (create, read, update, delete) screen is overhead. For a small internal admin screen it is fine to let the application layer map directly from the repository to a DTO, as long as the ORM entity never leaves the persistence module.
Design an extensible plugin architecture for a platform where customers or internal teams can add custom business logic (for example, custom validation rules or scoring logic) without modifying the core service. Cover how plugins are structured and versioned, how you'd validate and sandbox them for safety, whether they can be hot-reloaded, how you'd enforce backward compatibility across plugin versions, and how you'd test a new plugin without risking the core service.
Sample Answer
Direct answer
I would build a microkernel (plugin) architecture: a small core that owns the data, the workflow and every side effect, plus a fixed set of named extension points (hooks such as validate_order or score_lead) where plugins are allowed to run. Plugins are pure functions over a versioned input/output contract, compiled to WebAssembly (Wasm, a portable, low-level bytecode format that runs inside its own sandboxed virtual machine with its own private block of memory, so plugin code cannot read or write anything outside that block unless the host explicitly allows it) and executed in a sandbox with hard limits on CPU, memory and host access. The core never links plugin code into its own process memory (never loads it as a native library or script sharing the core's own memory space, the way an in-process plugin would); each plugin instead gets its own separate memory and only a narrow, explicit channel back to the core. It loads plugins from a signed, versioned registry, can swap them at runtime without a restart, and only promotes a new plugin version after it has passed a conformance suite and a shadow run (running the new version on a copy of real traffic, off to the side of the request that is actually served, and comparing its output against the current version without letting it affect anything yet) against real traffic.
The key design stance: plugins compute, the core acts. A plugin returns a decision ("reject with reason X", "score = 72"); it never writes to the database or calls the network itself. That single rule is what makes safety, versioning and testing tractable.
Requirements I am designing to
- Tenants (customers) and internal teams add validation rules and scoring logic; the core team must not need to redeploy for that.
- Hot path: validation runs inline on each order submission. Budget assumption: the core's own p99 (99th percentile) latency target is 200 ms, and I give all plugins on one request a combined budget of 20 ms.
- A buggy or malicious plugin must not crash the core, read another tenant's data, or exfiltrate data over the network (secretly send it out to somewhere the plugin's author controls).
- Hundreds of tenants, each with a handful of plugins, so plugin count is in the low thousands, not millions.
Architecture
flowchart LR
Dev[Plugin author] --> CI[Plugin pipeline: build, sign, conformance tests]
CI --> Reg[(Plugin registry: artifact + manifest + version)]
Reg --> Loader[Core: plugin loader]
Req[Order request] --> Core[Core service]
Core --> Host[Plugin host: Wasm sandbox]
Loader --> Host
Host --> Core
Core --> DB[(Core database)]
1. How a plugin is structured
Each plugin is an artifact plus a manifest:
| Manifest field | Example | Why it exists |
|---|---|---|
name, version | acme-credit-check, 2.3.1 | Semantic versioning (MAJOR.MINOR.PATCH) for the plugin's own releases |
api_version | validation/v2 | Which version of the core's extension contract it targets |
hooks | ["validate_order"] | Which extension points it implements |
capabilities | ["read:order", "config:thresholds"] | Everything it asks the host for; nothing else is granted |
limits | timeout_ms: 5, memory_mb: 32 | Requested ceilings, capped by platform maximums |
owner, signature | tenant id, signing key id | Accountability and tamper detection |
The contract is a small SPI (service provider interface: the interface the core publishes for plugins to implement), defined in an interface definition language rather than in one programming language. For Wasm, that is a WIT file (WebAssembly Interface Types, the interface language of the Wasm Component Model: a standard way for Wasm modules to describe and exchange typed interfaces, so a module written in one language can be called correctly from a host, or from another module, written in a different one), which lets authors write plugins in Rust, Go, JavaScript or Python and still meet one contract. Input is a read-only snapshot the core prepares (the order, the tenant's config), output is a typed decision.
2. Versioning
Two version numbers matter, and conflating them is the common mistake:
- Contract version (
validation/v1,validation/v2) is owned by the core. Within a major version, only additive changes are allowed: new optional input fields, new optional output fields. Removing or retyping a field means a new major version. - Plugin version (
2.3.1) is owned by the author. The registry is immutable:2.3.1can never be overwritten, only superseded, so every decision in the audit log can be traced to exact bytes.
Each tenant pins a plugin version per hook (acme-credit-check@2.3.1). Upgrading is an explicit promotion, never "latest".
3. Validation and sandboxing
Validation happens before a plugin can run anywhere near production:
- Signature check: the artifact is signed by the pipeline; the loader refuses unsigned or re-signed bytes.
- Static checks: the Wasm module only imports functions the host actually offers for its declared capabilities. An import the manifest did not declare fails the upload.
- Conformance suite: the core team publishes contract tests per hook: golden inputs (a fixed set of sample inputs with known-correct expected outputs, which every plugin version is checked against) and required behaviours such as never throws on an empty cart, always returns a reason with a rejection, handles unknown optional fields. Every plugin version must pass them in CI.
At runtime, the sandbox enforces what the manifest promised:
- Memory: each instance gets a fixed linear-memory cap (32 MB here): linear memory is Wasm's name for a module's own private, contiguous block of memory, the one described above, and a Wasm module can only ever address bytes inside its own block. So it cannot read the core's heap (the pool of memory the core process itself uses to hold its own data) even by accident.
- CPU: the runtime meters execution. Wasmtime, for example, offers fuel consumption (a deterministic instruction budget) and epoch interruption (a periodic deadline check); either stops an infinite loop at the time budget.
- Host access: no filesystem, network or clock unless the manifest's capabilities grant a specific host function. Tenant isolation comes from the core only ever passing that tenant's own snapshot in.
- Failure policy per hook, decided up front: a validation hook that times out or traps fails closed (the order is held for review, because silently accepting an invalid order is worse), a scoring hook fails to a default (a neutral score) and emits an alert. A plugin that fails more than a threshold (say 5% of calls in 5 minutes) is auto-disabled for that tenant and its owner is paged.
4. Hot reload
Yes, and it is cheap because plugins are stateless and sandboxed:
- The loader watches the registry for promotion events.
- It fetches and compiles the new version in the background.
- It atomically swaps the tenant's pointer from the old compiled module to the new one. Requests already running keep the old instance; new requests get the new one.
- The old module is dropped once its in-flight count reaches zero.
Rollback is the same swap in reverse, so it takes seconds and needs no deploy. What I would not hot-reload is the contract itself: a new api_version ships with a core release.
5. Backward compatibility across plugin versions
- The core supports the current and previous contract major version (
v2andv1) at the same time, using an adapter that translates av2input down tov1shape for old plugins. - Every contract change runs the entire registry's current pinned plugins against the new core in CI. If a change to the core breaks any pinned plugin, the core change fails, not the plugin.
- Deprecation has a published window (for example two quarters), a dashboard of which tenants are still on
v1, and direct outreach before removal.
6. Testing a new plugin without risking the core
Progressive exposure, each stage automated:
| Stage | What runs | Risk to production |
|---|---|---|
| Local SDK | Author runs the plugin against a local host binary with sample inputs | None |
| Conformance CI | Contract tests plus fuzzed inputs (automatically generated random or malformed inputs meant to find edge cases a human author would not think to write by hand) | None |
| Replay | New version runs over a recorded day of that tenant's (anonymised) inputs; decisions are diffed against the current version | None |
| Shadow | New version runs on a copy of live inputs alongside the pinned version, off the request path (after the response is sent); its output is logged but ignored | Extra compute only; no user-visible latency |
| Canary (a canary release: rolling the new version out to a small, real slice of traffic before trusting it with all of it) | New version's decisions take effect for 5% of that tenant's requests | Limited to one tenant, one slice |
| Full | Promoted | Rollback is one pointer swap |
Worked example: does the latency budget hold?
Assume a tenant has 3 validation plugins on validate_order, run sequentially because later rules may depend on earlier outcomes, with a per-plugin timeout of 5 ms.
- Worst case with every plugin hitting its timeout: 3 × 5 ms = 15 ms, inside the 20 ms plugin budget.
- A 4th plugin lands exactly on the line: 4 × 5 ms = 20 ms, equal to the budget, not over it. I treat the budget as inclusive (a request may use up to and including 20 ms of plugin time), so a 4th plugin is technically allowed, but it leaves zero margin, so in practice I would cap tenants at 3 sequential plugins and require anything beyond that to run in parallel instead.
- A 5th plugin makes that explicit: 5 × 5 ms = 25 ms, which exceeds 20 ms even before any margin. The platform either rejects the configuration or requires the tenant to mark independent rules (ones that do not depend on each other's output) as parallel, in which case the worst case is max(5 ms) plus scheduling overhead, the extra time to dispatch several plugin calls at once and wait for the slowest one to finish, typically well under a millisecond, rather than the sum of every plugin's timeout.
That calculation is why the manifest carries limits and why the platform, not the author, caps them: the budget is a property of the whole request, not of one plugin.
Trade-offs and pitfalls
- Letting plugins perform side effects. The moment a plugin can write to the database or call an external API, you need transactions across plugin code, retries, and data-access policy per plugin. Keep plugins pure and have the core act on their decisions.
- In-process plugins "for speed". Loading tenant code as a native library or script inside the core process means one bad plugin can crash or read everything. Worth it only for internal code written and reviewed by the core team.
- An extension point per request. Every hook is a contract you must keep backward compatible forever. Start with a few well-chosen hooks and add more when there is real demand.
- "Latest" pinning. Auto-upgrading plugins turns every author release into an unreviewed production change for every tenant.
- Plugins that need heavy compute or long-running work (calling an ML model, minutes of processing) do not fit the inline Wasm model. Those belong out of process, behind an asynchronous job interface, which is a different extension point with different latency expectations.
What would change the design: if plugins were only ever written by the core company's own teams, I would drop Wasm and load them in-process behind the same SPI, trading isolation for easier debugging.
Unlock Full Question Bank
Get access to all 20 Architectural Patterns and Anti-Patterns interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.