Architectural Patterns and Anti-Patterns Questions
Architecture-level patterns and the anti-patterns that signal a wrong turn. Patterns: layered and n-tier architecture, including where cross-cutting concerns like authentication, rate-limiting and tracing belong, dependency injection trade-offs, and thin-versus-fat controller design; hexagonal (ports and adapters) and clean architecture; CQRS and event sourcing; backend-for-frontend; plugin (microkernel) extension models; and the coupling, cohesion, encapsulation and separation-of-concerns principles behind them, including when each applies and what it costs. Anti-patterns: distributed monolith, chatty services, shared-database coupling, cyclic service dependencies, leaky abstractions that expose internal schemas, and golden-hammer pattern adoption. Covers the detection signals (deploy coupling, call-graph fan-out, change amplification, trace evidence), incremental remediation, and architecture governance that keeps smells from recurring. This is about diagnosing and fixing the smell in an existing design, not the monolith-versus-microservices decision itself.
Design patterns can become anti-patterns when they're reached for out of habit rather than fit. Describe a realistic case where adopting a popular pattern (at the service or architecture level) created more complexity or risk than it solved. What made it the wrong fit for that context, and what would have signaled the mismatch earlier?
Sample Answer
Direct answer
A realistic case is a small product team that adopted backend-for-frontend (BFF), one dedicated API service per client type, because it was "the recommended pattern when you have multiple frontends". BFF solves a real problem: very different clients, owned by separate teams, that need very different data shapes. This team had none of those forces. Its web and mobile apps showed nearly the same screens and were built by the same engineers, so the pattern mostly bought triplicated code, drift between clients and extra operational load. The mismatch was visible early in measurable signals: nearly identical endpoints across the BFFs, business rules leaking into them, and every feature needing several pull requests.
This is the golden-hammer anti-pattern: choosing a familiar or fashionable tool before naming the problem it is meant to solve.
The case (story skeleton)
Context. A seven-engineer team builds a B2B (business-to-business) ordering product: a web app for office staff and iOS and Android apps for field sales. One core API service owns orders, pricing and customers.
Decision. Early on, the team adds three BFF services (web, iOS, Android), each aggregating calls to the core API and shaping responses for its client. The design doc cites the pattern and large companies using it; it does not state which client need the core API fails to meet.
What happened over the following months.
- Change amplification: a new field on the order screen needed a change in the core API plus the same change in three BFFs: four pull requests, four deploys, four sets of tests for one user story.
- Logic drift: "just for now", a discount-eligibility check was added in the web BFF to unblock a release. The mobile BFFs never got it, so mobile users briefly saw prices that the web app correctly refused. The bug came from business logic living in an aggregation layer, which a BFF is not supposed to hold.
- Operational load: three extra services to patch, scale, monitor and put on call, for a team with no dedicated platform engineer.
- Latency: every call now crossed an extra network hop for responses the core API could have served directly.
Remediation. The team collapsed the three BFFs into the core API, which gained optional field selection (clients request the fields they need) and a couple of screen-oriented composite endpoints. The discount rule moved into the pricing domain code where every client gets it. A single BFF was kept later, but for a genuinely different consumer: a partner integration with its own authentication, rate limits and response format.
Why it was the wrong fit
Every pattern resolves specific forces. Checking them against this context:
| Force BFF resolves | Present here? |
|---|---|
| Clients need substantially different data shapes or call patterns | No: the same screens on every client |
| Separate teams own separate clients and need to ship independently | No: one team built everything |
| Client-specific concerns (device-specific auth flows, payload limits, protocol differences) | Minimal |
| A general-purpose API that is hard to change | No: the same team owned the core API |
With none of the forces present, the pattern contributed only its costs: more deployables, more duplicated code and a new place for business logic to hide. The value of a pattern is always relative to the problem; out of context, what remains is its overhead.
Signals that would have exposed the mismatch earlier
- The design doc named the pattern before naming the problem. A good architecture decision record (ADR, a short document capturing context, decision and consequences) states the forces first. "Because Netflix does it" is not a force.
- Endpoint similarity. Diffing the three BFFs' routes and response shapes would have shown them to be nearly identical. If most endpoints are copies, the per-client split has no purpose.
- Change amplification per story. Tracking how many repositories or services a typical feature touches: one story needing four pull requests is a structural smell, not a process one.
- Business logic outside the domain. Any
ifabout pricing, eligibility or permissions in an aggregation layer is a warning; a lint rule or review checklist can flag it. - Ownership mismatch. In Sam Newman's original framing, a BFF is ideally owned by the team that owns the client. One team owning all of them is a hint the split is not buying autonomy.
- Operational cost against team size. Services per engineer rising without a matching rise in independent release cadence.
A lightweight guard against golden hammers
Before adopting any architecture-level pattern, answer four questions in writing:
- What specific force or pain does it resolve, and is that pain happening now?
- What does it cost to build, run and staff?
- What is the simplest alternative, and why is it insufficient?
- How hard is it to reverse, and what signal would tell us to reverse it?
If question 1 has no concrete answer, defer the pattern. Adding a BFF later, when a genuinely divergent client appears, is cheap; removing three that have accumulated business logic is not.
Trade-offs and pitfalls
- The overcorrection: refusing all patterns leads to the opposite smell, one general API contorted to serve incompatible clients. The point is fit, not avoidance.
- Fixing the symptom: moving the discount rule into all three BFFs "for consistency" would have hidden the drift while cementing the duplication.
- Sunk cost: teams keep an ill-fitting pattern because removing it feels like admitting a mistake. Treat reversals as a normal outcome of an ADR with an explicit reversal signal.
Explain dependency injection (DI) and the differences between constructor injection, setter injection, and the service-locator approach. When designing a layered backend, what are the advantages and the potential pitfalls of leaning on a DI framework in your service and repository layers?
Sample Answer
Direct answer
Dependency injection (DI) means an object receives the collaborators it needs (a repository, a clock, an HTTP client) from outside, instead of creating or looking them up itself. Constructor injection passes required dependencies as constructor arguments, so the object is complete the moment it exists. Setter injection assigns dependencies after construction through setter methods, which suits optional ones. A service locator is the opposite direction: the object asks a global registry for what it needs at run time, which hides its dependencies. Use constructor injection by default, setters only for genuinely optional dependencies, and keep service locators out of business code.
The three approaches compared
| Constructor injection | Setter injection | Service locator | |
|---|---|---|---|
| How dependencies arrive | Constructor parameters | Setter methods after construction | Object calls Registry.get("x") itself |
| Are dependencies visible? | Yes, in the signature | Partly (must read the setters) | No, buried in method bodies |
| Can the object exist half-built? | No | Yes, until every setter runs | Looks complete, fails later |
| Immutability | Fields can be final or read-only | Fields must stay mutable | N/A |
| When a dependency is missing | Fails at construction (or at application startup with a container) | Null reference at first use | Lookup error at first call, possibly in production |
| Good for | Required collaborators | Optional collaborators with a sensible default | Framework glue at the edge of the application |
Some frameworks also allow field injection (the container writes directly into a private field). It hides dependencies like a locator does and makes plain unit testing awkward; Spring's own reference documentation recommends constructor injection for mandatory dependencies and setters for optional ones.
Worked example: why hidden dependencies hurt
This runnable Python program builds the same service two ways.
class Registry: # a service locator: a global lookup table
_services = {}
@classmethod
def register(cls, name, obj):
cls._services[name] = obj
@classmethod
def get(cls, name):
return cls._services[name]
class OrderServiceWithLocator:
def total(self, order_id):
repo = Registry.get("order_repo") # hidden dependency
return sum(repo.line_amounts(order_id))
class OrderServiceWithConstructor:
def __init__(self, repo): # dependency is visible in the signature
self.repo = repo
def total(self, order_id):
return sum(self.repo.line_amounts(order_id))
class FakeOrderRepo:
def line_amounts(self, order_id):
return [1200, 350]
print("constructor:", OrderServiceWithConstructor(FakeOrderRepo()).total("o1"))
svc = OrderServiceWithLocator() # constructs fine, nothing looks wrong yet
try:
print("locator:", svc.total("o1"))
except KeyError as e:
print("locator failed at call time, missing:", e)
Registry.register("order_repo", FakeOrderRepo())
print("locator after global setup:", svc.total("o1"))
Output:
constructor: 1550
locator failed at call time, missing: 'order_repo'
locator after global setup: 1550
The constructor version tells you from its signature that it needs a repository, and a test simply passes a fake (1,200 + 350 = 1,550). The locator version constructs without complaint and fails only when total runs; to test it you must mutate global state, which leaks between tests that run in the same process.
Advantages of a DI framework in the service and repository layers
- Wiring large object graphs. A container (Spring, .NET's built-in container, Guice, Dagger) builds hundreds of objects in the right order so you do not hand-write the assembly code.
- Lifecycles and scopes. It manages which objects are shared for the whole application (singleton), created per request, or created each time.
- Environment-specific configuration. Swap the real payment client for a sandbox client in staging through configuration, not code changes.
- Declarative cross-cutting behaviour. Transactions or metrics applied by annotation (a marker attached to code, such as
@Transactional, that tells the framework to wrap it with extra behaviour automatically) around service methods. - Testability. Because services receive their repositories, tests hand them fakes, and the business logic never needs a database to be exercised.
Pitfalls to watch for
- Framework leaking into the core. Annotations and container types in domain classes couple your business rules to the framework; hexagonal and clean architecture (architectural styles that keep business rules as plain code with no framework dependencies, pushing frameworks, databases and other details to the outside) both require that dependencies point inward, toward framework-free code. Keep the core as plain objects and do the wiring in one composition root (the single place at startup where the object graph is assembled).
- Scope mismatch. A singleton service that receives a request-scoped object (one the container creates fresh for each incoming request and discards afterward) at startup keeps that one instance forever, so data from one request can bleed into another. Containers offer proxies (a lightweight stand-in injected into the singleton that, on every call, fetches whichever request-scoped instance is current) or providers (a small factory object injected instead of the real dependency, which the singleton calls each time it needs a fresh, current instance) for this; you have to know to use them.
- Circular dependencies. With constructor injection, a cycle (A needs B, B needs A) cannot be constructed and the container fails at startup, which is good: it exposes a design problem. Setter or field injection can let the container paper over the cycle.
- Hidden god classes (classes that quietly accumulate far more responsibility than they should, because adding one more dependency is so easy). Adding a dependency costs one more constructor parameter, so bloat creeps in silently. A service with 10 or more injected collaborators is a cohesion smell (a warning sign that a class is doing too many unrelated things to be well designed); split it.
- Magic and startup cost. Classpath scanning (the container inspecting every class in the application at startup to find the ones it should manage) and reflection (code that inspects or calls into other code at run time, by name, rather than at compile time) make it hard to answer "where does this object come from?" and slow startup, which matters for serverless (functions that start fresh for each invocation and shut down between them, so every millisecond of startup cost is paid repeatedly) or short-lived processes. Compile-time DI (Dagger) (wiring code generated when the project is built, instead of discovered by scanning and reflection at startup) or plain manual wiring avoids both.
- Interfaces for everything. Creating an interface for every class "for DI" adds ceremony. Introduce an interface where there is a real second implementation or a test double worth having.
Recommendation: constructor injection everywhere in service and repository code; a container only at the outer layer; plain manual wiring for small services, where a 20-line composition root is clearer than a framework.
A mature system has become 'chatty', with many synchronous service calls driving up latency. Propose a refactor plan to reduce the chattiness: the steps you'd take, how you'd stay backward-compatible along the way, the trade-offs of moving some calls to an event-driven or batched approach, how you'd preserve data correctness, and what KPIs would show it's working.
Sample Answer
Direct answer
A chatty system makes many small synchronous calls between services to serve one user request, so latency and failure probability add up across every hop. My plan: first measure the call graph per endpoint with distributed tracing, then apply the cheapest fixes in order (parallelise independent calls, add batch endpoints, add coarser-grained endpoints: fewer endpoints that each return several related things together instead of one endpoint per small piece), and move only read-heavy reference data to an event-driven local copy. Every change ships behind the old interface using an expand-then-contract sequence so no caller breaks. Data that must be correct at the moment of a decision (price charged, stock reserved) stays on a synchronous, authoritative check. KPIs are calls per request, end-to-end p99 latency (the 99th-percentile response time: how slow the worst 1% of requests are, not the average), replica staleness and reconciliation drift (how far a copied dataset has drifted from its source; see below).
Why chattiness hurts: worked example
A product-listing endpoint shows 20 items. Today it calls the catalogue service once for the list, then pricing once per item and inventory once per item, sequentially: 1 + 20 + 20 = 41 calls. Assume (as an illustration, not a measurement) about 10 ms per call including network: 41 x 10 = 410 ms of mostly waiting.
Latency is not the only problem. If each call independently has a 1% chance of being slow (say over 100 ms), the chance that a request hits at least one slow call is:
P(at least one slow)=1−0.99n n=41:1−0.9941≈0.338n=3:1−0.993≈0.030So roughly one request in three sees a slow dependency with 41 calls, versus about one in 33 with 3 calls. Chattiness multiplies tail latency (how slow the worst-affected requests are, as opposed to the average; p99, the 99th percentile, is the usual way to measure it) far more than it multiplies average latency.
After the refactor below: one catalogue call, then one batch pricing call and one batch inventory call in parallel. Assuming 10 ms for the catalogue call and 15 ms for each batch call: 10 + max(15, 15) = 25 ms, and 3 calls instead of 41.
Refactor plan
flowchart LR
A[1. Trace and rank endpoints by calls per request] --> B[2. Parallelise independent calls]
B --> C[3. Add batch endpoints]
C --> D[4. Coarse-grained or composite endpoints]
D --> E[5. Event-fed local read model for reference data]
E --> F[6. Revisit boundaries that remain chatty]
- Measure. Use distributed traces to compute, per endpoint, the number of downstream calls, the call depth, and which calls are repeated in a loop (the network version of the N+1 query problem: one call to fetch a list, then one more call per item in that list, when a single batched call would do). Rank endpoints by traffic x calls per request, and fix the top few first.
- Parallelise. Calls that do not depend on each other run concurrently. No contract change needed; often the largest early win for latency, though it does not reduce load or failure exposure.
- Batch. Add
POST /prices:batchGettaking up to, say, 100 IDs. Cap the batch size so one huge request cannot become its own tail-latency problem. Deduplicate IDs within a request. - Coarsen. If a caller always needs order + lines + shipment together, offer one endpoint that returns them, owned by the service that owns the data.
- Local read model for reference data. For data read far more often than it changes (product names, categories, tax classes), the owning service publishes change events and the consumer keeps its own read-only copy. Calls disappear entirely, at the price of eventual consistency (the copy lags the source by some seconds).
- Question the boundary. If two services still exchange dozens of calls per request after all this, they probably share one responsibility. That is a signal to raise, even if redrawing boundaries is a separate project.
Staying backward-compatible
- Expand, then contract. Add the batch endpoint next to the per-item one; migrate callers one at a time; remove the old endpoint only when its traffic has been zero for an agreed period.
- Consumer-driven contract tests so the provider knows exactly which callers depend on which fields.
- Feature flags per caller, so a new path can be enabled for 1% of traffic and rolled back instantly.
- Shadow comparison: for a period, compute the response both ways and log differences without serving the new one. The diff rate is the go/no-go signal.
- Additive-only event schemas: new fields optional, never repurpose a field.
Trade-offs: sync per item vs batched vs event-driven
| Approach | Freshness | Coupling | Failure mode | Complexity |
|---|---|---|---|---|
| Synchronous per item | Always current | High: caller fails when callee fails | Tail latency and cascading failures (one slow or failing call causing failures to spread to the services that called it) | Low |
| Synchronous batch | Always current | Still runtime-coupled, but 1 call | One slow batch delays the whole response | Low to medium |
| Event-driven local copy | Seconds stale | Runtime-decoupled; coupled to the event schema | Silent staleness or drift if consumption stalls | High: consumers, replays, schema versioning |
My default is batch first; I move to events only for reference data where seconds of staleness are acceptable and read volume is high. Event-driven is not a free latency fix: it trades a visible, loud failure (a timeout) for a quiet one (stale data).
Preserving data correctness
- Classify every field by how stale it may be. Product name: minutes is fine. Displayed price: seconds is fine. Charged price and stock reservation: must be authoritative, so checkout re-validates synchronously against the owning service even if the page used the local copy.
- Atomic publish. The owner writes the change and the event in one database transaction (the transactional outbox: the event is stored in an outbox table in the same transaction as the data change, then a separate relay process reads that table and publishes the event to the broker afterwards), so an update is never saved without its event or vice versa.
- Idempotent, ordered application. Consumers record the last applied version per entity and ignore duplicates or older versions, so redelivery or out-of-order delivery cannot roll data back.
- Reconciliation. A periodic job compares a sample (or all) of the local copy against the source and reports a drift count; drift above zero is investigated, not ignored.
- Batch semantics. Define partial results explicitly: a batch call returning prices for 98 of 100 IDs must say which 2 are missing, not silently omit them.
KPIs that show it is working
- Downstream calls per request (p50, the typical/median request, and p99, the value only the worst 1% of requests exceed) per endpoint: the direct measure of chattiness. Target for the example: 41 down to 3.
- End-to-end p99 latency of the refactored endpoints, and the share of requests touching a slow dependency.
- Error rate of those endpoints and the number of distinct dependencies each request touches.
- Replica staleness: event consumer lag (p99, in seconds) against the agreed freshness target.
- Reconciliation drift count, which should stay at zero.
- Traffic remaining on deprecated per-item endpoints, trending to zero.
- Load on the downstream services (requests per second), which should fall as batching lands.
Pitfalls
- Jumping to events for everything and discovering that "eventually" means minutes during a consumer backlog.
- Batch endpoints with no size cap, which move the latency problem inside one call.
- Removing the old endpoint based on "we think everyone migrated" instead of measured traffic.
In a layered backend, concerns like authentication, rate-limiting, and tracing apply to nearly every request. If each service implements them independently you get drift and duplicated bugs. Where architecturally should this kind of cross-cutting logic live, what are the trade-offs of the placement options, and how do you keep the behavior consistent as new services are added?
Sample Answer
Direct answer
Put each cross-cutting concern at the outermost place that has enough information to do it correctly, and make that place shared infrastructure rather than per-service code. Concretely: verify end-user identity and enforce per-client rate limits once at the edge (an API gateway), secure service-to-service traffic in the platform (a service mesh: a dedicated network layer, usually a fleet of per-service proxies plus a shared control plane, that handles this uniformly, or equivalent), and keep only the parts that need business context inside each service, delivered through one shared, versioned middleware library. Consistency then comes from a paved-road service template plus automated conformance checks, not from asking every team to remember.
Cross-cutting concern means behaviour that applies to nearly every request regardless of what the request does (identity, throttling, tracing, logging). Copying it into every service is how you get the drift the question describes.
The placement options
| Option | What it is | Strength | Cost |
|---|---|---|---|
| Copy per service | Each team writes its own | Full local control | Drift, duplicated bugs, N places to patch a security fix |
| Shared library / middleware | One package every service imports; runs in-process as request middleware | Sees business context; cheap at runtime | Tied to a language; upgrades roll out service by service, so versions skew |
| API gateway | A single entry point in front of all services that inspects every external request | One deploy fixes all services; blocks bad traffic before it costs anything | Extra hop; a shared choke point; tempts teams to put business logic in it |
| Sidecar / service mesh | A proxy deployed next to each service instance (a sidecar) that handles its network traffic; a service mesh is the fleet of these plus a control plane (the piece that configures every sidecar with routing rules, security policy and certificates, so operators manage one thing instead of each proxy individually); Istio and Linkerd are example service-mesh products that bundle both | Language-agnostic service-to-service security and telemetry | Operational complexity, per-pod (per Kubernetes pod, the group of containers deployed and scaled together) CPU and memory overhead, a new failure point |
| Managed platform service | Identity provider, managed secrets, managed tracing backend | Someone else runs the hard part | Vendor coupling, less control |
Where each of the three concerns should live
Authentication (who is calling).
- Edge: the gateway validates the end-user token (a JWT, JSON Web Token: a signed token carrying the user's identity and expiry) once, rejects bad ones, and forwards the verified identity.
- Between services: mTLS (mutual TLS, where both sides present certificates so each proves its identity) issued by the mesh, so a service knows which service is calling it.
- Inside the service: a thin shared middleware still checks that the forwarded identity is present and meant for this service. This follows zero trust: never trust a request merely because it arrived from inside the network.
- Authorization ("may this user read this invoice?") stays in the service, because only the service knows who owns the invoice. Pushing that into the gateway is how gateways grow into god services (services that have absorbed so many unrelated responsibilities that no one can change them safely).
Rate limiting.
- Per-client quotas ("each API key gets 600 requests per minute") belong at the gateway, backed by a shared counter.
- Self-protection limits ("this service accepts at most 200 concurrent requests") belong in the service or its sidecar, because they depend on that service's own capacity.
Tracing (recording each request's path across services as a tree of timed spans, where each span is one unit of work, such as one service handling one call, with its own start time and duration).
- It must be partly in-process. A sidecar can time each network hop, but it cannot see work inside the process, and it cannot connect an inbound request to the outbound calls it causes unless the application forwards the trace context header (the W3C
traceparentheader, which carries the trace ID and current span ID so the next hop can attach itself to the same trace). So the shared library must include OpenTelemetry (the vendor-neutral standard and SDK for traces, metrics and logs) instrumentation that propagates context automatically.
flowchart LR
C[Client] --> G[API gateway: verify JWT, per-client quota]
G --> P1[Sidecar: mTLS]
P1 --> S1[Service A: shared middleware, authz, tracing SDK]
S1 --> P2[Sidecar: mTLS]
P2 --> S2[Service B: shared middleware, authz, tracing SDK]
Worked example: two numbers that decide placement
Distributed rate limits. A 600 requests-per-minute limit per API key is configured on 3 gateway replicas, each counting locally. A client whose traffic is spread round-robin across replicas can send up to 3 x 600 = 1,800 requests per minute before any replica says no. Fixes: keep the counter in a shared store such as Redis (a fast, shared in-memory data store), or give each replica 600 / 3 = 200 and accept that uneven routing will throttle some clients early. Placing the limit "in the gateway" is not enough; where the counter lives is part of the design.
Patch latency. Say 40 services depend on an auth library and a token-validation bug is found. If teams upgrade at about 5 services a week (an assumption; measure your own), full coverage takes 40 / 5 = 8 weeks, during which some services are still vulnerable. The same fix in the gateway is one deploy. That asymmetry is why token validation belongs at the edge and the library only double-checks.
Keeping it consistent as services are added
- Paved road (a default, pre-wired path that is easier to follow than to deviate from). New services start from a template that already wires in the middleware, the tracing SDK and the mesh annotations (configuration labels attached to a service that tell the mesh's control plane how to handle its traffic). The easy path is the correct path.
- Minimum-version policy. CI (continuous integration) fails a build whose shared-library version is older than the supported floor, which turns skew into a visible, bounded window.
- Policy as code. Gateway routes, rate-limit rules and mesh authorization policies live in one reviewed repository, not in consoles.
- Conformance tests. An automated check hits every newly registered service without a token and expects 401, and checks that its spans appear in the trace backend. This is an architecture fitness function: an automated test of an architectural property.
- Clear ownership. A platform team owns the gateway, mesh and library; product teams own authorization and business rules.
Trade-offs and pitfalls
- Gateway creep. Once authentication lives in the gateway, teams add request transformation, then business rules, and the gateway becomes a god service every team must change. Keep it domain-agnostic.
- Edge-only security. If services trust any request that reached them, one compromised internal service can impersonate everyone. Keep mTLS and in-service identity checks.
- Mesh adoption for one concern. Adopting a service mesh only to get tracing is golden-hammer thinking (reaching for one favourite tool for every problem, regardless of fit); the SDK does most of that. Adopt a mesh when you need uniform mTLS and traffic policy (rules for how requests are routed, retried and load-balanced) across many languages.
- What would flip the recommendation: a small system with 3 services in one language does not need a gateway or mesh; a single shared middleware library plus a service template is enough.
Explain the Backend-for-Frontend (BFF) pattern: how it tailors APIs per client type (mobile vs. web), reduces over-fetching, and hides internal service complexity. What deployment, testing, and team-ownership trade-offs would you weigh before adopting BFFs?
Sample Answer
Direct answer
A Backend-for-Frontend (BFF) is a small server-side API built for one specific client experience, such as the mobile app or the web app, and usually owned by the team that builds that client. Instead of the client calling many internal services and trimming the results itself, it makes one call to its BFF, which fans out to the internal services, combines the results and returns exactly the shape that screen needs. That cuts round trips and over-fetching on slow networks and hides the internal service layout from clients. The price is an extra deployable per client, extra hops, duplicated logic risk and the need for clear ownership. Adopt it when client experiences genuinely differ and each has its own team.
How it works
flowchart LR
M[Mobile app] --> MB[Mobile BFF]
W[Web app] --> WB[Web BFF]
MB --> O[Orders service]
MB --> C[Catalog service]
WB --> O
WB --> C
WB --> R[Reviews service]
- Tailoring per client. The mobile BFF returns a compact list with a thumbnail URL and price; the web BFF returns richer detail with reviews. Each BFF can also handle the client's own concerns, such as cookie sessions for web and token refresh for mobile.
- Reducing over-fetching. Over-fetching means receiving far more data than the screen uses. The BFF drops unused fields server-side, on the fast internal network, instead of over the user's mobile connection.
- Aggregation. The BFF calls several services in parallel inside the data centre and merges the results, so the client pays one network round trip instead of several.
- Hiding internal complexity. Clients see one stable API. Internal services can be split, merged or renamed, and only the BFF changes.
Worked example: one product-list screen
Assumptions (illustrative, not measured): a mobile round trip of 200 ms on a cellular network; the screen needs data from 4 services; each internal call takes 25 to 45 ms; 20 products per page; the full product object is 12 KB, while the mobile screen uses about 1.5 KB of it.
| Client calls services directly | Through a mobile BFF | |
|---|---|---|
| Round trips over the mobile network | 4 sequential calls: 4 x 200 = 800 ms | 1 call: 200 ms |
| Internal calls | none | 4 in parallel, bounded by the slowest: about 45 ms |
| Network time, approximately | 800 ms | 200 + 45 = 245 ms |
| Payload to the device | 20 x 12 KB = 240 KB | 20 x 1.5 KB = 30 KB |
The payload shrinks by a factor of 240 / 30 = 8, and the waiting time drops by more than half, which is the whole case for the pattern on mobile. If the client could already make those calls in parallel, the direct figure would be closer to 200 ms plus the slowest call, and the latency argument weakens; the payload and coupling arguments remain.
Trade-offs to weigh before adopting
Deployment
- One more service per client experience to build, monitor, scale and keep available. If the BFF is down, that client is down.
- An extra hop on every request, small inside a data centre but not zero.
- Old mobile clients never fully disappear. Users do not all update, so the mobile BFF must serve several app versions at once. Plan versioned endpoints and a support window from day one.
Testing
- Two contracts per BFF: client to BFF, and BFF to each downstream service. Consumer-driven contract tests (the BFF publishes what it relies on; each downstream service verifies it in CI) keep a service change from silently breaking a client.
- The aggregation code needs tests for partial failure: what the screen shows when the reviews service fails but orders succeeds.
Team ownership
- The pattern works best as "one experience, one BFF, one team": the mobile team owns the mobile BFF and can change screen and API in one release.
- A central team owning every BFF turns them back into a shared bottleneck.
- Logic duplicated across the web and mobile BFFs will drift. Anything that is a business rule (pricing, eligibility) belongs in a downstream service; the BFF only shapes, combines and adapts.
When to choose it, and alternatives
- Choose a BFF when two or more clients need materially different data shapes or flows and are built by different teams.
- Choose one general-purpose API when there is one client, or clients are nearly identical: separate BFFs would be duplication.
- Consider a GraphQL gateway (a single API where the client's query names exactly the fields it wants) when there are many clients with overlapping needs and you prefer self-serve field selection to per-client code. It solves over-fetching without a BFF per client, but moves query-cost control onto the server team.
Pitfalls
- Business logic creeping into the BFF, which then must be duplicated in every other BFF.
- A BFF per device type (separate iOS and Android BFFs) when the experiences are the same; split by experience, not by platform.
- BFFs calling other BFFs, which recreates chatty, coupled services at the edge.
- Mirroring internal models. A BFF that returns downstream objects unchanged gives the client a pass-through and none of the insulation.
Unlock Full Question Bank
Get access to all 16 Architectural Patterns and Anti-Patterns interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.