Secure Architecture and Design Principles Questions
Designing systems that are secure by construction: core design principles (least privilege, separation of duties, fail-safe and fail-secure defaults, secure-by-default, attack surface reduction, assume-breach), defense-in-depth and layered control placement, classifying controls as preventive, detective and corrective, secure design patterns such as tenant isolation and blast-radius limiting, security architecture reviews and secure-by-design checklists, and reasoning about trade-offs between security, usability, performance and delivery speed when selecting and placing controls, including build, native or buy choices and making the secure option the easy one for developers. Covers enterprise-scale reference architecture, such as placing enforcement across hybrid and multi-cloud estates and giving many teams a consistent baseline, how security requirements shape system structure, designing safeguards to degrade safely when a dependency is down or in an emergency, and testing whether layers and isolation hold. Boundary: the mechanics of identity, cryptography, networking, threat models, detection, incident response and compliance evidence are covered elsewhere.
Your organization is breaking a monolith into microservices. What security properties did the old design give you for free that you will lose, and how would you rebuild them in the new architecture?
Sample Answer
Direct answer
A monolith gives you several security properties implicitly because everything runs in one process: calls between modules never cross a network, one code path enforces authorization, one log shows what happened, and one dependency set is patched at once. Splitting into microservices turns each of those into a distributed problem. I would rebuild each one explicitly, and I would put identity and service-to-service trust in place before the first service is carved out.
What is lost and how to rebuild it
| Property the monolith gave for free | What breaks in microservices | How I rebuild it |
|---|---|---|
| In-process trust: a function call cannot be spoofed or sniffed | Every call is a network request an attacker can forge or observe | Mutual TLS (both sides prove identity with certificates) and per-service identities; default-deny network policy |
| One authorization enforcement point | Many services, each could check permissions differently or not at all | Shared policy library or central policy engine (a service that answers "may this user do this?" from written rules), plus user identity passed in a signed token (a tamper-evident proof of who the user is) and checked at each service |
| Immediate session revocation | Tokens already issued stay valid across services until they expire | Short token lifetimes, revocation lists (lists of tokens cancelled before expiry) for high-risk actions, re-check sensitive operations online |
| One database and one transaction | Data spreads across services; multi-step operations can fail halfway, leaving security-relevant state inconsistent (for example, a refund is recorded but the step that revokes the customer's access never ran) | Each service owns its data with its own credentials; design sagas (a sequence of steps with compensating undo steps) and idempotent operations (safe to repeat: running the same request twice has the same effect as once, so a retry after a failure cannot double a refund) |
| One log and one audit trail | Actions span many services | Correlation IDs (one unique ID attached to a request and passed to every service it touches, so its trail can be reassembled) on every request, structured logs, central collection with tamper protection |
| One perimeter and one patch cycle | Many images, many dependencies, many endpoints | Standard base images, automated dependency scanning, an API gateway as the only public entry |
| One configuration and secret store | Many secrets, many places to leak | Vault with short-lived dynamic credentials per service |
If I could name only three in an interview, I would start with in-process trust (mutual TLS), the single authorization enforcement point, and revocation, because they are the hardest to retrofit once dozens of services call each other, while logging, patching and secret handling can be improved incrementally afterwards.
Worked example
In the monolith, refund(order) ran after the web layer checked the user's role in one place. In microservices, orders calls payments. If payments trusts anything from orders, an attacker who compromises orders can issue refunds for anyone. The fix is two checks: payments verifies the caller is really the orders service (mTLS), and verifies the end user's token authorizes a refund for this order. That avoids the confused deputy problem (a trusted service abused to use its own power for someone else).
Trade-offs
- Not everything should be split: keep security-critical, tightly coupled logic (such as the authorization decision) in few places.
- Each rebuilt control costs operations effort. Sequence the work: identity, then transport, then logging, then fine-grained policy. Splitting first and securing later leaves a long window with a flat, trusting internal network.
- Failure behavior now needs a decision: if the policy service is down, sensitive actions fail closed (deny, so that doubt means refusal rather than access), low-risk reads may use a short cache.
A product team brings you a new feature late in design. How do you decide which security requirements are must-haves before launch, which can follow, and which you would formally accept as risk?
Sample Answer
Direct answer. I sort every requirement by two questions: how bad is the harm if it is missing at launch, and how hard is it to fix later? High harm or hard-to-fix-later goes to must-have. Lower harm that is cheap to add afterward goes to a dated follow-up. Whatever remains and cannot be fixed in time is a formal risk acceptance, signed by someone with the authority to own that risk, with an expiry date.
Key terms. A risk acceptance is a written decision by a business owner to live with a known risk, recording what the risk is, why it is accepted, what limits it, and when it will be revisited. It is different from ignoring the risk. A compensating control is a substitute safeguard that reduces the same risk when the ideal control is not in place. A one-way door is a decision that is costly or impractical to reverse once it is live, such as a data model holding real customer data, where a trust boundary (the line where data moves between parts trusted differently) sits, or how identities work. A two-way door is easy to change later. A preventive control stops a bad event from happening; a detective control does not stop it but records and reveals it afterwards, such as an audit log with an alert.
The three buckets
| Bucket | Test | Who decides |
|---|---|---|
| Must-have before launch | Missing it can expose customer data, allow account takeover or cross-customer access, or the change is a one-way door (data model, trust boundary, identity design) that is costly to rework, or it is the only way to detect and investigate such exposure | Security architect, with the product owner informed |
| Can follow, dated | Harm is limited, a compensating control exists, and the fix does not change the architecture | Product owner commits to a date and owner; security tracks it |
| Formally accept | Harm is real but the fix is disproportionate or blocked, and the exposure is bounded | A business owner with authority over that risk level |
Worked example. The feature is "share a document through a public link". I would assess:
- Links must be unguessable and revocable (the owner can switch a link off so it stops working at once), and the shared view must not leak other documents. Must-have: a leaked link exposes customer data and the link scheme is hard to change after links exist in the wild.
- Link expiry defaults and an admin report of active links. Can follow: expiry is cheap to add later, and in the meantime the must-have revocation lets an owner switch off any link, which is the compensating control that limits exposure until expiry ships.
- Download-tracking and watermarking. Accept: nice to have, bounded harm, and the owner records that decision with a review date.
- Audit-log entries for link creation and link access. Must-have: it is the detective control (it cannot stop a leak, but with creation and access both logged it tells us who made a link and who opened it, and when), and without it we cannot investigate a leak.
How I handle a late design. Lateness changes the cost of rework, not the harm, so I protect the one-way doors first. I propose the smallest launch shape that makes the must-haves possible, for example limiting the feature to a pilot group for now, which narrows the exposure while the follow-ups land. I also say what would change my call: more sensitive data types, a larger audience, or a failed control test would move an accepted item back to must-have.
Pitfalls. Calling everything a must-have teaches teams to route around you. Accepting risk by email thread with no owner, scope, or expiry is not formal acceptance. A "can follow" with no date and owner is silent acceptance.
You are asked to create a secure-by-design checklist for architecture reviews of new services. What goes on it, what is mandatory versus advisory, and how do you stop it becoming a rubber stamp?
Sample Answer
Direct answer. A secure-by-design checklist is a short list of questions a new service must answer with evidence before launch. I would keep it to ten items, make seven of them launch blockers (mandatory), leave three as advisory, and stop it becoming a rubber stamp by demanding links to proof, giving the reviewer a real "not ready" outcome, and auditing the reviews themselves.
Key terms. An architecture review is a design-time meeting where a security architect examines a new service before it ships. Mandatory means the item blocks launch unless a named approver signs a written, time-limited exception. Advisory means the team answers in one line and the reviewer may recommend, not require. A rubber stamp is a review where every box gets ticked without anyone checking the claim. A trust boundary is a line on a diagram where data passes between parts that are trusted differently, such as browser to server, or your service to a third party. A negative test is a test that tries something that must be refused and passes only when it is refused. An egress allow-list is the list of outside destinations a service may call, with every other outbound connection blocked. Network segmentation means splitting the network into zones so that a compromised component cannot reach everything.
The checklist (10 items, 7 mandatory, 3 advisory)
| # | Item | Tier | Evidence the team must attach |
|---|---|---|---|
| 1 | Data inventory: what data enters, where it is stored, who can read it | Mandatory | Data-flow diagram with trust boundaries marked |
| 2 | Authentication and authorization: every endpoint requires identity, and access is checked on the server | Mandatory | Link to the authorization code path and a negative test (user A cannot read user B) |
| 3 | Least privilege: each service identity and human role holds only the permissions it uses | Mandatory | Exported role or policy definitions |
| 4 | Secrets and encryption: no secrets in code, keys in a managed store, data encrypted in transit and at rest | Mandatory | Secret-scan result and storage configuration |
| 5 | Logging and detection: security events (logins, permission changes, admin actions) reach the central log with a named alert owner | Mandatory | Sample log line and the alert rule |
| 6 | Failure behaviour: authentication and authorization fail closed (deny when the dependency is down), and the blast radius (how much one compromised component can reach) is stated | Mandatory | Test or design note showing the denied path |
| 7 | Build and dependency integrity: dependencies scanned, builds come from the pipeline, not a laptop | Mandatory | Pipeline run and scan report |
| 8 | Abuse controls: rate limits, quotas, bot protection | Advisory | One-line answer |
| 9 | Extra layers: egress allow-lists, tighter network segmentation | Advisory | One-line answer |
| 10 | A planned game day (a rehearsed attack or failure) within a quarter of launch | Advisory | Date on the calendar |
Seven mandatory plus three advisory is the full set of ten.
One row, bounced and accepted (illustrative). Take item 2 for a new orders service. Bounced: the form says "Authorization: Yes" and nothing else. Accepted: "Every route passes through the require_user check; ownership is checked in the order lookup (link to the code). The test user_a_cannot_read_user_b_order returns 403 and ran in pipeline run 4121 (link)." The reviewer opens the test and confirms it runs in the pipeline, not only on a laptop. The first version cannot be checked by anyone; the second can be checked in two minutes.
What decides the tier. An item is mandatory when failing it can expose customer data or let one tenant reach another, and when the fix is cheap before launch and expensive after. Anything that only adds a further layer of defense in depth (an additional, independent control behind the first) is advisory, because the service stays safe if it is missing. Two further items are mandatory by a second test: they are the controls that let us detect or limit damage when a first control fails. Logging and detection (item 5) cannot stop a breach, but without it nobody would notice one, and build integrity (item 7) decides whether what runs is what was reviewed. Abuse controls (item 8) are advisory only for a service with no public or unauthenticated endpoint; for a login page or public API, missing rate limits allow credential stuffing (trying stolen username and password pairs at volume), so the item becomes mandatory for that service. The tier is set per service, so the table shows the default for a typical internal service.
Five ways to stop it becoming a rubber stamp
- Evidence, not ticks. Each mandatory row needs a link to a diagram, configuration, or test result. "Yes" with no link counts as "no".
- Questions that force a story. Beside each item sits one prompt such as "describe what happens to a request when the identity provider is unreachable". A team that has not thought about it cannot answer fluently.
- A real fail outcome. The reviewer can return "not ready", and I track how often that happens. A checklist that has never blocked anything is decoration.
- Audit the reviews. Each quarter, pick a random sample of approved services and re-check the claims against the live configuration. Differences between what was claimed and what is deployed are the finding.
- Measure outcomes. Track the share of reviews that changed the design, the number and age of open exceptions, and incidents traced to reviewed services. If reviews change nothing, the review is too late or too shallow.
Trade-offs and pitfalls. A long list gets skimmed, so resist adding an item after every incident; retire or merge items instead. Tailor depth to risk: an internal tool with no sensitive data gets a lighter pass than a payment service. Run the review during design, not the week before launch, or "mandatory" turns into a fight about the date.
After an incident where stolen credentials led to lateral movement, leadership asks for an architecture roadmap so it cannot happen the same way again. What do you change, in what order, and why?
Sample Answer
Direct answer
I start from the attack chain of this incident and break every link in order of how cheaply and certainly I can: the credentials were stolen, they were usable, they carried too much privilege, they could reach other systems, and nobody noticed or stopped it in time. Lateral movement means an attacker who has got into one account or machine using it to reach others; the attack chain is the sequence of steps from first access to the damage. I do not know the incident's exact details, so the plan below states its assumptions. The order is: close the proven path now, remove standing privilege (admin rights that stay switched on permanently, whether or not anyone is using them) next, then change how identity and networks are built so one stolen credential is not enough.
Phase 1: contain and close the proven path (first 30 days)
- Rotate every credential the attacker could have touched and invalidate sessions and tokens. Reset service-account secrets reachable from the compromised machines.
- Require phishing-resistant multi-factor authentication (hardware keys or passkeys: a login method where a private key stays on the device and works only for the genuine website, so a fake login page cannot capture anything replayable, unlike texted codes) first for administrators, because their credentials are the prize.
- Remove obvious standing admin rights and shared accounts.
- Add detection for the movement that happened: alerts for logins from a new place, admin logons to servers by non-admin workstations, and sudden use of one account across many hosts.
Phase 2: shrink what one credential can do (30 to 120 days)
- Just-in-time privileged access: admin rights granted for a task and an hour, not permanently.
- Separate admin accounts and admin workstations from daily-use ones.
- Unique local administrator passwords per machine (so cracking the built-in admin password on one laptop does not open every other machine; tools such as Microsoft LAPS automate this) and no shared service-account passwords.
- Network segmentation between user, server and sensitive zones with default deny, so a foothold cannot reach everything.
Phase 3: change the design (4 to 18 months)
- Phishing-resistant authentication for everyone, then reduce password reliance.
- Short-lived workload credentials instead of long-lived keys.
- Fine-grained segmentation around the most valuable systems and per-service identity.
- Assume-breach engineering: per-system audit logs sent to a store admins cannot edit, and tested containment.
Replaying the incident step by step (illustrative)
| Attacker step | Before | After the phase that breaks it |
|---|---|---|
| Phish an employee's password | Works | Admins: fails after phase 1 (hardware key). Everyone: after phase 3 |
| Land on a workstation and find a cached administrator session | Works | Fails after phase 2 (no standing admin, just-in-time rights only) |
| Reuse one local admin password across servers | Works on every server | Fails after phase 2 (unique password per machine) |
| Reach the database from the user network | Works | Fails after phase 2 (default-deny segmentation) |
| Go unnoticed for days | 14 days | Phase 1 alert fires within minutes |
An example alert from phase 1: "service account svc-backup logged on to 14 distinct hosts within 10 minutes (normal: 2 hosts a day)". An example measure: 12 of 40 privileged accounts on phishing-resistant authentication is 30 percent at day 30, with the target 40 of 40; standing domain-admin-level accounts falling from 9 to 2. All figures are illustrative.
Why this order
Phase 1 is cheapest and targets what we know happened, so it reduces repeat risk fastest. Phase 2 reduces the damage of the next unknown theft. Phase 3 is expensive and slow but removes the cause class. Each step is chosen by risk removed per effort, and the dependencies matter: segmentation is of little use while standing admin credentials work everywhere.
Proof and communication
Define a measure per phase, for example the share of privileged accounts using phishing-resistant authentication and the number of standing domain-admin-level accounts (accounts that can administer every machine in a Windows Active Directory domain, or the top-level admin role in a cloud account). After each phase, a red team (testers who play the attacker) replays the original path and reports where it now fails. Leadership gets the chain, the phase, the owner, the cost and the residual risk, and decides what to accept. Say plainly that this reduces the chance and the damage, and that no roadmap makes recurrence impossible.
Pitfalls
Buying a detection product without removing privilege, a roadmap with no replay test, and a plan that fixes only the exact path while leaving the class open.
In a shared relational database serving many tenants, how do you make sure a bug in application code cannot return one tenant's rows to another? Where do you place the enforcement, and what does it cost in performance and operations?
Sample Answer
Direct answer
Put the enforcement in the database, below the application, using row-level security (RLS): the database attaches a tenant filter to every query itself, so a forgotten WHERE clause in application code returns no other tenant's rows. The cost is some query overhead and real operational discipline, mainly around connection pooling, privileged roles and new tables.
Placement: three layers, each stronger than the one above
- Application layer (ORM scopes, automatic tenant filters added by the object-relational mapper library, and middleware, shared code that runs on every request): convenient but is the layer that has the bug, so it cannot be the only guard.
- Database policy: RLS on every tenant table, using a role the app connects as.
- Data protection underneath the policy: per-tenant encryption keys for the most sensitive columns, if the threat includes a database admin.
Postgres example (behavior checked on PostgreSQL 16)
How to read the policy below: the first line turns row-level security on for the orders table. The second makes the policy apply even to the role that owns the table (normally owners skip it). USING is the filter for rows you may read, update or delete. WITH CHECK is the same test applied to rows you write, so you cannot insert a row for another tenant. current_setting('app.tenant_id', true) reads a per-connection setting named app.tenant_id (the true means return nothing instead of an error if it was never set). NULLIF(x, '') turns an empty string into NULL, and ::uuid converts the text to the uuid type used by tenant_id; comparing to NULL matches no rows.
ALTER TABLE orders ENABLE ROW LEVEL SECURITY;
ALTER TABLE orders FORCE ROW LEVEL SECURITY;
CREATE POLICY tenant_isolation ON orders
USING (tenant_id = NULLIF(current_setting('app.tenant_id', true), '')::uuid)
WITH CHECK (tenant_id = NULLIF(current_setting('app.tenant_id', true), '')::uuid);
Per request, inside a transaction, the app runs SET LOCAL app.tenant_id = '<tenant uuid>' (SET LOCAL sets the value only until the current transaction ends; the function call set_config('app.tenant_id', value, true) does the same and accepts a bound parameter), then its queries. In my test, orders held one row for tenant 1 and one for tenant 2:
- With no tenant set, the app role saw zero rows.
- With tenant 1 set, it saw only tenant 1's row, and inserting a row for tenant 2 failed with a row-level security violation.
- A superuser connection still saw all rows even with FORCE, because superusers and roles with BYPASSRLS skip policies. FORCE applies the policy to the table's owner.
- After a transaction that used SET LOCAL ended, the setting reads as an empty string on that connection, and a bare uuid cast errored; the NULLIF in the policy turns that into zero rows.
Operational rules
- The application connects as a role that is not the owner, not a superuser and not BYPASSRLS (a role attribute that makes a role skip all row-level policies). Not being a superuser or BYPASSRLS is the non-negotiable part, because those roles skip every policy; not being the owner matters because an owner can disable row-level security or drop the policy (FORCE only makes the policy apply to the owner's own queries). Together with the transaction-scoped setting below, these are the core rules; the rest harden the design. Migrations and admin tools use a separate role.
- Behind a transaction-mode connection pooler (such as PgBouncer, a proxy that lends one database connection to different clients between transactions), session-level SET can leak a tenant to the next client, so use transaction-scoped settings only.
- Every new table needs a tenant column and a policy. A CI check queries the catalog and fails on a tenant table without RLS.
- Put tenant_id first in indexes and in composite unique keys, so uniqueness is per tenant and tenant filters stay fast.
- Reporting and background jobs that cross tenants run as a named, audited role.
Performance and operations cost
The policy adds a predicate to every query. With tenant_id leading the indexes the planner can use it, but check with EXPLAIN (a command that shows the query plan the database's planner chose) on real queries, since cost depends on the schema and policy shape. Debugging gets harder (a query "returns nothing" because of context, not data), and each request pays a small extra round trip to set context unless it is bundled with the first statement.
Trade-offs
If the threat includes insiders or a strict compliance demand, schema-per-tenant or database-per-tenant is stronger but costs migration and connection overhead per tenant.
Unlock Full Question Bank
Get access to all 47 Secure Architecture and Design Principles interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.