Secure Architecture and Design Principles Questions
Designing systems that are secure by construction: core design principles (least privilege, separation of duties, fail-safe and fail-secure defaults, secure-by-default, attack surface reduction, assume-breach), defense-in-depth and layered control placement, classifying controls as preventive, detective and corrective, secure design patterns such as tenant isolation and blast-radius limiting, security architecture reviews and secure-by-design checklists, and reasoning about trade-offs between security, usability, performance and delivery speed when selecting and placing controls, including build, native or buy choices and making the secure option the easy one for developers. Covers enterprise-scale reference architecture, such as placing enforcement across hybrid and multi-cloud estates and giving many teams a consistent baseline, how security requirements shape system structure, designing safeguards to degrade safely when a dependency is down or in an emergency, and testing whether layers and isolation hold. Boundary: the mechanics of identity, cryptography, networking, threat models, detection, incident response and compliance evidence are covered elsewhere.
You have a small team where the same two engineers write the code, review it, deploy it and hold the production secrets. How would you build separation of duties into provisioning, review, deployment and secret access when you cannot add headcount?
Sample Answer
Direct answer
With two people I cannot get full separation of duties (SoD: no single person can complete a risky action end to end), so I aim for two achievable things: no single account, and no single mistake, can change production or read secrets unnoticed. I do that by moving privileges from the humans to a pipeline, making each engineer the other's gatekeeper, keeping secret values out of human hands, and sending the evidence to someone outside the pair. Honest limit: two engineers can collude, which is why one outside reviewer remains part of the design. Break-glass, used below, means an emergency access route that is normally closed, opened only with approval, time-limited and loudly logged.
The four duties (engineers A and B)
| Duty | Control | Who |
|---|---|---|
| Provisioning (create cloud resources) | All infrastructure through infrastructure-as-code (infrastructure described in files in a repository, so every change is reviewed like code). Humans have read-only cloud roles day to day. Only the pipeline's role can create or change resources. | Pipeline role |
| Review | Branch protection (repository settings that enforce rules on the main branch): the author cannot approve their own pull request, one approval from the other engineer is required, direct pushes blocked, and the rule cannot be bypassed by admins without an alert. | A reviews B's work and B reviews A's |
| Deployment | Deploy runs only from the pipeline on the protected main branch. Production deploy needs an approval click by the engineer who did not author the change. | Pipeline plus the non-author |
| Secret access | Secrets live in a secret manager (a managed vault that stores values and releases them to authorised callers) and are read by workloads through their own identity (an identity given to the software itself, not to a person). No human role can read production values. Rotation is automated. | Workload identity |
Compensating controls (the safeguards that replace a missing person)
- Break-glass. If a human really needs production access, they use a separate emergency role that is time-limited (illustrative: one hour), needs the other engineer to approve, and pages an outside person.
- Logs the pair cannot edit. Cloud audit logs and pipeline logs go to a separate account whose owner is someone else (a manager or a fractional security contractor, meaning a part-time outside security specialist hired by the day), with write-only access for the pair.
- Monthly sample review. The outside person reviews all break-glass uses and a sample of merged changes. The two engineers are not marking their own work.
- Separate admin identities. Daily accounts have no admin rights. The only human route to administrator rights is the break-glass role above, assumed only when needed, so the rule that only the pipeline role creates or changes resources holds on a normal day.
What the flow looks like for engineers A and B (illustrative). The rule on main reads: require 1 approving review, ignore the author's own approval, dismiss approvals when new commits arrive, block force-pushes, apply to administrators. A opens a pull request and the pipeline runs tests and scans. B reads it and approves. Merge deploys to staging automatically. The production job then waits for an approval click from the engineer who did not write the change, here B. If A is away, the outside reviewer, who holds an approve-only role (it can approve pull requests and the production approval click, but cannot push code, run the pipeline or read secrets), approves instead. Each step leaves a record naming who did it.
At pipeline scale across several teams
The same ideas extend: build jobs assume short-lived roles via federated identity (the code host vouches for the job with a signed token and the cloud hands it a role that expires, so no long-lived keys exist), a policy-as-code gate (security rules written as automated checks, for example checking infrastructure plans against rules before apply) blocks risky changes automatically, and every deployment writes an audit record of who approved what, which commit, and which artifact.
Trade-offs and pitfalls
- When one engineer is away, the other cannot approve alone. Plan a named backup approver (the outside reviewer with the approve-only role described above), not a bypass.
- Shared accounts defeat everything above, because logs cannot say who acted.
- What would change my call: a regulated customer requiring independent approval, which would justify buying outside review capacity.
What does secure by default mean, and how would you apply it to a service you ship to other teams? Describe the defaults you would choose for a cloud-hosted service and how you handle teams who need to loosen them.
Sample Answer
Direct answer
Secure by default means that when someone uses the service with zero configuration, the result is already the safe one. Insecure behavior must be something a team deliberately asks for, never something they must remember to turn off. For a cloud service shipped to other teams, the safe path should also be the easiest path.
Defaults I would choose for the service
| Area | Default |
|---|---|
| Network | No public endpoint, no inbound access except from named callers; outbound traffic limited to what the service needs |
| Authentication | Required on every endpoint; no default or shared passwords; single sign-on (SSO) for humans, per-service identities for machines |
| Access | Least privilege roles; the creator is not automatically an administrator |
| Data | Encryption at rest on (using a managed key, one held and rotated by the cloud provider's key service, with every use logged), TLS (encrypted transport) required with an old-version floor (refuse old protocol versions, for example anything below TLS 1.2, because they have known weaknesses), storage never public |
| Logging | Audit logging on and shipped to a separate account the service team cannot edit |
| Updates | Automatic patching of images and dependencies |
The same idea applied to a Linux host image and a fresh cloud VM
- OS: a minimal, hardened, centrally built image with only needed packages; patched on a schedule.
- Services: nothing listening that was not asked for; SSH with keys or a managed session tool, password login off, root login off.
- Network: the default firewall rule (security group) allows no inbound traffic; the instance metadata service requires session tokens (IMDSv2 on AWS) so a web bug cannot easily steal credentials. The instance metadata service is a local address that a cloud VM can query for facts about itself, including temporary cloud credentials. In its first version a plain web request to that address returns them, so a bug that makes the application fetch an attacker-chosen URL (server-side request forgery) can leak them. IMDSv2 requires first obtaining a session token with a special request, which that kind of bug usually cannot make.
- Storage: disks encrypted, object storage with public access blocked.
Handling teams who need to loosen a default
Make loosening possible but visible: a self-service exception request names the owner, the reason, the compensating control (an alternative safeguard, such as an IP allow list when a port must open), and an expiry date. Low-risk loosening is approved by the service owner. High-risk loosening, such as a public endpoint or turning off encryption, needs a security sign-off and may be accepted by the business owner who carries the risk. Exceptions expire automatically and are reviewed.
Worked example
A team needs a public endpoint for a partner webhook. Instead of opening the service, they get a time-limited exception that allows exactly one path through the gateway with a WAF rule (a filter limiting that path to expected methods, sizes and request shapes) and request signing (the partner attaches a signature computed with a shared secret over the request body, so the service can verify who sent it and that it was not altered), approved for 90 days and renewed only if still used.
Pitfalls
Defaults that break every team's first deploy get disabled wholesale, so test them with real users. Track the number of active exceptions as a signal: a rising count means a default is wrong, not that teams are careless.
Here is an architecture: public web servers, a single application tier, and one database, all sitting behind one firewall. Where does the failure of a single control turn into a full compromise, and what would you change first to bound the damage?
Sample Answer
Direct answer
In this design one firewall is the only boundary, so any single failure turns into full compromise at three main points (the first three rows of the table below; the last three rows are amplifiers that make any breach worse): a compromised public web server sits on the same flat network as the app and the database; the single app tier holds a database credential that is probably broad; and the firewall itself is one policy, one admin plane and one failure point. I would change the network first: split the three tiers into separate zones with deny-by-default rules between them, because it is quick and limits every later failure.
Where one failure becomes everything
| Weak point | What fails | Result |
|---|---|---|
| Flat network behind one firewall | A web server is exploited | Attacker can scan and reach the app and database directly |
| Firewall rules too broad or misconfigured | One bad rule | All tiers exposed at once, no second boundary |
| App tier holds a broad database login | App is exploited (for example SQL injection) | Attacker can read and write all data, as the app can |
| One database for all data | Any database access | Everything is in the blast radius (the amount of damage one failure can cause) |
| No egress (outbound) filtering | Any host compromised | Data can leave freely; the attacker can fetch tools |
| Logs only on the same hosts | Host compromised | Attacker deletes the evidence |
What I would change, in order
- Segment into three zones: web servers in a DMZ (demilitarized zone, a network segment for internet-facing hosts), app tier and database in separate private segments. Allow only web to app on one port and app to database on one port; everything else denied. A web-server compromise now yields a web-server compromise, not the database. Enforce the zone boundaries at separate enforcement points rather than as more rules on the one existing firewall: for example host-level or cloud security-group rules on each tier, administered from a different account or console than the edge firewall, or a second internal firewall with its own admin credentials and rule source. If one device and one admin plane enforce every boundary, a single bad rule or one stolen admin login still removes all of them.
- Shrink the database login: a dedicated account for the app with only the operations on the tables it uses, no administrative rights. This bounds what a hacked app can do.
- Restrict outbound traffic from each zone to what it needs, and ship logs off-host to a separate system.
- Then separate sensitive data into its own store and add a second control type (such as a WAF in front of the web tier).
What would change my order
If the known weakness is SQL injection in the app, I would do the database account first. If the firewall is the bigger worry (for example a poorly governed rule set), the segmentation comes first anyway, provided it is enforced at separate points (see step 1), because then it adds a second boundary.
Pitfall
Adding more rules to the same single firewall is not a second layer; the layers must fail independently.
You are the first architect on a greenfield SaaS where performance and scale matter and enterprise customers will soon expect proof of security maturity. How would you decide which security controls to build into the architecture from day one, how would you keep them from hurting performance or developer speed, and how would you estimate and defend their ongoing cost?
Sample Answer
Direct answer. On day one I would build the controls that are cheap now and very expensive to retrofit: identity, tenant isolation, encryption and key handling, logging, and an automated build pipeline. I would defer controls that are expensive to run and easy to add later. I would deliver them as platform defaults so developers get them for free, and defend the cost as a small, explicit share of engineering spend tied to the enterprise deals it unlocks.
1. Deciding what goes in from day one. Score each candidate control on three questions: How hard to retrofit once data and customers exist? How much risk does it remove (what could we lose)? Do enterprise buyers ask for it? Retrofit pain is the main sorting rule.
| Build now | Why | Defer or keep light |
|---|---|---|
| Central sign-in with multi-factor and least-privilege roles | Hard to untangle later | Fine-grained attribute-based access for every feature (permissions decided by attributes such as department, region and data label rather than a few fixed roles) |
| Tenant isolation in the data layer | Retrofitting after launch is a rewrite | Separate infrastructure per tenant (offer later, as a paid tier) |
| Encryption in transit and at rest, managed keys | Cheap with cloud defaults | Customer-managed keys (the customer holds the encryption key and can cut off access by revoking it) |
| Audit logging with tenant and user context | Evidence and detection depend on it | Full security analytics platform |
| Infrastructure as code (the cloud setup written as reviewed text files instead of clicked together), dependency and secret scanning in the pipeline | Cheap, catches issues early | Formal bug bounty (a paid program inviting outside researchers to report vulnerabilities) |
| Backups, tested restore | Resilience baseline | Multi-region active-active (full copies running in several regions, all serving traffic at once) |
The right-hand column holds real options that a specific trigger brings forward: a bank customer asks for customer-managed keys, an outage review demands multi-region, a regulator asks for finer access rules. The other three deferred options get triggers too: separate infrastructure per tenant when a customer will pay for a dedicated tier, a formal bug bounty once the first external audit or penetration test has been closed out, and a full security analytics platform when log volume makes manual search unworkable. Each deferred option has a named trigger, so deferring is a scheduled decision rather than neglect.
2. Without hurting performance or speed.
- Secure by default (paved road): a service template with authentication, tenant scoping, logging and secure headers (response headers that make browsers enforce protections such as HTTPS-only) already wired in. The secure path is the easiest path.
- Cheap on the hot path: token verification and authorisation checks happen at the gateway with caching; heavy work (scanning, analytics) is asynchronous, off the request path. Measure latency in load tests with controls on, and set a budget agreed with the performance owner.
- Fast feedback: pipeline checks run in minutes and block only high-severity findings; the rest become tickets. Exceptions need an owner and an expiry date.
3. Evidence as a by-product. Because controls are code and pipeline steps, they leave records automatically: access reviews from the identity system, change approvals from pull requests, configuration state from infrastructure as code, logs from the platform. When an enterprise customer asks for a SOC 2 report (an independent auditor's attestation report on controls, measured against the AICPA Trust Services Criteria), much of the evidence already exists.
4. Estimating and defending cost (illustrative numbers). Assume 20 engineers at a loaded cost of 200,000 each (salary plus benefits, equipment and overhead, not salary alone), so engineering costs 4,000,000 a year. If the paved road and review process consume 3 percent of engineering time, that is 120,000. Add 60,000 for tooling. Total 180,000, which is 4.5 percent of engineering cost. Compare it with revenue at stake: if enterprise security reviews are blocking two deals of 250,000 annual value each, that is 500,000 a year. The point is not the figures (replace them with your own) but the shape: a named cost, a named benefit, and a review each quarter of whether the spend still earns its place. The 180,000 excludes the SOC 2 audit fee and an annual penetration test, which I would add as separate lines once an audit is scheduled, and the 500,000 is revenue, so I would compare the cost with the margin on it rather than the revenue. Retrofitting later is the alternative cost, and it is usually larger and arrives during a deal.
Pitfalls. Buying a tool stack before knowing the risks; adding a heavy gate that developers route around; promising maturity you cannot evidence. What would change my call: if the first customers are regulated, move audit logging and key management forward; if the data is low-sensitivity, defer more.
How do you think about preventive, detective and corrective controls when you design a system? Take a customer-facing web application and show how you would balance the three, and what a dangerous gap in the mix looks like.
Sample Answer
Direct answer
Controls come in three jobs. Preventive controls stop a bad event from happening. Detective controls notice that it is happening or has happened. Corrective controls limit the damage and restore normal operation. A sound design uses all three because prevention is never perfect, so you need to see failures and be able to act on them. The dangerous gap is a mix with no working detection or no working correction: you cannot respond to what you cannot see, and seeing without a way to act only gives you better information about the loss.
Customer-facing web app, worked example: credential stuffing (attackers replaying passwords leaked elsewhere)
| Class | Control | Stack tier |
|---|---|---|
| Preventive | MFA (multi-factor authentication), rate limits, check passwords against known breached lists | Application and edge |
| Detective | Alert when failed-login volume or distinct-IP spread passes a baseline; alert on a sudden rise in password-reset emails | Logging and monitoring |
| Corrective | Force reset for affected accounts, revoke active sessions, block abusive IP ranges, notify users | Application and operations |
More examples across the stack:
- Cloud storage: a rule that blocks public buckets (preventive), an alert when a bucket policy changes (detective), an automated job that reverts the change (corrective).
- Data: least-privilege database accounts (preventive), alerts on unusually large exports (detective), restore from tested backups (corrective).
- Code: dependency checks in the build pipeline (preventive), runtime anomaly monitoring (detective), a one-step rollback (corrective).
How I balance them
Weight toward prevention where it is cheap and reliable (patching, secure defaults). Weight toward detection and correction where prevention cannot be guaranteed: an attacker holding a valid stolen password looks like a normal user, so detection and fast session revocation matter more than another input filter.
Dangerous gaps
- Preventive only: a breach can run for months unseen.
- Detective without correction: alerts fire but nobody can disable the account or restore data quickly.
- Detective without an owner: the alert goes to an unwatched mailbox at 2 a.m.
- Corrective never tested: the backup that fails on restore day.
How a tester validates each class
Preventive: try to bypass it. Detective: run the attack in a safe test and confirm the alert fired, how fast, and that a human saw it. Corrective: run a restore or revocation drill and time it.
Unlock Full Question Bank
Get access to all 41 Secure Architecture and Design Principles interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.