Secure Architecture and Design Principles Questions
Designing systems that are secure by construction: core design principles (least privilege, separation of duties, fail-safe and fail-secure defaults, secure-by-default, attack surface reduction, assume-breach), defense-in-depth and layered control placement, classifying controls as preventive, detective and corrective, secure design patterns such as tenant isolation and blast-radius limiting, security architecture reviews and secure-by-design checklists, and reasoning about trade-offs between security, usability, performance and delivery speed when selecting and placing controls, including build, native or buy choices and making the secure option the easy one for developers. Covers enterprise-scale reference architecture, such as placing enforcement across hybrid and multi-cloud estates and giving many teams a consistent baseline, how security requirements shape system structure, designing safeguards to degrade safely when a dependency is down or in an emergency, and testing whether layers and isolation hold. Boundary: the mechanics of identity, cryptography, networking, threat models, detection, incident response and compliance evidence are covered elsewhere.
A small company leaked customer data because a storage bucket was publicly readable. Which design-time controls should have made that impossible rather than merely detectable, and how would you prioritize them across prevention, detection and recovery?
Sample Answer
Direct answer
The leak should have been blocked by controls that make public exposure impossible to configure, not by an alert after exposure. Priority: prevention first (a guardrail that cannot be switched off by a developer), then fast detection for anything that slips past, then recovery so the damage is bounded and the response is rehearsed. For a small company this is mostly a handful of one-time settings.
Design-time controls that make it impossible
- Block Public Access, enforced at account or organization level. Amazon S3 (AWS's object storage) has settings that reject public ACLs (access control lists, per-object permission grants) and public bucket policies, and can be applied to a whole account or an AWS Organization (the AWS grouping of many accounts under one owner), overriding bucket-level settings. AWS documents that new buckets do not allow public access by default. Other clouds offer the same idea as a storage-account or organization-level switch that forbids public access; the principle is to enforce it above the level where a developer works.
- Stop people from turning it off. Use an SCP (service control policy, an organization-wide permission ceiling) that denies the actions which change the public-access settings, so only a small platform role can. An illustrative statement:
{
"Effect": "Deny",
"Action": ["s3:PutAccountPublicAccessBlock", "s3:PutBucketPublicAccessBlock"],
"Resource": "*",
"Condition": {"ArnNotLike": {"aws:PrincipalArn": "arn:aws:iam::*:role/platform-admin"}}
}
Every request to loosen those settings is refused unless it comes from the platform-admin role, however much access the developer otherwise has.
3. Policy-as-code in the deployment pipeline. Policy-as-code means the security rules are written as automated checks, and infrastructure-as-code means the cloud setup is defined in template files that go through review. The check fails the build if a template contains a public ACL or a bucket policy with principal "*" (meaning anyone on the internet, signed in or not) and no restricting condition. This is the kind of statement it rejects:
{"Effect": "Allow", "Principal": "*", "Action": "s3:GetObject", "Resource": "arn:aws:s3:::customer-exports/*"}
Adding a condition that pins access to a fixed value, such as one specific network or one account, makes it non-public.
4. Serve public files differently. If the site needs public images, put a content delivery network in front and keep the bucket private, so there is no reason for any bucket to be public.
5. Separate data by sensitivity: customer data in its own bucket, encrypted with a key whose access is limited, so a mistake in one place does not expose it.
Prioritization across prevention, detection and recovery
| Order | Control type | Example | Why this rank |
|---|---|---|---|
| 1 | Preventive | Org-wide Block Public Access, SCP, CI check | Cheap, global, removes the failure class |
| 2 | Detective | IAM Access Analyzer (an AWS feature that reports resources reachable from outside your account or organization) external-access findings, config-change alerts routed to a person | Catches exposures that arrive via a path you did not block |
| 3 | Corrective | Runbook to block access, rotate keys, review access logs to see who read what; versioning and backup | Limits and measures harm after the fact |
The difference matters in time: a daily scan could leave a bucket exposed for up to a day, while the preventive setting rejects the change at the moment it is attempted.
What would change my call
If the business genuinely hosts public downloads, I would allow one dedicated public bucket by exception, with an owner, no sensitive data and an expiry review, while the block stays on everywhere else.
When a security component fails, should it fail open or fail closed? Walk me through your choice for an authentication service, a payment gateway and an operational monitoring pipeline, and tell me what would change your answer.
Sample Answer
Direct answer
The answer depends on what the failure would harm. Fail closed (deny when the check cannot run) when failing open would expose data or money. Fail open (allow when the check cannot run) when closing would stop something that must keep working and the check itself is not the thing protecting the asset. The aim is to choose deliberately for each component, then make the failure visible. Terminology: "fail-secure" means the secure state on failure (usually closed); "fail-safe" means the state that minimizes harm to people, which can be open.
The three components
| Component | Choice | Reasoning | What I add |
|---|---|---|---|
| Authentication service | Fail closed | Failing open lets anyone in as anyone | Multi-zone replicas; already-issued tokens stay valid until they expire so existing sessions keep working; break-glass admin path |
| Payment gateway | Fail closed for normal and high-value payments; optionally a small capped fallback | Approving payments with no authorization or fraud check turns an outage into a fraud loss | Any fallback is a business decision: low per-transaction cap, daily total cap, signed off by the risk owner and logged |
| Operational monitoring pipeline | Fail open for the product, loudly | Taking the product down because the log shipper broke trades a visibility gap for an outage | Buffer data locally, alert on gaps, and fail closed for the narrow set of actions where an audit record is required (for example privileged admin changes) |
What would change my answer
- Life safety: in a medical device, fail-safe beats fail-secure. An infusion pump that cannot reach its authentication server should continue or move to a safe state rather than stop therapy; building door locks release on power loss so people can leave.
- Data sensitivity and regulation: the more sensitive the asset, the stronger the case for closed.
- Business impact: compare the cost of an hour of downtime with the exposure from allowing unchecked access; this is for the owner of the risk to decide, not only security.
- Dependency design: a check with a cached recent result or a local fallback lets you stay safe without a hard outage.
Feature flags
Make the failure mode a runtime setting (a feature flag such as closed or open-with-limits), so the team can change it during an incident without a deploy. Default each component to the choice in the table above (closed for authentication and payments, open with local buffering for the monitoring pipeline), require approval and logging to change it, and test both modes with deliberate fault injection (breaking the dependency on purpose in a test) so the fallback is known to work.
Explain the principle of least privilege and give one concrete way you would enforce it for human users and one for machine identities in a cloud environment. Where does it usually erode over time?
Sample Answer
Direct answer
Least privilege means every person, program or service gets only the access it needs to do its current job, for only as long as it needs it, and nothing more. The point is to shrink the damage when an account is stolen or misused: a stolen account that can read one folder is a small incident, one that can administer everything is a large one.
Humans: just-in-time elevation (JIT)
Day to day, an engineer holds a low-privilege role through group membership. When they need production access, they request it for a stated reason and a short window, a second person or an automated rule approves, and the access expires on its own (for example after an hour). Everything done during the window is logged. The standing "always admin" account disappears, so a phished laptop yields little.
Machine identities: one scoped, short-lived identity per workload
A machine identity is the credential a program uses (a service account or cloud role). Give each service its own role, name the exact actions and resources, and issue short-lived credentials from the platform rather than long-lived keys pasted into config. Example for an invoice service that only reads PDFs (this is an AWS-style policy; other clouds use the same ideas under different syntax):
{
"Effect": "Allow",
"Action": ["s3:GetObject"],
"Resource": ["arn:aws:s3:::acme-invoices/*"]
}
Line by line: Effect: Allow says this rule grants access; Action lists the one operation permitted, s3:GetObject, meaning read a stored file; Resource names exactly what it applies to, here every file in the one storage bucket called acme-invoices (the long arn:aws:s3:::acme-invoices/* string is AWS's unique name format for a resource, and the * means any file inside it). Anything not allowed is denied by default. It cannot list other buckets, write, or touch the database. A related capability-based pattern (holding the token is itself the permission) hands out a narrow token for one operation instead of broad credentials. For example, a pre-signed link is a web address with a signature built in that lets whoever has it download one specific file for a few minutes, with no account needed.
Where it erodes over time
- Privilege creep: Priya moves from the payments team to search but keeps her payments database access, because nobody removes access on a transfer.
- Temporary becomes permanent: an emergency wildcard policy ("Action": "*", where the asterisk means every possible action) added during an outage and never narrowed.
- Orphaned service accounts and keys from retired projects.
- Copy-pasted policies that carry more permissions than the new service needs.
- Request friction: if getting the right permission is slow, people ask for the broad one.
Counter-measures
Tie access to team membership so it moves with the person; expire grants by default; compare granted permissions to permissions actually used (cloud providers report last-used data, and AWS IAM Access Analyzer, one example tool, can generate findings for unused roles, keys and permissions; other platforms have comparable reports) and remove the difference.
Explain defense in depth to me as you would to a new engineer, then show how you would apply it to an enterprise web application running in a hybrid cloud. What makes layers genuinely independent rather than merely redundant?
Sample Answer
Direct answer
Defense in depth means protecting an asset with several different controls in a row, so that when one fails (and eventually one will) the attacker still faces the next. Think of an airport: ID check, ticket check, security scan, boarding gate. A forged ticket gets past one step but not all of them. Layers are only worth having if they are independent: an attacker who beats layer one by some method should gain nothing toward beating layer two.
What makes layers independent rather than merely redundant
Redundant layers fail together; independent layers fail for different reasons. Check four things for any pair of layers:
- Different failure mode: a WAF (web application firewall, a filter that inspects web requests for attack patterns) is defeated by an encoding trick; parameterized database queries are not affected by that trick. For example, a WAF rule may block the text
' OR 1=1 --(a classic SQL injection string). An attacker sends the same characters percent-encoded (%27%20OR%201%3D1--, or encoded twice) in a way the WAF does not decode but the application does, so the attack slips through. A parameterized query sends the SQL text and the user's value to the database separately (SELECT * FROM users WHERE name = ?with the value supplied on its own), so the database treats whatever arrives as plain data and never as SQL, however it was encoded. - Different trust boundary and credentials (a trust boundary is the line where data or requests pass from a less-trusted zone into a more-trusted one, such as internet to web tier, or web tier to database): if one admin password or one management console controls both layers, compromising it removes both.
- Different technology or rule source: two firewalls from one vendor with one copied rule set share the same bug and the same misconfiguration.
- Different owner or change path: one bad deployment should not weaken both.
Quick test to apply in a design review: "Assume layer N is completely bypassed. What does the attacker still need to do?" If the answer is "nothing", the layers were redundant.
Worked example: SQL injection against an enterprise web app in a hybrid cloud (on-prem datacenter plus a public cloud)
| Layer | Control | Still protects if the layer above failed | Pentest check (a penetration test is an authorized, simulated attack by a tester) |
|---|---|---|---|
| Edge | WAF and DDoS protection (flooding attacks) | n/a | Send encoded payloads; then hit the app directly, skipping the WAF |
| Network | Separate segments for web, app and database; deny by default between them, in the datacenter and in the cloud network | A compromised web host cannot reach the database | From a web-tier foothold, try to connect to the database port |
| Identity | MFA (multi-factor authentication) for users; one separate identity per service | Stolen password alone is not enough | Replay a stolen session or password |
| Application | Parameterized queries, input validation | Injection fails even with no WAF | Test injection with the WAF disabled |
| Data | Database account limited to the tables and operations the app needs; encryption at rest | A successful injection reads little, not everything | Run injected queries and see what the account can reach |
| Detection | Central logging to a SIEM (security information and event management, a system that collects and alerts on logs) from both environments | Attack is noticed within minutes | Confirm an alert fires on the test attack |
In the hybrid case, keep the layers consistent across both sides: the same segmentation rules expressed as code in the datacenter and the cloud, one identity provider but separate break-glass (emergency) admin accounts per environment, and one log destination so the picture is not split.
The same idea in a streaming pipeline (Kafka brokers feeding Spark jobs)
Kafka is a system that carries streams of records between programs (brokers are its servers; producers write records and consumers read them), and Spark jobs process those records. The network zone limits who can reach the brokers; TLS plus authentication on every producer and consumer; per-topic ACLs (access control lists) so each job reads only its topics; schema validation at ingest so poisoned records are rejected; encryption at rest; audit logs of topic access.
Pitfalls
- Counting layers rather than testing independence.
- Layers nobody monitors: a layer that fails silently is no layer.
- Validating only from outside. Test each layer assuming the previous one failed (an assume-breach test, where the tester starts with an internal foothold).
How do you think about preventive, detective and corrective controls when you design a system? Take a customer-facing web application and show how you would balance the three, and what a dangerous gap in the mix looks like.
Sample Answer
Direct answer
Controls come in three jobs. Preventive controls stop a bad event from happening. Detective controls notice that it is happening or has happened. Corrective controls limit the damage and restore normal operation. A sound design uses all three because prevention is never perfect, so you need to see failures and be able to act on them. The dangerous gap is a mix with no working detection or no working correction: you cannot respond to what you cannot see, and seeing without a way to act only gives you better information about the loss.
Customer-facing web app, worked example: credential stuffing (attackers replaying passwords leaked elsewhere)
| Class | Control | Stack tier |
|---|---|---|
| Preventive | MFA (multi-factor authentication), rate limits, check passwords against known breached lists | Application and edge |
| Detective | Alert when failed-login volume or distinct-IP spread passes a baseline; alert on a sudden rise in password-reset emails | Logging and monitoring |
| Corrective | Force reset for affected accounts, revoke active sessions, block abusive IP ranges, notify users | Application and operations |
More examples across the stack:
- Cloud storage: a rule that blocks public buckets (preventive), an alert when a bucket policy changes (detective), an automated job that reverts the change (corrective).
- Data: least-privilege database accounts (preventive), alerts on unusually large exports (detective), restore from tested backups (corrective).
- Code: dependency checks in the build pipeline (preventive), runtime anomaly monitoring (detective), a one-step rollback (corrective).
How I balance them
Weight toward prevention where it is cheap and reliable (patching, secure defaults). Weight toward detection and correction where prevention cannot be guaranteed: an attacker holding a valid stolen password looks like a normal user, so detection and fast session revocation matter more than another input filter.
Dangerous gaps
- Preventive only: a breach can run for months unseen.
- Detective without correction: alerts fire but nobody can disable the account or restore data quickly.
- Detective without an owner: the alert goes to an unwatched mailbox at 2 a.m.
- Corrective never tested: the backup that fails on restore day.
How a tester validates each class
Preventive: try to bypass it. Detective: run the attack in a safe test and confirm the alert fired, how fast, and that a human saw it. Corrective: run a restore or revocation drill and time it.
Unlock Full Question Bank
Get access to all 10 Secure Architecture and Design Principles interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.