Threat Modeling and Attack Surface Analysis Questions
Systematically identifying how a system can be attacked and where its exposure lies. Covers structured methodologies (STRIDE, PASTA, DREAD, OCTAVE, attack trees), enumerating and reducing attack surface, mapping trust boundaries and data flows via DFDs, profiling likely threat actors, and prioritizing identified threats by likelihood and impact during design. Includes applying this methodology to specific architectural substrates (cloud-native and serverless, microservices, ML/AI systems, IoT, CI/CD pipelines, cryptographic subsystems) and operationalizing it as a recurring program (SDLC integration, governance, tooling, KPIs). The proactive 'think like an attacker before you build' discipline: distinct from live penetration testing (the adversarial validation of a built system), from runtime detection/monitoring (recognizing an attack already in progress), and from implementing the resulting security controls (a separate design-and-build discipline).
Given this architecture description, identify hidden trust boundary misconfigurations and propose design changes:
- Mobile app (public) communicates with an API Gateway (public L7) → routes to API service running in the same VPC as the database
- API service connects directly to Database on port 5432 with a single shared DB credential
- Admin UI deployed in the same cluster as the API uses the same DB credentials and exposes an admin route on a public load balancer
- CI system can run arbitrary scripts that have network access to the cluster
List at least five issues, explain the associated risks, and propose fixes prioritized by impact and effort.
Sample Answer
Direct answer
The architecture has one structural flaw repeated five different ways: components that should sit on opposite sides of a trust boundary are instead sharing a boundary, most visibly a single shared database credential used by both the API service and the admin UI, and a build system with unrestricted network reach into the same cluster that credential lives in. None of these five issues individually requires an exotic attack; each is a direct consequence of "everything in the same VPC/cluster is treated as equally trusted," which is precisely the assumption a trust-boundary review exists to catch. Fix order should prioritize removing the admin UI's public exposure and splitting the shared database credential first, since both are high impact and comparatively low effort, before the harder network-segmentation and CI-hardening work.
Structured elaboration
Issue 1: Admin UI exposed on a public load balancer
Risk: an administrative interface, by definition capable of higher-privilege actions than the customer-facing API, is reachable directly from the internet with the same exposure as the public mobile-app-facing path. This collapses a trust boundary that should exist between "anyone on the internet" and "an authenticated administrator," turning any vulnerability in the admin UI itself (an authentication bypass, a vulnerable dependency) into a directly internet-exploitable path to administrative access, rather than requiring an attacker to first gain internal network access.
Fix: move the admin UI behind a private network path (a VPN, a bastion, or an internal-only load balancer, plus a separate strong authentication mechanism such as multi-factor authentication) so it is not reachable from the public internet at all. Impact: high (removes a direct internet-to-admin path entirely). Effort: low-to-medium (a load balancer/routing change plus, ideally, adding a stronger auth requirement; does not require restructuring the application itself).
Issue 2: Single shared database credential across API service and admin UI
Risk: because both components authenticate to the database as the same identity, the database cannot distinguish "the customer-facing API doing a normal customer-scoped query" from "the admin UI doing an administrative action," which means (a) neither component's database access can be scoped to only what it actually needs (least privilege is impossible when the credential is shared), (b) a compromise of either component grants the SAME database access as compromising the other, so the admin UI's public exposure (Issue 1) effectively also exposes whatever the API service's database access allows, and (c) there is no way to audit which component performed a given database action, since the credential identity is identical for both.
Fix: issue distinct, narrowly-scoped database credentials per component (the API service gets exactly the customer-data permissions it needs; the admin UI gets its own, separately-scoped and separately-audited credential), ideally via a secrets manager issuing short-lived, per-component credentials rather than static long-lived shared ones. Impact: high (this is the single change that most directly limits blast radius across the whole architecture, since it decouples the two components' compromise consequences from each other). Effort: medium (requires a database-permissions redesign and a secrets-management integration, but does not require re-architecting network topology).
Issue 3: API service and database share the same network segment with no isolation
Risk: the API service connects directly to the database on its default port with no indication of a network-layer boundary (a dedicated database subnet with restrictive security-group rules, a network policy limiting which services can reach the database at all) separating them. This means ANY other workload that ends up running in the same VPC, whether through a future deployment, a compromised unrelated service, or (concretely, per Issue 4) a CI job with cluster network access, can potentially reach the database directly, not just the API service that is supposed to be its only legitimate client.
Fix: place the database in its own subnet with security-group or network-policy rules that allow inbound connections ONLY from the specific API service's network identity, not from the VPC broadly. Impact: high (this is the control that actually enforces "only the API service talks to the database," which the architecture currently only achieves by accident, since nothing structurally prevents another workload from doing the same). Effort: medium (network-policy or security-group changes, testable incrementally, but requires care not to break the legitimate path while restricting everything else).
Issue 4: CI system has arbitrary script execution and network access to the cluster
Risk: this is the most severe issue in the set, because it means the build system, which by design runs code from pull requests, dependency updates, and build scripts, sits INSIDE the same trust boundary as production, with network reachability to everything else described above (the database, given Issue 3's lack of isolation, and potentially the admin UI's cluster-internal path even if Issue 1 is fixed). A single compromised dependency or a malicious pull request effectively grants an attacker the same network position as any other cluster workload, without needing to compromise the API service or admin UI at all.
Fix: run CI build agents in a separate, isolated network with no direct path to production systems; where a CI job genuinely needs to interact with production (a deployment step), route that through a narrowly-scoped, audited deployment mechanism rather than general network reachability, and use ephemeral, least-privilege build agents rather than long-lived ones with broad access. Impact: high (closes what is otherwise a direct path from "arbitrary code an attacker can influence" to "the same network as the database"). Effort: high (typically requires a genuine network-topology redesign separating CI infrastructure from the production cluster, which is more disruptive than the other fixes and needs careful sequencing to avoid breaking legitimate deployment paths).
Issue 5: No stated authentication/authorization boundary between the API Gateway and the API service
Risk: the description states the API Gateway routes to the API service, with no mention of the API service independently verifying that a request actually came through the gateway with proper authentication applied, versus trusting any request that reaches it on the internal network. If the API service implicitly trusts anything that reaches it over the internal network (which Issue 3's lack of network isolation makes broader than intended), then any other workload on that network, again including a compromised CI job, could call the API service directly, bypassing whatever authentication the gateway is supposed to enforce.
Fix: have the API service independently verify a signed assertion of gateway-applied authentication (rather than implicitly trusting all internal traffic), so the internal network position alone is not sufficient to act as an authenticated caller. Impact: medium-to-high (closes a defense-in-depth gap; the specific severity depends on what the gateway's authentication would otherwise be relied on to fully prevent). Effort: medium (requires the API service to add a verification step, generally a contained code change rather than an infrastructure change).
Prioritized order
| Priority | Issue | Impact | Effort |
|---|---|---|---|
| 1 | Admin UI public exposure | High | Low-medium |
| 2 | Shared database credential | High | Medium |
| 3 | CI arbitrary network access to cluster | High | High |
| 4 | No network isolation around the database | High | Medium |
| 5 | Implicit internal trust between gateway and API service | Medium-high | Medium |
Issues 1 and 2 lead because they combine high impact with comparatively achievable effort and can each be fixed largely independently of the others; Issue 4, the CI system's network access, is sequenced third despite being arguably the most severe root cause, specifically because its fix (a genuine network-topology redesign separating CI from production) is the most disruptive, and it benefits from Issues 1-2 already having reduced what a CI-originated compromise could reach in the meantime. Issue 3's database isolation then follows at priority 4, cheaper than the CI redesign but narrower in what it closes on its own.
Worked example
Tracing a single realistic compromise path through the CURRENT architecture, using the issues above, to show them compounding rather than being independent findings:
- An attacker compromises a transitive dependency pulled in during a routine build (a supply-chain attack, not requiring any flaw in the application code itself).
- Because the CI system has network access to the cluster (Issue 4) with no isolation from production, the malicious build step can attempt to reach cluster-internal services directly.
- Because the database has no network-layer isolation restricting it to only the API service (Issue 3), the compromised build step can reach the database directly on port 5432.
- Because the credential is shared between the API service and the admin UI, and effectively usable by anything that can reach the database on the network (Issue 2), the attacker does not need to steal a distinct "admin" credential separately; the one credential in use grants full access regardless of which component it was originally issued for.
- The result: a compromised dependency in a routine build reaches full database access, without ever needing to exploit the mobile app, the API gateway, the API service's application logic, or the admin UI's public exposure at all. This is exactly why Issue 4 (CI network access) is arguably the root cause connecting the others, even though Issues 1 and 2 are fixed first for practical sequencing reasons.
Trade-offs and pitfalls
- Fixing issues in isolation without checking for exactly this kind of compounding path is a common mistake. Each of the five issues looks individually moderate in isolation; the worked example shows why they need to be evaluated as a connected graph, not five independent checklist items, since the actual severity comes from the chain, not any single link.
- "Same VPC" is often mistaken for "trusted network," when it should mean nothing about trust on its own. A VPC boundary is a network-routing construct; treating co-location within it as equivalent to a trust boundary, rather than defining trust boundaries explicitly with their own enforcement (network policies, distinct credentials, distinct authentication), is the single assumption underlying all five issues here.
- Sequencing by effort alone, ignoring impact, risks fixing the cheap issues while leaving the most severe one (CI network access) unaddressed indefinitely because it is always the hardest and therefore always gets deprioritized; the prioritization above deliberately keeps Issue 4 at priority 3, not last, specifically to avoid that trap.
- A partial fix to the shared-credential issue (rotating the credential without actually splitting it into per-component identities) looks like progress but does not change the underlying problem: the database still cannot distinguish which component is acting, so blast radius and auditability are unimproved even though "the password changed."
Walk through a threat modeling exercise for a new cloud-native microservice that accepts file uploads and stores them in object storage. Use an explicit framework (e.g., STRIDE) to identify assets, actors, threats, attack paths, and mitigations. List the artifacts you'd produce (data flow diagram, threat list, prioritized mitigations) and one example detection control for a critical threat.
Sample Answer
Direct answer
A threat-modeling exercise for a cloud-native file-upload microservice produces three concrete artifacts: a data-flow diagram (DFD) that names every asset, actor, and trust boundary; a threat list built by walking each element of the DFD against STRIDE (Spoofing, Tampering, Repudiation, Information Disclosure, Denial of Service, Elevation of Privilege); and a prioritized mitigation list ranking those threats by realistic impact and likelihood. For this specific design, the highest-priority threat is Elevation of Privilege through the asynchronous processing worker, since a compromised file-processing step inherits whatever cloud permissions that worker's identity holds, and the concrete detection control below is built around exactly that threat.
Structured elaboration
Assets, actors, and the data-flow diagram. Assets: the uploaded file content itself, the object storage bucket it lands in, the metadata database recording upload and processing state, and the upload API's authentication tokens. Actors: the authenticated end user (legitimate uploader), an external attacker (unauthenticated, or holding a stolen or forged token), and two internal service identities, the upload API and the asynchronous processing worker, each of which holds its own cloud permissions.
flowchart LR
User[Authenticated End User]
Attacker[External Attacker]
subgraph Edge["Edge: public-facing"]
API[Upload API]
end
subgraph Internal["Internal: service network"]
Queue[[Processing Queue]]
Worker[Async Processing Worker]
MetaDB[(Metadata Database)]
end
subgraph Storage["Object Storage"]
Bucket[(Object Storage Bucket)]
end
User -->|upload request plus token| API
Attacker -.->|forged or stolen token| API
API -->|validated file| Bucket
API -->|enqueue job| Queue
Queue --> Worker
Worker -->|reads object| Bucket
Worker -->|writes result metadata| MetaDB
API -->|writes upload record| MetaDB
Threat list (STRIDE walked against the diagram above):
| STRIDE category | Threat | Where |
|---|---|---|
| Spoofing | Attacker uses a stolen or forged token to call the upload API as a legitimate user | Upload API edge boundary |
| Tampering | Uploaded object is modified after storage by an actor with broader-than-intended bucket write access | Object storage bucket |
| Repudiation | A user who uploaded malicious content denies doing so, with no verifiable record tying the upload to their authenticated session | Upload API to metadata database |
| Information Disclosure | Overly broad bucket policy, or an overly long-lived pre-signed URL, exposes stored files to unintended readers | Object storage bucket |
| Denial of Service | An attacker uploads very large files, many small files rapidly, or a decompression-bomb-style file that consumes excessive resources when the worker processes it | Upload API and processing worker |
| Elevation of Privilege | A malicious file exploits a vulnerability in the processing worker's file-handling logic (an image, document, or archive parser), and the worker's cloud identity has broader permissions than the processing task needs, letting the compromise reach other cloud resources | Async processing worker |
Attack path for the highest-priority threat. The Elevation of Privilege path runs: attacker uploads a crafted file that passes the upload API's basic validation (correct declared content type, acceptable size) but is actually built to exploit a parsing vulnerability in whatever library the worker uses to process it (image library, document parser, archive extractor); the worker picks the job off the queue, reads the object, and processing triggers the exploit; if the worker's cloud identity holds permissions beyond what processing strictly requires (for example, broad read/write across all buckets rather than just the one it processes, or permissions to call unrelated cloud application programming interfaces, APIs), the compromised worker process can pivot to reading or modifying data well outside the original upload's scope.
Prioritized mitigations, ranked by the combination of how likely the path is and how much damage it enables:
- Least-privilege identity for the processing worker (addresses Elevation of Privilege, ranked highest because it is the one threat here whose worst case is otherwise unbounded: every other entry on the list has a blast radius confined to one upload, one bucket, or one log record, while a compromised worker holding broad cloud permissions reaches resources that have nothing to do with file uploads at all. Ranking it first is not the same as it being sufficient, and it is worth saying which entries it does not touch: least-privilege scoping on the worker does nothing for Repudiation, nothing for Information Disclosure through an over-long pre-signed URL, and nothing for resource exhaustion, which is why items 2 through 5 are requirements rather than nice-to-haves): scope the worker's cloud identity to only the specific bucket paths and operations processing requires, with no broad cross-bucket or administrative permissions.
- Content validation beyond declared type (addresses Elevation of Privilege and Denial of Service): validate actual file content (magic-byte/content sniffing, not just the client-declared content type or file extension), enforce size limits before the file is fully accepted, and guard against decompression bombs by capping expanded size during any extraction step.
- Short-lived, narrowly scoped upload tokens and pre-signed URLs (addresses Spoofing and Information Disclosure): tokens tied to a specific authenticated session with a short expiry, and any pre-signed URLs generated for reading objects scoped to minutes, not days.
- Bucket policy least privilege plus encryption (addresses Information Disclosure and Tampering): default-deny bucket policy with explicit, narrow grants, and server-side encryption so a misconfigured policy is not the only line of defense.
- Signed, immutable audit logging of upload events (addresses Repudiation): record each upload tied to the authenticated identity and a content hash, in a log the uploading service itself cannot retroactively edit.
Worked example
One example detection control for the highest-priority threat, Elevation of Privilege via the processing worker: alert on any API call made by the processing worker's cloud identity that falls outside its expected, narrow allow-list, most importantly any call touching a bucket other than the one it is scoped to process, or any call to an unrelated service (identity and access management, compute control-plane APIs, and so on). Because the least-privilege mitigation above already constrains what the worker's identity is supposed to be able to do, any call outside that expected set is a strong, low-noise signal, not a fuzzy heuristic: a correctly-behaving worker should never generate one. Concretely, this means shipping the cloud provider's own API audit log (for example, an AWS-style CloudTrail equivalent) for the worker's service identity to a monitoring pipeline with a rule that fires the moment that identity's calls deviate from its documented allow-list, which catches exactly the pivot step in the attack path above (the compromised worker attempting to read or write outside its intended scope) even if the initial exploit itself was never directly observed.
Trade-offs and pitfalls
The most common mistake is validating only the client-declared content type or file extension and treating that as sufficient input validation; an attacker fully controls both of those fields, so real validation has to inspect actual file content. A second is scoping the worker's cloud identity broadly "to avoid permission issues later," which is precisely the choice that turns a contained parsing-library exploit into a cross-resource compromise; least-privilege scoping has real operational cost (more explicit configuration, more friction when the processing logic legitimately needs a new resource) but that cost is the point, since it forces each new permission to be a deliberate decision rather than a default. A third pitfall is treating the DFD, threat list, and mitigation list as one-time deliverables produced once at design time and never revisited; this pipeline's processing logic and dependencies will change, and a new library version or a new processing step reopens the STRIDE walk for at least the elements it touches, not the whole system from scratch, but not nothing either.
Perform a STRIDE analysis for a typical CI/CD pipeline composed of a source repository, CI runners/build agents, artifact registries, deployment service, and production environment. Identify specific attack vectors (compromised credentials, dependency poisoning, malicious build steps), propose mitigations (artifact signing, least-privilege runners, ephemeral agents), and recommend detection mechanisms.
Sample Answer
Direct answer
Walk the pipeline stage by stage (source repository, CI runners/build agents, artifact registry, deployment service, production environment) and apply STRIDE (Spoofing, Tampering, Repudiation, Information disclosure, Denial of service, Elevation of privilege) at each one, because a CI/CD pipeline's real danger is that a single compromised stage inherits the trust of every stage downstream: a build agent that can push a signed artifact is functionally equivalent to a production credential. The dominant threats in practice are compromised credentials (spoofing an identity the pipeline trusts), dependency and build-step poisoning (tampering with what actually gets built), and over-privileged long-lived runners (elevation of privilege), and the mitigations that matter most are artifact signing, least-privilege ephemeral runners, and treating the pipeline itself as a production system with its own threat model, not as tooling exempt from one.
Structured elaboration
STRIDE walk across the five pipeline stages
| STRIDE category | Source repository | CI runners / build agents | Artifact registry | Deployment service | Production environment |
|---|---|---|---|---|---|
| Spoofing | Stolen developer or bot credentials push code as someone else | A malicious PR triggers a runner that assumes the identity of a trusted CI job | A pusher spoofs a trusted publisher identity to upload a look-alike package | Deployment service impersonated to push to the wrong cluster | A compromised pod assumes the service identity of another workload |
| Tampering | Force-push rewrites history; unsigned commits alter approved code post-review | Malicious build step injected into a build script modifies output silently | Dependency poisoning: a compromised or typosquatted upstream package is pulled and baked into the artifact | Deployment manifest altered in transit to deploy a different image than what was approved | Runtime patching of a container after deployment, bypassing the pipeline entirely |
| Repudiation | No signed commits; cannot prove who authored a change | Build logs mutable or not retained; cannot prove what actually ran | No provenance record; cannot prove which build produced which artifact | No immutable deployment record; cannot prove which artifact went where, when | No audit trail tying a running workload back to its deploying pipeline run |
| Information disclosure | Secrets committed to the repo history | Build logs or environment variables leak secrets (API keys, cloud credentials) printed to console | Registry misconfigured for public read access exposes proprietary images | Deployment credentials logged or exposed via a debug endpoint | Environment variables or mounted secrets readable by a compromised sibling container |
| Denial of service | Repository or webhook flooded, blocking legitimate builds | Runner pool exhausted by an attacker submitting expensive or infinite build jobs | Registry quota exhausted or artifact deleted, blocking deploys | Deployment pipeline locked or stuck on a malformed manifest | Resource-exhaustion attack against production from within the cluster |
| Elevation of privilege | A CI configuration file (.yml) change grants the pipeline broader permissions than intended | Compromised credentials: a runner's long-lived cloud credential is stolen and reused outside the pipeline | A build step embeds a supply-chain implant that runs with the registry push permission | Deployment service account has standing production-admin rights it does not need per-deploy | A workload's service account is broader than that single workload requires |
The two attack vectors the question calls out by name, compromised credentials and dependency poisoning, both concentrate at the CI runner stage: credential compromise is fundamentally a Spoofing/Elevation problem (whoever holds the credential is trusted as the pipeline), and dependency poisoning is fundamentally a Tampering problem (the input to the build is not what the developer intended it to be), which is why runner design gets the most mitigation attention below.
Mitigations
- Artifact signing (the question's named mitigation): every build produces a signed attestation binding the artifact's hash to the specific source commit, build definition, and builder identity that produced it (a SLSA-style provenance approach, or simpler code-signing if the org is not ready for full provenance). Deployment refuses to run anything unsigned or whose signature does not match an allow-listed builder identity. This directly closes the Tampering and Repudiation gaps in the table: an altered artifact fails signature verification, and every deployed artifact has a traceable origin.
- Least-privilege runners (the question's named mitigation): a build job gets exactly the permissions its specific build needs (scoped, short-lived tokens minted per-job), never a standing broad cloud credential shared across all jobs. This directly bounds the blast radius of the Elevation-of-privilege row: a compromised job can do only what that job's narrow scope allows, not everything the CI system as a whole can do.
- Ephemeral agents (the question's named mitigation): each build runs on a fresh, disposable runner instance destroyed immediately after the job, rather than a long-lived, reused machine. This closes a persistence path an attacker would otherwise use (implant something on the runner once, benefit from every subsequent build) and limits how long a compromised runner's stolen credential remains valid, since ephemeral credential issuance is naturally tied to ephemeral instance lifetime.
- Container-image-specific mitigations (folding the absorbed CI/CD-for-containers angle): image signing at push time (so the deployment orchestrator, commonly Kubernetes, can enforce an admission-control policy that only accepts signed, provenance-verified images), a dedicated secrets store (a managed secrets manager rather than pipeline environment variables) with short-lived, per-deployment credentials injected at runtime instead of baked into the image or the pipeline config, and scanning images for known-vulnerable base layers before they are allowed to reach the registry.
- Dependency-poisoning-specific mitigations: pin dependencies to exact, hash-verified versions rather than floating ranges, generate and diff a software bill of materials (SBOM, an inventory of every component that went into a build) between builds to catch an unexpected new or changed dependency, and restrict which package registries a build is allowed to pull from.
Detection: signal requirements, not the alerting system
The detection mechanisms the question asks for are best specified as the signals a pipeline stage must emit, handed to whatever system owns alerting and triage; designing that alerting pipeline itself is a separate discipline. The signal requirements this threat model implies:
- A signature-verification failure at deployment time (an artifact that does not match its claimed provenance) must be emitted as a distinct, high-priority signal, not silently retried.
- A build job requesting a broader credential scope than its historical baseline, or reaching an external destination outside an allow-list, must be flagged as anomalous.
- An SBOM diff introducing a dependency not present in the prior N builds must be surfaced before the artifact is promoted, not after.
- Any write to CI configuration files that changes granted permissions must be signaled distinctly from an ordinary code change, since it is a privilege-relevant event even when the diff is small.
Worked example
Trace a single dependency-poisoning incident through the model to show the mitigations actually intercepting it, rather than asserting they would in the abstract:
- An attacker compromises a small, indirectly-depended-upon open-source package and publishes a malicious new version containing a credential-exfiltration payload that runs at install time.
- The CI pipeline's build step (Tampering row, artifact registry column upstream input) pulls the malicious version because the dependency was pinned to a floating range (
^2.3.0) rather than an exact hash. - Without the mitigations above: the payload executes on the build runner with that runner's standing cloud credential (a long-lived, broadly-scoped credential shared across all jobs), exfiltrates it, and the poisoned artifact is signed by the pipeline's own signing key (since signing happens automatically post-build with no provenance check) and deployed to production. Nothing distinguishes this build from a legitimate one in the logs.
- With the mitigations above: the dependency is pinned to an exact hash, so the malicious version is never pulled in the first place; even if it had been, the build runner's credential is short-lived and scoped only to this job's specific artifact-push permission, so exfiltrating it yields the attacker a token that expires within the job's runtime and cannot reach anything outside this one build's narrow scope; and the SBOM diff signal (an unexpected new transitive dependency version) fires before promotion, giving a human a chance to catch it even if the pin had been missed.
This is the concrete version of the direct answer's claim that a compromised stage inherits downstream trust: step 3 shows exactly how a single unpinned dependency plus a single over-broad credential chains all the way to a signed, deployed, production artifact with no distinguishing signal, and step 4 shows each of the three named mitigations breaking a different link in that same chain.
Trade-offs and pitfalls
- Signing without provenance verification is theater. Automatically signing every build output, with no check on what fed into the build, proves an artifact came from "the pipeline" but says nothing about whether the pipeline itself was tricked into building something malicious. Provenance (binding the signature to a specific verified commit and build definition, not just "any successful build") is what actually closes the gap; bare signing alone is a common half-measure.
- Ephemeral runners are not free of cold-start cost. Provisioning a fresh instance per job adds build latency and infrastructure cost versus a warm, reused pool; the trade-off is worth it for anything that touches production credentials or artifact publishing, but a low-risk internal linting job may not need the same treatment. Applying ephemerality uniformly everywhere is a legitimate cost/benefit call, not an automatic best practice.
- Least privilege is easy to state and hard to maintain. Scopes drift toward "broad enough that nothing ever breaks" under delivery pressure unless someone owns periodically re-deriving the minimum scope each job actually uses; treat the per-job permission set as something that needs the same drift review as the rest of the threat model, not a one-time setup task.
- Detection signals without an owner are noise. Specifying signal requirements (as this answer does) is only half the job; each signal needs a named owner and a defined response, or it becomes another unread log line. Confusing "we emit the signal" with "we detect the attack" is the most common gap between a paper threat model and one that actually works in production.
Write a Python script or clear pseudocode that reads a CSV file named 'vulnerabilities.csv' with columns: id, cvss (0.0-10.0), asset_criticality (1-5). Compute a normalized risk score defined as risk = (cvss/10.0) * (asset_criticality/5.0). Output the top 10 vulnerabilities sorted by risk descending, printing id and score. The solution should handle large files without loading everything into memory at once.
Sample Answer
Direct answer
Read the CSV in a streaming fashion (never load the whole file into memory), compute risk = (cvss/10.0) * (asset_criticality/5.0) per row, and maintain only the top 10 by score using a small fixed-size min-heap rather than sorting the full dataset. The heap is what actually delivers the 'without loading everything into memory' requirement; reading line-by-line but then sorting a full in-memory list defeats the purpose.
Structured elaboration
The formula itself is simple (both factors are already 0-1 normalized: cvss/10 and criticality/5), so the interesting engineering decision is the top-k-under-memory-constraint pattern. A bounded min-heap of size 10 gives O(n log 10) time and O(10) space regardless of how large the file is: for each row, if the heap has fewer than 10 items, push it; otherwise compare the new score against the heap's minimum and replace-and-reheapify only if the new score is larger. Python's csv.reader over an open file handle is itself already a streaming iterator (it does not read the whole file at once), so combining it with heapq and never materializing a full list is sufficient to satisfy the memory constraint even on a multi-GB file.
Worked example (executed)
import csv, heapq, io
def top_n_risks(file_obj, n=10):
heap = [] # min-heap of (score, id) so heap[0] is always the current smallest kept
skipped = [] # never skip silently: a dropped row is a vulnerability nobody ranks
reader = csv.DictReader(file_obj)
for row in reader:
try:
cvss = float(row["cvss"])
crit = float(row["asset_criticality"])
except (ValueError, KeyError, TypeError) as exc:
# TypeError is the one people forget: DictReader fills a TRUNCATED row's
# missing column with restval (None), and float(None) raises TypeError,
# not ValueError. Catching only ValueError/KeyError lets a short row kill
# the whole batch, which is exactly the crash this handler exists to stop.
skipped.append((row.get("id"), type(exc).__name__))
continue
score = (cvss / 10.0) * (crit / 5.0)
entry = (score, row["id"])
if len(heap) < n:
heapq.heappush(heap, entry)
elif score > heap[0][0]:
heapq.heapreplace(heap, entry)
return sorted(heap, key=lambda e: -e[0]), skipped
# Simulate a CSV as a text stream (in real use this would be an open() file handle)
csv_text = "id,cvss,asset_criticality\n" + "\n".join(
f"V{i},{cvss},{crit}" for i, (cvss, crit) in enumerate([
(9.8, 5), (7.2, 3), (5.0, 5), (9.1, 4), (3.3, 2), (8.8, 5), (6.0, 1),
(9.9, 5), (2.1, 4), (7.7, 5), (8.0, 3), (4.4, 4), (9.5, 2), (6.6, 5),
])
)
# Two deliberately malformed rows so the skip path is actually exercised, not just written:
# V14 has an empty cvss field (ValueError), V15 is truncated to two columns (TypeError).
csv_text += "\nV14,,4\nV15,8.0"
result, skipped = top_n_risks(io.StringIO(csv_text), n=10)
for score, vid in result:
print(f"{vid}: {round(score, 3)}")
print(f"skipped {len(skipped)} malformed rows: {skipped}")
Actual output (16 input rows, 2 of them deliberately malformed, correctly returns exactly 10):
V7: 0.99
V0: 0.98
V5: 0.88
V9: 0.77
V3: 0.728
V13: 0.66
V2: 0.5
V10: 0.48
V1: 0.432
V12: 0.38
skipped 2 malformed rows: [('V14', 'ValueError'), ('V15', 'TypeError')]
Hand-check on the top and bottom kept rows: V7 is (cvss 9.9, criticality 5): (9.9/10)(5/5) = 0.99, matches. V12 is (cvss 9.5, criticality 2): (9.5/10)(2/5) = 0.95*0.4 = 0.38, matches. With 14 input rows and n=10, exactly 4 rows must be excluded; a full unbounded sort (computed separately as a check on the heap) confirms the 4 excluded are V11 (0.352), V8 (0.168), V4 (0.132), and V6 (0.12), precisely the 4 lowest scores in the input set, which is exactly what a correct bounded top-10 heap should discard. The last output line is the malformed-row path proving it runs rather than merely existing: V14 (empty cvss) raises ValueError and V15 (a truncated two-column row) raises TypeError, and both are skipped and counted instead of killing the batch. That TypeError is the whole reason the row is in the test set. csv.DictReader fills a short row's missing column with None, so float(None) raises TypeError, and a handler that catches only ValueError and KeyError, the obvious pair to reach for, still crashes on the single most common real-world CSV defect.
Trade-offs and pitfalls
A min-heap only helps if n is small and fixed (here 10); for a 'top 5% of a huge file' requirement where the result set itself is large, external sorting or a streaming approximate top-k (e.g., Space-Saving) would be more appropriate. csv.DictReader is convenient but re-parses the header dict on every row; for very high-throughput ingestion, csv.reader with fixed column indices is faster. Always wrap the numeric parse in a try/except and skip rather than crash on one malformed row, and enumerate the exception types against how the parse can actually fail rather than against the two that come to mind: one bad CVSS field should not take down a nightly batch job scanning a million-row file. Count and report the skips too. A silent continue on a wrong-schema file returns an empty top-10 that looks like a clean run with no findings, which is a worse failure than a crash because nobody investigates it.
Draw or describe a Data Flow Diagram (textual description acceptable) for a microservices-based payment system composed of: frontend, API gateway, payment service, order service, user service, third-party payment processor, and database. Identify trust boundaries and list three components that should be considered privileged or high-value targets, explaining why.
Sample Answer
Direct answer
A data flow diagram (DFD) for this payment system needs to show every component the question names (frontend, API gateway, payment service, order service, user service, third-party payment processor, database) plus where trust level changes between them. Of the seven components, the payment service, the database, and the API gateway are the three that deserve privileged-target status, because each of them, if compromised, gives an attacker either money movement, bulk sensitive data, or a foothold that spoofs every request behind it.
Structured elaboration
Data flow diagram
flowchart LR
Client[Customer Browser or App]
GW[API Gateway]
PAY[Payment Service]
ORD[Order Service]
USR[User Service]
DB[(Shared Database)]
TPP[Third-Party Payment Processor]
Client -->|HTTPS request| GW
GW --> PAY
GW --> ORD
GW --> USR
PAY --> DB
ORD --> DB
USR --> DB
PAY -->|signed transaction request| TPP
TPP -->|authorization callback| PAY
subgraph Public["Untrusted: public internet"]
Client
end
subgraph Edge["Semi-trusted: edge, enforces authentication"]
GW
end
subgraph Internal["Trusted: internal service network"]
PAY
ORD
USR
end
subgraph Data["Data tier: shared database, per-service least privilege"]
DB
end
subgraph External["External trust domain: payment partner"]
TPP
end
Trust boundaries (the four points above where trust level actually changes, not just where a network hop happens):
- Customer to API gateway: the public internet meets the first piece of infrastructure the company controls. Nothing here is authenticated yet from the system's point of view; the gateway's job is to establish identity before anything crosses further in.
- API gateway to internal services: the gateway is semi-trusted (it enforces authentication and rate limiting) but the services behind it are fully trusted internally. A request that reaches the payment, order, or user service is implicitly assumed to have already passed the gateway's checks, which is exactly why the gateway itself is a privileged target (see below).
- Internal services to the shared database: a trusted-to-trusted boundary in principle, which is exactly why the diagram above puts the database in its own zone rather than folding it in with the three services. "Trusted" does not mean "uniformly privileged": the order and user services should not need the same database permissions as the payment service, even though all three reach it over the same internal network.
- Payment service to the third-party payment processor: crossing into a separate organization's trust domain entirely. This boundary needs its own authentication (signed requests, mutual TLS, or an equivalent) independent of whatever authenticates traffic inside the company's own network, because the processor cannot rely on the company's internal trust assumptions and vice versa.
Three privileged / high-value targets, and why
- Payment service. It's the only component that both initiates money movement and holds (or brokers access to) payment credentials or tokens. Compromise here doesn't just leak data, it enables direct financial fraud: unauthorized charges, refund abuse, or manipulated transaction amounts. It also sits at the boundary to the external processor, so it's the natural place an attacker would target to pivot from "inside the network" to "able to move money externally."
- Shared database. It aggregates payment records, order history, and user identity data in one place. In a microservices design this is often the component that quietly violates the trust boundaries drawn above, because if payment, order, and user services all write to the same database with similar credentials, a compromise of any one service's database access can read data that belongs to a different service's domain. Its blast radius is the largest of any single component here.
- API gateway. It is the single enforcement point for authentication and authorization for every request into the system. A compromise or misconfiguration here doesn't threaten one data flow, it threatens all of them simultaneously, since every downstream service trusts that a request reaching it has already been authenticated by the gateway.
Worked example
Trace one concrete failure through the diagram: an attacker finds a server-side request forgery bug in the order service that lets them make the order service issue arbitrary internal requests. The order service itself holds no payment data, so on its own this looks low-severity. But because trust boundary 3 above (internal services to database) is not actually enforced with per-service least privilege, the order service's database credentials happen to have read access to the payment table too, a common shortcut when all three services share one database and one connection pool. The attacker uses the order-service bug to query payment records directly. This is exactly the kind of finding a DFD with clearly drawn, and separately verified, trust boundaries is meant to catch before launch: the diagram would prompt the question "does the order service's database credential actually need payment-table access," and the answer is no.
Trade-offs and pitfalls
The most common mistake in this exercise is drawing every network hop as a trust boundary, which produces a diagram with so many boundaries marked that none of them stand out; the four boundaries above are worth marking specifically because trust level changes there, not merely because a network call happens there. A second common mistake is treating "privileged target" as synonymous with "handles the most requests": the API gateway and order service both see more traffic than the payment service, but traffic volume is not the same as blast radius, which is why payment service, database, and gateway rank above order and user service here despite lower call volume. Finally, a shared database across services is a legitimate early-stage architecture choice, but it collapses the trust boundary between services at the data layer even when the service layer looks properly separated; a senior answer flags that collapse explicitly rather than assuming service-level separation implies data-level separation.
Unlock Full Question Bank
Get access to all 20 Threat Modeling and Attack Surface Analysis interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.