Vulnerability Assessment and Management Questions
Finding, prioritizing, and remediating vulnerabilities across systems. Covers vulnerability assessment methodologies, scanning and automation, interpreting and validating scan results, vulnerability classification and scoring (CVSS), prioritization based on exploitability and business impact, and driving remediation to closure. The operational vulnerability-lifecycle discipline, distinct from adversarial penetration testing.
What is a compensating control in vulnerability management? Give concrete examples (network, application, cloud/endpoint) and explain how you'd verify their effectiveness and document them for audit.
Sample Answer
Direct answer
A compensating control is a safeguard that reduces the risk from a vulnerability without actually removing the underlying flaw, used when you cannot patch or fix the root cause right away. It buys time by making the vulnerability harder to reach or harder to exploit, not by eliminating it.
Why they exist
Sometimes a fix is not available yet, no vendor patch exists, or applying it would take unacceptable downtime or break something else, so a compensating control mitigates the risk while the real fix is pending, or where a fix genuinely is not possible on a reasonable timeline. This is different from ignoring the risk: a good compensating control is documented, time-boxed where the plan is to eventually remediate, and periodically re-verified.
Concrete examples by layer
| Layer | Example compensating control | What it does |
|---|---|---|
| Network | Segmentation or firewall rules restricting which systems can even reach the vulnerable service | Reduces the pool of attackers who can reach the flaw at all |
| Application | A web application firewall (WAF) rule blocking the specific exploit pattern | Blocks the known attack payload before it reaches the vulnerable code |
| Cloud or endpoint | Disabling the vulnerable feature or API, or tightening an IAM (identity and access management) policy so only a trusted service account can invoke it | Removes the exposed attack surface without touching the underlying code |
Verifying effectiveness
Actively test that the control blocks the specific exploit technique, not just that it is turned on. A WAF rule that does not actually match the real attack pattern gives false confidence. Re-test periodically, since a firewall rule or WAF configuration can get quietly reverted, and the underlying vulnerability is still there waiting if the control ever lapses. Where possible, run the same detection the original finding came from, a rescan, to confirm the control changes the observed result, not just that a control exists on paper.
Documenting for audit
Record what the underlying vulnerability is, why it cannot be remediated directly right now, exactly what control is in place and how it reduces risk, who approved accepting the residual risk, and a review date. Auditors care most about the last two: that someone with the authority to accept the risk actually signed off, and that it is not an indefinite, unreviewed exception.
Trade-offs and pitfalls
The biggest failure mode is treating a compensating control as a permanent substitute for the real fix; it should always have a plan and a date to actually remediate, even if that date is generous. A control that is not tested against the specific exploit technique can create a false sense of security that is worse than knowing the risk is unmitigated, because it gets deprioritized as "already handled."
Given a finding description, walk through how you'd map it to a CVSS v3.1 vector and base score: state your AV/AC/PR/UI/S/C/I/A choices and justify each.
Sample Answer
Direct answer: Mapping a finding to a CVSS version 3.1 vector means answering eight questions about how the vulnerability is reached and what it does, one per Base metric, then plugging the resulting letters into the published formula. Walking through a concrete finding makes the choices concrete: an unauthenticated file-upload endpoint that lets an attacker execute arbitrary commands on the server.
Structured elaboration, metric by metric:
- Attack Vector (AV): can the attacker reach it over the network with no special positioning? Yes, it's a public web endpoint, so
AV:N(Network). - Attack Complexity (AC): does exploitation require specialized conditions outside the attacker's control (a race condition, a specific unpatched intermediary)? No, uploading a malicious file and triggering it is reliably repeatable, so
AC:L(Low). - Privileges Required (PR): does the attacker need to be logged in first? No, the endpoint is unauthenticated, so
PR:N(None). - User Interaction (UI): does a victim have to click or open something? No, the attacker triggers it directly, so
UI:N(None). - Scope (S): does a successful exploit let the attacker affect resources beyond the vulnerable component itself, crossing a security boundary it doesn't control? Here the compromised web process is itself the whole blast radius (no sandbox escape into a separate security authority), so
S:U(Unchanged). - Confidentiality / Integrity / Availability (C / I / A): arbitrary command execution as the web server's user means the attacker can read data, modify data, and take the service down, so all three are High:
C:H,I:H,A:H.
Worked example: the full vector is AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H. Applying the CVSS version 3.1 Base formula to those values:
Before substituting, here is the letter-to-number lookup that the CVSS version 3.1 specification assigns to each value used here: a High impact (each of C:H, I:H, A:H) = 0.56; Attack Vector Network (AV:N) = 0.85; Attack Complexity Low (AC:L) = 0.77; Privileges Required None (PR:N) = 0.85; and User Interaction None (UI:N) = 0.85. The 6.42 and 8.22 are fixed constants defined by the same specification (6.42 scales the impact term when Scope is Unchanged, and 8.22 scales the exploitability term). With those numbers in hand, every substitution below is just arithmetic.
ISS=1−(1−C)(1−I)(1−A)=1−(1−0.56)(1−0.56)(1−0.56)=1−(0.44)3=0.9148
Impact=6.42×ISS=6.42×0.9148=5.873(Scope unchanged)
Exploitability=8.22×AV×AC×PR×UI=8.22×0.85×0.77×0.85×0.85=3.887
BaseScore=roundup(min(5.873+3.887, 10))=roundup(9.760)=9.8
That 9.8 is the same number you'd get by pasting the vector into any standard CVSS calculator, which is exactly the point of walking through it by hand: it lets you sanity-check a tool's output or defend a manual score in a ticket instead of treating the calculator as a black box.
Trade-offs and pitfalls: the metric people get wrong most often is Scope: it's not asking "does this affect more than one server," it's asking whether the vulnerable component and the impacted component are governed by the same security authority. A container escape that lets an attacker affect the host is a Scope change (S:C); a web app bug that only ever affects that same web app's own data, even across many rows in its database, is not. Getting Scope wrong changes which formula branch you use and can shift the score by a full point or more.
How would you use SBOMs and SCA tooling to prioritize and remediate vulnerabilities in transitive dependencies, including handling version drift and deciding when to remove a library outright?
Sample Answer
Direct answer. For transitive dependencies specifically, a software bill of materials (SBOM) tells you what you depend on several levels deep, and software composition analysis (SCA) tooling tells you which of those have known vulnerabilities, but what you actually do about it, upgrade, override, or remove, hinges on two things a raw vulnerability list doesn't give you: whether the vulnerable code path is reachable from your application, and how much version drift separates you from a safe version.
Structured elaboration
- SBOM and SCA together, plus reachability. The SBOM, ideally including the full transitive tree rather than just direct dependencies, is the map. SCA tooling cross-references every node in that map against vulnerability databases, and increasingly performs call-graph reachability analysis to tell you whether your own code actually invokes the vulnerable function several levels down. This matters a great deal: an unreachable transitive vulnerability at a given severity score is a much lower real priority than a reachable one at the same score, even though a severity-only view would rank them identically.
- Handling version drift. When the vulnerable transitive dependency is several major versions behind what a fix requires, there are generally three paths, in increasing order of maintenance burden: upgrade the direct dependency that pulls it in, if a newer version of that direct dependency also bumps its own transitive pin; force or override the transitive version directly through your package manager's resolution mechanism, accepting the risk that the direct dependency was never tested against that overridden version; or fork or vendor a patched version yourself if no upstream fix exists at all yet. Try the first path first, use the second as a bridge, and reserve the third for something you genuinely cannot wait on.
- When to remove a library outright rather than patch it. Removal is appropriate when the library is unmaintained, with no release addressing the vulnerability and no active maintainer response; when it duplicates functionality your stack already has, such as a small utility pulled in for one function you could inline yourself; or when the attack surface it introduces, a network-facing parser, a deserialization routine, a template engine, is disproportionate to the value it provides. Removal is strictly stronger than patching, because it also eliminates every future vulnerability in that library, not just the current one.
Worked example. Consider two findings at the same severity score, both in the same transitive dependency, reached through one intermediate direct dependency. In the first, reachability analysis shows the application never calls the library's vulnerable deserialization function, only its unrelated date-parsing utility, so the fix is scheduled as a routine, non-emergency upgrade in the next normal release. In the second, the vulnerable function sits directly on the application's live request-handling path, so the same severity score gets remediated the same week. The severity number alone gives no reason to treat these differently; the reachability distinction is what actually separates them.
Trade-offs and pitfalls. Reachability analysis is a real and useful capability in modern tooling, but it isn't perfect: dynamic dispatch or reflection can produce a false "not reachable" result that static analysis can't trace. A "reachability says safe" result should lower a finding's priority, not remove it from the backlog entirely. Removing a library outright is attractive in principle but risks becoming a large, disruptive refactor if treated as the default response to every transitive alert; the criteria above are meant to bound when removal is worth that cost, not to make it the first option considered.
A regulator issues a 30-day remediation mandate for vulnerabilities on a regulated system, but remediating on that timeline would cause unacceptable downtime to revenue-critical services. Draft an executive-level plan that balances compliance and business continuity.
Sample Answer
You don't have to pick only compliance or only uptime here: buy time with compensating controls that measurably reduce risk, formally document a time-boxed exception with the regulator instead of silently missing the deadline, and run the actual remediation on a phased schedule that respects the business's real maintenance windows.
Immediate risk reduction (do this regardless of how the timeline negotiation goes)
Deploy compensating controls that cut exploitability without touching the vulnerable service itself: a WAF (web application firewall) or IPS (intrusion prevention system) rule that virtually patches the specific attack pattern, network segmentation or access restriction that shrinks who can reach the vulnerable component, and stepped-up monitoring so an exploitation attempt gets caught quickly even while the underlying flaw is still present.
Formal documentation and regulator engagement
Don't let the deadline pass silently, since an undocumented miss reads as plain non-compliance. Proactively request a documented extension or risk acceptance: what the vulnerability is, why the 30-day timeline causes unacceptable harm to revenue-critical services, exactly which compensating controls are in place in the meantime, and a firm revised remediation date. Many regulatory frameworks already have a named mechanism for this: a POA&M (Plan of Action and Milestones) in NIST- or FedRAMP-aligned programs, or a documented compensating-controls justification under PCI DSS (Payment Card Industry Data Security Standard). Even without a named mechanism, getting executive and legal sign-off on the decision converts an unexplained gap into a governed, time-boxed exception.
Phased remediation that limits downtime
Patch during a scheduled maintenance window instead of an all-at-once cutover, use a canary or rolling deployment so only a slice of capacity is ever down at once, and sequence the work by exposure: internet-facing and highest-risk instances first, lower-exposure instances later within the extended window.
Executive framing
Present this as a resourced trade-off, not an excuse. Name the business-continuity cost the 30-day timeline would cause, weigh it honestly against the residual risk carried during the extension, and get an explicit executive sign-off on that trade-off, since accepting residual risk on a revenue-critical, regulated system is a business decision, not a call security should make alone.
Worked example
Say the regulated system is a payment-processing cluster where a full patch requires taking the whole cluster offline for several hours, and the business can only safely absorb that during one low-traffic maintenance window per quarter. Rather than force an unsafe cutover inside 30 days, the team deploys a WAF rule blocking the specific request pattern that exploits the flaw within 48 hours, opens a formal exception request citing that mitigation and the added monitoring, and commits in writing to the real patch at the next available low-traffic window, with a hard date attached.
Trade-offs and pitfalls
A compensating control is a mitigation, not a fix: if it's ever misconfigured, unmonitored, or the attack pattern shifts slightly, the underlying vulnerability is still exploitable, so time-box the exception and revisit it rather than letting "temporary" become permanent. Silence is the worst option of all, since a documented, communicated exception is defensible in a way that a quiet missed deadline never is. And because downtime cost and regulatory or reputational exposure aren't measured in the same units, this decision belongs with executives and legal, with security's job being to hand them an accurate picture of the risk, not to make the call unilaterally.
Design a validation harness that automatically re-tests high-priority vulnerabilities after remediation, minimizing risk to production, and integrates with ticketing to close or reopen issues.
Sample Answer
Direct answer
Build the harness around a simple rule: verify high-priority findings automatically wherever it's safe to do so, but keep the riskiest verification techniques (anything resembling active exploitation) out of production entirely, favoring passive, read-only checks there and reserving destructive-style retesting for staging. The harness's job is to produce trustworthy evidence and update the ticket (close or reopen) based on that evidence, never to silently assume success.
Structured elaboration (architecture)
flowchart LR
A[Ticket moves to<br/>Fixed, pending verification] --> B[Trigger]
B --> C[Test-case registry<br/>lookup by finding type]
C --> D{Safety gate:<br/>passive or active check?}
D -->|Passive, safe for prod| E[Execution engine<br/>runs against live asset]
D -->|Active or destructive| F[Execution engine<br/>runs in staging replica only]
E --> G[Evidence captured]
F --> G
G --> H[Decision engine:<br/>match against original finding]
H -->|Confirmed fixed| I[Ticketing API:<br/>close with evidence]
H -->|Still vulnerable| J[Ticketing API:<br/>reopen with evidence]
H -->|Inconclusive/unreachable| K[Flag for manual review,<br/>do not auto-close or auto-reopen]
- Trigger: fires when a ticket transitions to a "fixed, pending verification" state, or on a scheduled poll for anything that's sat in that state too long.
- Test-case registry: for each finding type (mapped by a CVE identifier, a Common Vulnerabilities and Exposures number, a CWE category, Common Weakness Enumeration, or a scanner plugin ID), a pre-approved, non-destructive check: a version or banner check, a safe HTTP probe confirming an expected response change, a read-only configuration assertion via an agent. Production-safe checks favor passive verification (does the version string, response header, or config value now match the patched state) over anything that replays an actual exploit technique.
- Safety gate: routes each check by risk level. Passive, read-only checks can run directly against production. Anything that looks like active exploitation, even a "safe" PoC (proof-of-concept) replay, only ever runs against a staging replica that mirrors the production configuration, never against the live asset.
- Execution engine: a rate-limited job runner with a circuit breaker (halting a batch if error rates spike) and an allow-list of approved check types; no ad hoc or newly-added check type runs without going through the same review the rest of the registry did.
- Decision engine: matches the fresh evidence against the original finding identifier, the same discipline as manual verification (a clean-looking result that doesn't map to the specific original finding isn't proof of anything). Three outcomes: confirmed fixed, still vulnerable, or inconclusive (target unreachable, check timed out, credentials unavailable). Inconclusive results must never auto-close or auto-reopen; they route to a human.
- Ticketing integration: writes the outcome back with the evidence attached, closing with proof or reopening with the reason, always leaving an audit trail rather than a silent status change.
Worked example
A high-priority finding (an outdated TLS library on a public-facing service) moves to "fixed, pending verification." The registry maps this finding type to a passive check: a TLS handshake probe confirming the negotiated protocol version and certificate details reflect the patched configuration. Because this check is read-only and safe for production, the safety gate routes it directly against the live asset. The execution engine runs it, captures the raw handshake evidence, and the decision engine confirms it matches what a successful fix should look like; the ticket closes automatically with that evidence attached. A different, higher-risk finding (an authentication bypass) maps instead to an active check that actually attempts the bypass technique; the safety gate routes that one exclusively to a staging replica, and its result only closes the ticket after a human reviews the staging evidence, since a bypass-style check is never approved to run against production directly regardless of automation.
Trade-offs, failure modes & safety gates
- Flaky checks: a single failed check shouldn't immediately mean "reopen"; use a retry with backoff, and after repeated inconsistent results, mark inconclusive for human review rather than oscillating the ticket open and closed.
- Passive checks can be spoofed or masked: a proxy or load balancer in front of the real service can report a version string that doesn't match what's actually running behind it, so for anything high-stakes, combine more than one type of evidence rather than trusting a single passive signal.
- Full automation for every severity is the wrong target: reserve fully automatic close/reopen for medium and high findings where the evidence is strong and passive, and require a human sign-off step for critical or crown-jewel-asset findings even when the automated check comes back clean, since the cost of a false "fixed" on those is much higher than the cost of a manual review.
- Asset unreachable during a maintenance window should defer the check and retry later, never treated as either "still vulnerable" (fail-closed) or "confirmed fixed" (fail-open) by default.
Unlock Full Question Bank
Get access to all Vulnerability Assessment and Management interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.