Penetration Testing Methodology and Execution Questions
Running structured penetration-testing engagements end to end. Covers the pentest lifecycle, reconnaissance and information gathering, network scanning and enumeration (Nmap, service/version detection), tool selection and usage (Metasploit, Burp Suite), engagement scoping and planning, testing across target types, and findings reporting. The methodical offensive-assessment workflow.
Describe the core categories of penetration testing tools across the testing lifecycle: reconnaissance, vulnerability scanning, exploitation, post-exploitation, lateral movement, traffic analysis, and reporting. For each category provide 2-3 representative tools (open-source and commercial) and one common limitation to be aware of when using tools in that category.
Sample Answer
Direct answer
Across a pentest lifecycle, tooling sorts into seven natural categories, reconnaissance, vulnerability scanning, exploitation, post-exploitation, lateral movement, traffic analysis, and reporting, each with its own representative open-source and commercial tools and its own structural blind spot. The meta-skill an interviewer is actually probing is knowing which category a task belongs to and which tool's limitation you're willing to accept, not memorizing a tool list.
Structured elaboration
| Category | Representative tools | Common limitation |
|---|---|---|
| Reconnaissance | theHarvester, Shodan, Recon-ng | Passive data can be stale (cached DNS records, old scan snapshots) and doesn't reflect the target's current state |
| Vulnerability scanning | Nessus, OpenVAS (Greenbone), Qualys | High false-positive rate without manual verification, and largely blind to custom application logic |
| Exploitation | Metasploit Framework, Cobalt Strike | Well-known, canned modules are increasingly fingerprinted and blocked by modern EDR (Endpoint Detection and Response) tooling |
| Post-exploitation | Mimikatz, BloodHound | Generates significant forensic artifacts (new processes, credential material in memory) that a well-monitored environment may detect even if initial access wasn't |
| Lateral movement | Impacket, Evil-WinRM | Relies on legitimate administrative protocols, so a well-instrumented network can flag unusual host-to-host traffic even though the protocol itself is normal |
| Traffic analysis | Wireshark, tcpdump | Encrypted (TLS) traffic limits visibility unless you're also positioned to decrypt it, such as through an authorized intercepting proxy |
| Reporting | Dradis, PlexTrac | Organizes evidence and templates the structure, but doesn't write the risk narrative or judge severity for you |
Worked example
On a single internal engagement you might move theHarvester output into your target list, use Nessus to identify a credentialed, confirmed vulnerability on a host, exploit it with Metasploit to gain initial access, use Mimikatz and BloodHound to harvest credentials and map a path to a more privileged account, use Impacket to move onto that next host with the harvested credentials, watch the traffic with Wireshark to confirm nothing unexpected is happening on the wire, and log every command into Dradis throughout so the final report writes itself from already-organized evidence.
Trade-offs and pitfalls
- Defaulting to the flashiest tool in a category, an exploitation framework, before doing the cheaper reconnaissance and scanning work first wastes engagement time and is usually a sign of checklist-driven testing rather than judgment.
- Commercial tools often justify their cost through better false-positive tuning, support, or reporting integration, not through fundamentally different technical capability; know when a client's budget or existing tooling makes the open-source equivalent the better choice.
- A tool's limitation is not a reason to skip the category, it's a reason to plan a manual step around it, such as verifying a scanner finding by hand.
Scenario: you're asked to perform a penetration test inside an Industrial Control Systems (ICS) environment that supports manufacturing and cannot tolerate downtime. The scope includes HMIs, PLCs, and an isolated engineering network. Propose a phased testing plan and tool selection that minimizes risk to control systems while delivering actionable findings for OT security teams.
Sample Answer
A penetration test against an Industrial Control System (ICS), the computers and physical equipment that run a factory floor, must treat availability and physical safety as the overriding constraint, ahead of coverage. The right shape is a passive-first, safety-gated phased plan: build the asset picture without sending unsolicited packets to fragile devices, widen to narrow rate-limited active testing only inside an approved maintenance window with an operational technology (OT) engineer on standby, and reserve any live exploitation for a lab twin of the actual hardware rather than the production line.
Phased plan
-
Pre-engagement and safety scoping. Sit down with the plant engineering team before writing the scope document. Carve out the Safety Instrumented System (SIS), the machinery whose only job is shutting the plant down safely in an emergency (for example an emergency shutdown valve controller), as fully off limits: a false trip there can itself cause a safety incident. Agree on a maintenance window for anything beyond passive listening, name a single escalation contact who can physically walk to the equipment, and write down the abort condition in advance: any alarm, unexpected relay click, Human-Machine Interface (HMI, the touchscreen or workstation an operator uses to watch and adjust the process) freeze, or comms drop halts testing immediately and unconditionally.
-
Passive asset discovery, zero active packets to OT devices. Use a SPAN/mirror port or a network tap to copy traffic without touching the OT switches' forwarding path, and build the inventory purely from what devices say on their own: IP/MAC, protocol, apparent role. Identify the industrial protocols in play, for example Modbus (a decades-old protocol most Programmable Logic Controllers, PLCs, speak, with no authentication built in at all) or DNP3, and separate HMIs from PLCs (small ruggedized computers wired directly to sensors and actuators like a valve or motor).
-
Narrow, low-rate active testing, only in the approved window. Scan one host at a time at a deliberately slow rate, not a default fast scan timing template, because many embedded PLC network stacks have shipped for a decade or more and are known in the field to crash or hang on malformed or bursty traffic that a modern server operating system would shrug off. Prefer read-only protocol queries (reading a Modbus holding register) over anything that writes, and document that a write function code is reachable and unauthenticated without ever sending an actual mutating write to production equipment.
-
IT/OT boundary and Windows-based assets. Test the firewall, data diode, or jump host separating corporate IT from OT with full normal-strength technique: this is the one place typical pentest posture is appropriate, since it is exactly the path a real attacker would use to reach the fragile side. HMIs and engineering workstations usually run standard Windows, so apply ordinary host and patch-level testing there; this is often the fastest real path to control-plane compromise.
-
Exploitation only against a twin, never the live line. If a critical finding needs proof of exploitability rather than just "reachable and unauthenticated," do it on an offline test rig or vendor-provided digital twin running the same PLC model and firmware. Without a twin, stop at documented evidence and let that stand as proof of risk.
-
Reporting for an OT audience. Pair every finding with a plant-language impact statement (what physical process could be disrupted or made unsafe), not only a technical severity score, because the budget sign-off often sits with a plant manager, not a security engineer. Acknowledge that OT patch cycles are frequently measured in years due to vendor certification and change control, so compensating controls (segmentation, monitoring, boundary allow-listing) usually matter more than "patch the PLC."
Worked example
Say the passive capture from step 2 shows a PLC at an internal address responding to Modbus TCP on port 502 with zero authentication, and the firewall rule review from step 4 shows the corporate IT subnet can reach that port directly. That is a complete, reportable finding built entirely from passive listening and firewall-config review: unauthenticated protocol, writable function codes present, IT-to-OT reachability confirmed. The impact statement ("a compromised IT host can issue write commands to a controller on the production line") and remediation (segment the boundary, restrict the rule to a jump host) never required sending a single packet at the PLC itself.
Trade-offs and pitfalls
The biggest pitfall is applying ordinary IT pentest instincts, an aggressive default scan or an automatic exploitation framework, to OT devices; documented ICS incidents in the field involve PLCs and HMIs crashing from ordinary vulnerability scans, not just deliberate exploit attempts. There is a genuine tension between proving exploitability, which makes a stronger report, and the outage risk of doing so on a live line; the senior judgment call is knowing when documented-but-unexploited evidence is the more responsible answer, and saying so explicitly in the report instead of silently under-testing. Segmentation testing at the IT/OT boundary is usually the highest-value, lowest-risk work on this kind of engagement and deserves more time than the instinct to "get into" the OT segment itself would suggest.
Design an automated penetration test harness for API gateways and WAF rules that performs black box testing. The harness should: enumerate endpoints, send a curated set of malicious payloads and evasions, record which requests were blocked or allowed, and measure rule coverage and false positive rates. Describe how to automate repeated runs safely and how to use results to tune WAF policies.
Sample Answer
Direct answer
An automated black-box harness for testing API gateway and WAF (Web Application Firewall, the filtering layer in front of the app that inspects incoming requests and blocks the ones matching known attack patterns) rules needs three things a manual test rarely bothers to formalize: a labeled payload corpus (a curated, labeled collection of test payloads) organized by attack class with matching evasion variants, a way to reliably tell a WAF block apart from a normal application response, and a second, separate corpus of legitimate-looking requests to measure false positives, since coverage without a false-positive check just rewards a WAF tuned to block everything.
Design
Enumerate endpoints. Seed the target list from the OpenAPI specification when one exists, the same reasoning as fuzzing any other API: it is ground truth for what endpoints and methods exist. Supplement it with a passive crawl or proxy capture of real traffic through the gateway to catch undocumented routes the spec omits, and re-sync this target list on every run rather than treating it as a one-time snapshot, since endpoints change.
Curated payloads and evasions. Organize a labeled corpus by the injection class it targets: boolean, error-based, and time-based SQL injection probes, reflected cross-site scripting markers, path traversal sequences, command injection canaries, deserialization markers, using well-known, non-destructive detection strings for each class, the same "detect, do not exploit" probes used in manual testing, such as a boolean SQL injection pair like ' OR '1'='1 against ' OR '1'='2 to detect a behavioral difference, or a benign reflected-XSS marker. Pair each canonical payload with WAF-evasion variants, double URL-encoding, case variation, comment insertion, Unicode normalization tricks, since the specific purpose of this harness is measuring whether the WAF catches the obfuscated form, not just the textbook payload a WAF vendor's default rule set was written against.
Recording blocked versus allowed. Send each payload through the actual gateway and WAF path, never directly to the backend, and classify the response: a WAF block usually looks distinct (its own error page, a header such as a block-reason identifier, or a connection reset), while an allowed request either reaches the backend, confirmed through a backend-side correlation marker the harness can check, or is rejected by the backend itself for an unrelated reason, which must be distinguished from a WAF block or the harness will undercount real bypasses. Log the full request, the matched rule identifier if the WAF exposes one, and the classification for every payload.
Measuring rule coverage and false-positive rate. Coverage is the fraction of the malicious-payload corpus that gets blocked, reported per injection class so you learn, for example, that SQL injection coverage is strong while coverage for server-side request forgery (SSRF, where the server itself is tricked into making an attacker-directed request) is weak, rather than one blended number that hides the gap. False-positive rate needs a separate corpus of known-benign-but-suspicious-looking requests, a legitimate user bio field containing the word "select," a legitimate query string with several ampersands, run through the same harness; a well-tuned WAF blocks most of the malicious corpus and passes most of the benign corpus, and both numbers should be reported together, since a WAF tuned to maximize coverage alone can trivially do so by blocking almost everything, which destroys usability.
Automating repeated runs safely. Rate-limit the harness's own traffic well below anything that could become a denial-of-service against the gateway itself or trip unrelated infrastructure alarms. Run against a dedicated staging instance of the gateway and WAF configuration wherever possible, since firing thousands of malicious-looking requests at a shared production WAF risks polluting production security logs and consuming the WAF's own CPU or rate-limit budget. Version the payload corpus and the WAF rule set together so a scheduled run's results are diffable against the prior run, did coverage regress after a rule was "simplified," did a new false positive appear after a rule was tightened, rather than being a one-off snapshot with nothing to compare it to.
Using results to tune policy. Feed per-class coverage gaps back to whoever owns the WAF rule set as a prioritized list, the class with the lowest coverage and highest business risk first, and feed false positives back as concrete before-and-after examples so the rule owner can narrow a match pattern rather than being told only "false positives exist." Treat the labeled corpus itself as a regression suite for the WAF configuration going forward: every rule change gets run against it before deployment, turning WAF tuning from an ad hoc, reactive process into something with its own test suite.
Trade-offs & pitfalls
A harness that only reports coverage, with no benign corpus and no false-positive number, will systematically push a WAF owner toward over-blocking, since coverage is the only metric being rewarded. It is also easy to under-invest in the evasion-variant half of the corpus and end up only proving the WAF's default rule set works, which the vendor already tested; the evasion variants are what make this exercise worth automating in the first place.
You need to test for reflected XSS with Burp. Outline how you'll find injection points using passive and active techniques, how to craft payloads for different contexts (HTML body, attribute, JS literal, URL), how to use Repeater to confirm, and how to demonstrate impact to stakeholders. Mention DOM XSS differences and how Burp can help detect them.
Sample Answer
I find candidate injection points passively by watching what the app reflects back unmodified as I browse it normally, then confirm and refine with Burp Repeater using a harmless, unique marker before ever sending a real payload, and I always tailor the payload's syntax to the exact spot in the page where my input lands.
Discovery: passive then active
- Passive: proxy all traffic through Burp, and watch for any parameter whose value reappears verbatim in the HTML body, an attribute, or inline JavaScript.
- Active: fuzz parameters with a distinctive marker string (something like
zzTESTzz) and search responses for that marker, noting whether it comes back encoded or raw.
Context-specific payloads
- HTML body:
<script>alert(1)</script>, or<img src=x onerror=alert(1)>if script tags are filtered. - HTML attribute: break out of the attribute first, for example
" onmouseover=alert(1) x=". To trace it: if your input lands inside<input value="HERE">, the leading"closes thevalueattribute,onmouseover=alert(1)adds an event handler that runs when the mouse moves over the field, and the trailingx="opens a harmless leftover attribute so the tag stays valid HTML. - JavaScript string literal: close the string and add a statement, for example
');alert(1);//. To trace it: if your input lands inside a call likesearch('HERE'), the leading')closes the string and the function call,;alert(1);runs as its own statement, and the trailing//comments out the leftover')so the script does not throw a syntax error. - URL or href context:
javascript:alert(1), or an attribute breakout like"><svg onload=alert(1)>.
Confirming with Repeater
Send the crafted request in Repeater and inspect the raw response for the payload appearing unencoded exactly where expected; if it's being encoded or filtered, adjust case or encoding and retry before concluding the point isn't exploitable.
Demonstrating impact
A benign alert(1) popup, captured with a screenshot alongside the exact request, is the conventional proof for cross-site scripting (XSS); the point is proving script execution happened, not building a full attack chain against a real user.
DOM XSS versus reflected XSS
Reflected XSS round-trips through the server; the payload appears somewhere in the HTTP response Burp can see. Document object model (DOM) XSS never touches the server at all; it's entirely client-side JavaScript writing attacker-controlled input into a dangerous sink like innerHTML. Burp's passive scanner and its DOM Invader browser extension can flag likely DOM sinks, but confirming a DOM XSS finding means reading the page's client-side JavaScript, not just watching HTTP traffic.
Worked example
A search box reflects my marker zzTESTzz unencoded inside <div>results for zzTESTzz</div>. I replace it with <script>alert(1)</script> in Repeater; the raw response shows the tag unencoded, and loading that request in a browser pops the alert, confirming HTML-body-context reflected XSS.
Trade-offs and pitfalls
A common pitfall is trying a <script> payload everywhere regardless of context, missing that an attribute or JavaScript-literal context needs a different breakout syntax entirely. Another is confusing "the marker reflected" with "the payload executed," since output encoding can make a parameter reflect visibly while still being completely safe.
Write a Burp extension design (high-level, in Java or Python) that automatically detects and reports insecure CORS configurations (wildcard origins, wildcard with credentials, overly-permissive Access-Control-Allow-Origin). Include detection logic, confidence scoring, and how you would present remediation steps. Explain how you'd test and validate the extension across multiple sites.
Sample Answer
Direct answer
Design this as a Java extension on Burp's current Montoya API: a passive HttpHandler that inspects every response's CORS headers against the request's own Origin header for cheap, no-traffic-added detection, backed by an active follow-up probe (resending the same request with two unrelated, attacker-controlled Origin values) that upgrades a passive guess into a confirmed finding, which is what makes the confidence scoring meaningful rather than cosmetic.
Detection logic
By default a browser enforces the same-origin policy: JavaScript running on one site cannot read the responses another site returns, which is what stops a random page you visit from quietly reading your logged-in webmail. Cross-Origin Resource Sharing (CORS) is the mechanism a server uses to deliberately opt specific other origins out of that block, by sending Access-Control-Allow-* response headers. A credentialed cross-origin request is one that carries the victim's own cookies along with it, so if the server wrongly opts an attacker's origin in for credentialed requests, the attacker's page can read the victim's authenticated responses. With that baseline in place, the detection logic below explains itself.
The dangerous CORS pattern is not simply Access-Control-Allow-Origin: *. A literal wildcard combined with Access-Control-Allow-Credentials: true is actually rejected by browsers at the fetch-spec level for credentialed requests, so that exact combination is a misconfiguration worth flagging as low-severity noise, not the finding that matters most. The genuinely exploitable pattern is a server that reflects whatever Origin header the request sent back as the value of Access-Control-Allow-Origin, which passes the browser's same-origin check because it is not literally *, it exactly matches the requesting origin, while combined with Access-Control-Allow-Credentials: true this lets any origin, including an attacker's, make a credentialed cross-origin request and read the response.
public class CorsAuditExtension implements BurpExtension {
private MontoyaApi api;
@Override
public void initialize(MontoyaApi api) {
this.api = api;
api.extension().setName("CORS Misconfiguration Auditor");
api.http().registerHttpHandler(new CorsHandler(api));
}
}
class CorsHandler implements HttpHandler {
private final MontoyaApi api;
CorsHandler(MontoyaApi api) {
this.api = api;
}
@Override
public RequestToBeSentAction handleHttpRequestToBeSent(HttpRequestToBeSent request) {
return RequestToBeSentAction.continueWith(request);
}
@Override
public ResponseReceivedAction handleHttpResponseReceived(HttpResponseReceived response) {
HttpRequest request = response.initiatingRequest();
String requestOrigin = request.headerValue("Origin");
String acao = response.headerValue("Access-Control-Allow-Origin");
String acac = response.headerValue("Access-Control-Allow-Credentials");
if (requestOrigin == null || acao == null) {
return ResponseReceivedAction.continueWith(response);
}
boolean reflectsOrigin = acao.equalsIgnoreCase(requestOrigin);
boolean allowsCredentials = "true".equalsIgnoreCase(acac);
boolean isWildcard = "*".equals(acao);
if (isWildcard && allowsCredentials) {
report(response, "Wildcard origin combined with credentials",
AuditIssueSeverity.INFORMATION, AuditIssueConfidence.CERTAIN,
"Server sends Access-Control-Allow-Origin: * together with "
+ "Access-Control-Allow-Credentials: true, a combination browsers reject "
+ "for credentialed requests; usually a sign of a misunderstood CORS config "
+ "rather than an exploitable bypass on its own.");
} else if (reflectsOrigin && allowsCredentials && !isWildcard) {
AuditIssueConfidence confidence = confirmReflection(request, requestOrigin);
report(response, "Reflected-origin CORS with credentials enabled",
AuditIssueSeverity.HIGH, confidence,
"Server reflects the request's Origin header back as "
+ "Access-Control-Allow-Origin and allows credentials, letting any origin "
+ "read authenticated responses. Confirmed via active differential probing.");
}
return ResponseReceivedAction.continueWith(response);
}
private AuditIssueConfidence confirmReflection(HttpRequest original, String observedOrigin) {
String probeOriginA = "https://cors-probe-a.invalid";
String probeOriginB = "https://cors-probe-b.invalid";
HttpRequestResponse resultA = api.http().sendRequest(
original.withUpdatedHeader("Origin", probeOriginA));
HttpRequestResponse resultB = api.http().sendRequest(
original.withUpdatedHeader("Origin", probeOriginB));
String acaoA = resultA.response().headerValue("Access-Control-Allow-Origin");
String acaoB = resultB.response().headerValue("Access-Control-Allow-Origin");
boolean reflectsA = probeOriginA.equalsIgnoreCase(acaoA);
boolean reflectsB = probeOriginB.equalsIgnoreCase(acaoB);
if (reflectsA && reflectsB) {
return AuditIssueConfidence.CERTAIN;
} else if (reflectsA || reflectsB) {
return AuditIssueConfidence.FIRM;
}
return AuditIssueConfidence.TENTATIVE;
}
private void report(HttpResponseReceived response, String name,
AuditIssueSeverity severity, AuditIssueConfidence confidence, String detail) {
AuditIssue issue = AuditIssue.auditIssue(
name, detail,
"Replace origin reflection with a strict server-side allow-list of known "
+ "origins; never combine a wildcard or reflected origin with "
+ "Access-Control-Allow-Credentials: true.",
response.initiatingRequest().url(),
severity, confidence,
"Cross-Origin Resource Sharing misconfigurations can let an attacker-controlled "
+ "site read authenticated responses on behalf of a victim's browser.",
"Maintain an explicit allow-list of trusted origins server-side and echo back only "
+ "an exact match from that list, never the raw incoming Origin header.",
AuditIssueSeverity.HIGH,
response
);
api.siteMap().add(issue);
}
}
Confidence scoring
The two-phase design is what makes the confidence levels honest rather than arbitrary. A single passive observation, one response reflecting one Origin, is TENTATIVE, since a small, legitimate allow-list could coincidentally include that one origin. Resending the request with two unrelated, synthetic origins (cors-probe-a.invalid and cors-probe-b.invalid, domains no real allow-list would ever contain) and observing that the server reflects one of them raises confidence to FIRM; reflecting both is CERTAIN, since no legitimate, deliberately-configured allow-list would ever include two unrelated placeholder domains.
Remediation presentation
Each AuditIssue carries both an instance-specific detail (the actual origin observed being reflected on this endpoint) and a general remediationBackground describing the fix pattern once: replace origin reflection with a fixed, server-side allow-list of known trusted origins, and never combine a wildcard or reflected origin with Access-Control-Allow-Credentials: true. Surfacing findings through siteMap().add(...) puts them directly in Burp's normal issue list, next to the built-in scanner's findings, rather than a separate custom output panel a reviewer has to remember to check.
Testing and validation across multiple sites
Before trusting the extension's output on a real engagement, run it against a small internal test matrix: at least one service with a genuinely strict, small allow-list (to confirm no false positive fires), one deliberately misconfigured service reflecting any origin with credentials enabled (to confirm CERTAIN fires correctly), and one service serving only Access-Control-Allow-Origin: * with no credentials (to confirm it is scored as low-severity, not flagged the same as the credentialed case). Cross-check a sample of the extension's FIRM and CERTAIN findings by hand in Repeater, sending the same synthetic-origin probe manually, to confirm the tool's classification is not systematically miscalibrated before relying on it across a full engagement.
Complexity and edge cases
The passive path is O(1) per response, a handful of header comparisons, and only escalates to the two extra active probe requests when a passive candidate is actually found, so it does not meaningfully add to overall proxy traffic. Edge cases worth handling explicitly: requests with no Origin header at all (same-origin requests never send one, and should be skipped rather than treated as a non-match); multiple Access-Control-Allow-Origin headers on one response (a specification violation in itself, worth its own low-confidence flag); case-sensitivity in header name lookups (headerValue should already be case-insensitive per the API, but origin VALUE comparison should stay exact-match rather than case-insensitive, since scheme and host casing differences are meaningful); and non-credentialed endpoints returning genuinely public data behind a wildcard, which is only a real problem if the response also contains something sensitive despite requiring no cookies, so severity, not just presence, still needs the analyst's judgment.
Trade-offs & pitfalls
The active-probe phase doubles the request count for every passive candidate found, so on a very chatty API this needs its own rate limit to avoid becoming a mini load-test of the target. It is also worth remembering this design only catches cookie or session-based CORS abuse; an API that authenticates purely via a bearer token in a custom header rather than cookies is not exploitable through CORS in the same way, since a cross-origin script cannot read or set that header without already having the token, so the tool's findings should be read alongside how the endpoint actually authenticates, not in isolation.
Unlock Full Question Bank
Get access to all Penetration Testing Methodology and Execution interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.