Penetration Testing Methodology and Execution Questions
Running structured penetration-testing engagements end to end. Covers the pentest lifecycle, reconnaissance and information gathering, network scanning and enumeration (Nmap, service/version detection), tool selection and usage (Metasploit, Burp Suite), engagement scoping and planning, testing across target types, and findings reporting. The methodical offensive-assessment workflow.
What's the difference between automated vulnerability scanning and manual penetration testing? For each, describe its strengths, weaknesses, and typical deliverables, and explain how the two complement each other in a real security program.
Sample Answer
Direct answer
Automated vulnerability scanning runs signature and configuration checks against a target at scale and speed, producing a list of known, pattern-matched issues; it's cheap, repeatable, and good at breadth. Manual penetration testing puts a human in the loop who chains findings, applies business context, and looks for logic flaws no signature can describe; it's expensive and time-boxed but finds the things that actually cause a breach. A mature security program runs both: continuous or frequent automated scanning for baseline coverage, and periodic manual testing that validates the scanner's output and goes after what it structurally cannot see.
Structured elaboration
| Automated vulnerability scanning | Manual penetration testing | |
|---|---|---|
| Strengths | Fast, cheap, repeatable, covers a large asset inventory continuously, strong at known-CVE (Common Vulnerabilities and Exposures) and misconfiguration coverage | Context-aware, chains low- and medium-severity issues into a real attack path, finds business-logic and authorization flaws, proves exploitability rather than guessing at it |
| Weaknesses | High false-positive rate without manual triage, blind to custom application logic and multi-step authorization flaws, limited to what's already a known signature | Expensive, a snapshot of a single point in time, quality depends heavily on the individual tester's skill, doesn't scale to daily or weekly coverage of a large environment |
| Typical deliverable | A scored, often noisy list of findings mapped to CVEs or a compliance checklist | A narrative report showing proven exploit chains, business impact, and prioritized remediation guidance |
How they complement each other in a real program: scanning gives continuous, broad coverage across an entire fleet of assets at a cadence manual testing could never afford, weekly or even daily, which surfaces known issues quickly and cheaply. Manual pentesting is scheduled periodically (commonly quarterly external, annual internal, or triggered by a major release) and does two things scanning cannot: it triages the scanner's findings to separate real exploitable issues from noise, and it spends focused human time hunting for the class of flaw a scanner is structurally blind to, like an authorization check that's present on one endpoint but missing on a near-identical one.
Worked example
A scanner flags a server as running an outdated TLS (Transport Layer Security) library version with a critical CVE. A human tester manually verifies that the specific vulnerable code path is actually reachable from the internet-facing interface before treating it as a real critical risk, since some scanners flag a version string without confirming the vulnerable feature is even enabled. Separately, that same manual test finds a "view invoice" endpoint that accepts an arbitrary invoice ID from any authenticated user with no ownership check, a purely logic-driven authorization flaw no vulnerability scanner's signature database would catch, because there's no known "bad" pattern to match, just a missing business rule.
Trade-offs and pitfalls
- Treating scanner output as a report deliverable on its own, common in immature programs, drowns real risk in false positives and misses the logic flaws that actually cause incidents.
- Treating manual pentesting as a substitute for continuous scanning wastes expensive human time re-finding issues a much cheaper automated tool would have caught between engagements.
- The two are sequenced, not interchangeable: scanning first for breadth and triage, manual testing second for depth on what matters most.
Describe how you would structure a 20-minute executive briefing after a penetration test. Include slide topics and time allocation (e.g., summary, top risks, remediation roadmap, cost/impact), what material to present verbally versus in appendices, and an approach for handling difficult executive questions about legal exposure or remediation cost estimates.
Sample Answer
I treat 20 minutes as a fixed budget spent almost entirely on the three to five findings that actually change a business decision, push technical depth into an appendix, and rehearse the two or three hard questions I already know are coming, so I'm not improvising the organization's risk posture live in the room.
Time allocation
| Segment | Time | Content |
|---|---|---|
| Opening and context | 1-2 min | Scope, test window, objective, one-line risk posture |
| Executive summary | 3-4 min | Top 3-5 findings by business impact, a likelihood-times-impact framing, headline risk rating |
| Top risks, deep dive | 6-7 min | One slide per top finding: plain-language description, attack path in one sentence, business consequence |
| Remediation roadmap and cost/impact | 4-5 min | Quick wins versus structural fixes, rough effort tiers, expected risk reduction |
| Compliance and legal note (if relevant) | 1-2 min | Factual regulatory exposure only, no legal interpretation |
| Q&A and next steps | 2-3 min | Owners assigned per finding, retest date |
What's verbal versus in the appendix
The exploit-chain narrative and the business-impact framing get said out loud, in plain language. Raw evidence, full CVSS breakdowns, request/response dumps, and detailed remediation task lists live in an appendix the executives can hand straight to their engineering leads afterward, rather than being read aloud to a room that doesn't need that level of detail.
Handling difficult executive questions
- Legal exposure: stay strictly factual ("we demonstrated X was possible under Y conditions"), and explicitly defer any interpretation of liability to the client's counsel rather than characterizing exposure yourself.
- Remediation cost: give effort tiers (small, medium, large) grounded in what the finding actually requires to fix, rather than fabricating a precise dollar figure, and offer to loop in engineering leadership for a firmer estimate after the meeting.
Worked example
For an unauthenticated admin endpoint, the verbal line is something like: "This lets anyone on the internet who finds this URL take over the admin account without a password; it took my team about ten minutes to find." The appendix carries the full request/response evidence and the exact severity breakdown, but only the plain-language version ever reaches the slide.
Trade-offs and pitfalls
A common pitfall is spending the first ten minutes walking through methodology and tooling that executives neither need nor remember, leaving too little time for the findings that actually matter to them. Another is putting a full technical severity score or vector string directly on an executive slide; it reads as noise to that audience and belongs in the appendix instead.
Identify the legal and compliance notices and statements that should appear in a penetration test report. For each item, explain why it's important and provide a short sample phrasing suitable for inclusion in the report.
Sample Answer
Direct answer: A penetration test report needs a small set of standard front-matter statements that protect both the client and the testing firm legally and set expectations about what the report is and isn't. Missing them is a common gap that undermines an otherwise strong technical report.
Structured elaboration
- Authorization and scope statement. Confirms the test was authorized, states the dates and systems tested, and protects the tester from being mistaken for an actual attacker. Sample: "This assessment was conducted under signed authorization from [Client], limited to the assets and time window defined in the Rules of Engagement; testing outside that scope was not performed and is not covered by this report."
- Confidentiality notice. The report contains exploitable vulnerability detail, so it states who may see it. Sample: "This document is confidential and intended solely for [Client]. It must not be distributed outside authorized personnel."
- Point-in-time disclaimer. A pentest is a snapshot, not an ongoing guarantee, which manages expectations if a new vulnerability appears the next day. Sample: "Findings reflect the environment as tested between [dates]. This report does not guarantee the absence of vulnerabilities introduced after the assessment period."
- Methodology reference. States the framework followed (for example, the Penetration Testing Execution Standard, PTES, or NIST Special Publication 800-115) so the reader can judge rigor. Sample: "This engagement followed a methodology aligned with PTES, covering reconnaissance, scanning, exploitation, and reporting."
- Evidence handling and retention statement. States what happens to captured evidence after delivery. Sample: "All evidence and test artifacts, including any data samples extracted to demonstrate impact, will be securely destroyed 30 days after report acceptance."
- Ownership notice. States who owns the report content. Sample: "This report is the property of [Client] upon delivery."
- Classification marking. A short marking (often the Traffic Light Protocol, TLP, a simple color-coded sharing standard) so recipients know how far they may forward it. Sample: "TLP:AMBER, limited disclosure, restricted to participants' organizations."
Worked example. Together, these appear as a short block on the report's cover or first page: a scope-and-authorization paragraph, a confidentiality line, the point-in-time disclaimer, the methodology reference, and the classification marking, all before the executive summary begins.
Trade-offs and pitfalls. Don't copy a generic disclaimer template without checking it against the actual signed contract for this engagement; a mismatch in dates or scope undermines the protection the notice is meant to provide. Don't overload the front matter with dense legal language duplicating the separately-signed contract; keep it short and let the contract carry the full detail. The evidence-retention statement is the most commonly forgotten item, and it's the one most likely to matter if a client later asks what happened to the data you extracted.
Design an automated penetration test harness for API gateways and WAF rules that performs black box testing. The harness should: enumerate endpoints, send a curated set of malicious payloads and evasions, record which requests were blocked or allowed, and measure rule coverage and false positive rates. Describe how to automate repeated runs safely and how to use results to tune WAF policies.
Sample Answer
Direct answer
An automated black-box harness for testing API gateway and WAF (Web Application Firewall, the filtering layer in front of the app that inspects incoming requests and blocks the ones matching known attack patterns) rules needs three things a manual test rarely bothers to formalize: a labeled payload corpus (a curated, labeled collection of test payloads) organized by attack class with matching evasion variants, a way to reliably tell a WAF block apart from a normal application response, and a second, separate corpus of legitimate-looking requests to measure false positives, since coverage without a false-positive check just rewards a WAF tuned to block everything.
Design
Enumerate endpoints. Seed the target list from the OpenAPI specification when one exists, the same reasoning as fuzzing any other API: it is ground truth for what endpoints and methods exist. Supplement it with a passive crawl or proxy capture of real traffic through the gateway to catch undocumented routes the spec omits, and re-sync this target list on every run rather than treating it as a one-time snapshot, since endpoints change.
Curated payloads and evasions. Organize a labeled corpus by the injection class it targets: boolean, error-based, and time-based SQL injection probes, reflected cross-site scripting markers, path traversal sequences, command injection canaries, deserialization markers, using well-known, non-destructive detection strings for each class, the same "detect, do not exploit" probes used in manual testing, such as a boolean SQL injection pair like ' OR '1'='1 against ' OR '1'='2 to detect a behavioral difference, or a benign reflected-XSS marker. Pair each canonical payload with WAF-evasion variants, double URL-encoding, case variation, comment insertion, Unicode normalization tricks, since the specific purpose of this harness is measuring whether the WAF catches the obfuscated form, not just the textbook payload a WAF vendor's default rule set was written against.
Recording blocked versus allowed. Send each payload through the actual gateway and WAF path, never directly to the backend, and classify the response: a WAF block usually looks distinct (its own error page, a header such as a block-reason identifier, or a connection reset), while an allowed request either reaches the backend, confirmed through a backend-side correlation marker the harness can check, or is rejected by the backend itself for an unrelated reason, which must be distinguished from a WAF block or the harness will undercount real bypasses. Log the full request, the matched rule identifier if the WAF exposes one, and the classification for every payload.
Measuring rule coverage and false-positive rate. Coverage is the fraction of the malicious-payload corpus that gets blocked, reported per injection class so you learn, for example, that SQL injection coverage is strong while coverage for server-side request forgery (SSRF, where the server itself is tricked into making an attacker-directed request) is weak, rather than one blended number that hides the gap. False-positive rate needs a separate corpus of known-benign-but-suspicious-looking requests, a legitimate user bio field containing the word "select," a legitimate query string with several ampersands, run through the same harness; a well-tuned WAF blocks most of the malicious corpus and passes most of the benign corpus, and both numbers should be reported together, since a WAF tuned to maximize coverage alone can trivially do so by blocking almost everything, which destroys usability.
Automating repeated runs safely. Rate-limit the harness's own traffic well below anything that could become a denial-of-service against the gateway itself or trip unrelated infrastructure alarms. Run against a dedicated staging instance of the gateway and WAF configuration wherever possible, since firing thousands of malicious-looking requests at a shared production WAF risks polluting production security logs and consuming the WAF's own CPU or rate-limit budget. Version the payload corpus and the WAF rule set together so a scheduled run's results are diffable against the prior run, did coverage regress after a rule was "simplified," did a new false positive appear after a rule was tightened, rather than being a one-off snapshot with nothing to compare it to.
Using results to tune policy. Feed per-class coverage gaps back to whoever owns the WAF rule set as a prioritized list, the class with the lowest coverage and highest business risk first, and feed false positives back as concrete before-and-after examples so the rule owner can narrow a match pattern rather than being told only "false positives exist." Treat the labeled corpus itself as a regression suite for the WAF configuration going forward: every rule change gets run against it before deployment, turning WAF tuning from an ad hoc, reactive process into something with its own test suite.
Trade-offs & pitfalls
A harness that only reports coverage, with no benign corpus and no false-positive number, will systematically push a WAF owner toward over-blocking, since coverage is the only metric being rewarded. It is also easy to under-invest in the evasion-variant half of the corpus and end up only proving the WAF's default rule set works, which the vendor already tested; the evasion variants are what make this exercise worth automating in the first place.
An application uses WebSockets for real-time features such as chat and notifications. Describe the specific security tests you would run for WebSocket endpoints, including authentication and origin checks, message schema validation, rate-limiting, and how you would intercept, modify, and replay WebSocket frames during testing.
Sample Answer
Direct answer
WebSocket endpoints need their own test plan because the browser protections you rely on for normal HTTP (Same Origin Policy blocking cross-site reads, Cross-Origin Resource Sharing (CORS) preflight, the browser's automatic pre-check request that asks the server's permission before making certain cross-origin calls) do not apply to the WebSocket handshake the same way, and Burp does not scan them by default the way it scans HTTP requests. You have to test the handshake (authentication and origin), the message stream (schema and injection), the connection's resilience (rate limiting), and you have to do the actual interception by hand.
Testing approach
-
Authentication and origin checks. A WebSocket connection starts as a normal HTTP request with an
Upgrade: websocketheader, and if the app authenticates that handshake using a cookie, the browser will attach the cookie automatically even to a request initiated by a page on a completely different site. This is Cross-Site WebSocket Hijacking (CSWSH): host a page on an attacker-controlled origin that opens anew WebSocket(...)to the target, and check whether the server accepts the connection using the cookie alone. A correctly defended server validates theOriginheader on the handshake against an explicit allow-list (CWE-346, Origin Validation Error) rather than a substring or regex match that a lookalike domain likevictim-app.attacker.comcould pass. If the app instead authenticates with a short-lived token passed at connect time, check that the token is scoped to that origin and cannot be replayed from a captured handshake. -
Message schema validation. Once the socket is open, there is usually no schema gate equivalent to an OpenAPI-validated REST body: send malformed JSON, wrong field types (a string where a number is expected), oversized frames, and unexpected extra fields to see whether the server validates each inbound message the same way it would validate an HTTP body. Any field that gets reflected back to other connected clients (a chat message, a notification payload) needs the same injection testing as an HTML body field, since stored/reflected cross-site scripting delivered over a WebSocket message is still XSS (OWASP Top 10:2025 A05, Injection) once the frontend renders it unsanitized.
-
Rate limiting. A WebSocket is a single long-lived connection, so the per-request throttles that protect HTTP endpoints often do not apply to it at all. Test how many messages per second a single open socket can send before the server pushes back, and separately test reconnect-storm behavior (rapid open/close cycles), since both are common ways a WS-heavy feature becomes a DoS vector nobody load-tested.
-
Intercept, modify, and replay frames. Burp's WebSockets history tab captures every frame on a connection once the handshake goes through Burp's proxy. Turn on intercept to pause an individual frame before it is sent, edit the JSON payload in place (for example, changing a
room_idfield to a room the test account should not have access to, testing broken access control the same way you would test an IDOR (Insecure Direct Object Reference, swapping an object identifier in a request to reach a resource that is not yours) in a REST body), and forward it. To test replay, capture a previously sent frame and resend it later in the same or a new connection to check whether the server treats messages as idempotent when they should not be (for example, resending a "mark notification read" frame, or a payment-confirmation-style message, to see if it is processed twice).
Worked example
A chat app authenticates its WebSocket purely off the session cookie and does not check Origin. A page hosted on attacker.example opens new WebSocket("wss://target.example/chat"); the browser attaches the victim's session cookie automatically; the server accepts the handshake and starts streaming the victim's private messages to the attacker's page in real time. That single finding (CSWSH via missing origin validation) is worth more than a dozen message-schema nitpicks, because it defeats the entire access control model for the feature.
Trade-offs & pitfalls
Automated scanners routinely skip WebSocket traffic entirely, so a report that only lists what the scanner found will miss this whole surface; budget dedicated manual time with Burp's WebSocket tooling. Watch for teams who point to wss:// (WebSocket Secure, running over TLS, Transport Layer Security) as proof the channel is secure: transport encryption says nothing about origin validation or message-level authorization, which are the actual gaps this class of testing targets.
Unlock Full Question Bank
Get access to all Penetration Testing Methodology and Execution interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.