Penetration Testing Methodology and Execution Questions
Running structured penetration-testing engagements end to end. Covers the pentest lifecycle, reconnaissance and information gathering, network scanning and enumeration (Nmap, service/version detection), tool selection and usage (Metasploit, Burp Suite), engagement scoping and planning, testing across target types, and findings reporting. The methodical offensive-assessment workflow.
Explain the OWASP Top 10 web application risks (current release). For each of the Top 10, provide: 1) a one-sentence description of the issue; 2) a concrete testing technique or tool you would use to validate it during a penetration test; and 3) one practical mitigation recommendation. Structure your answer as a list and assume a modern cloud-hosted web app.
Sample Answer
Direct answer
As of the OWASP (Open Web Application Security Project) Top 10:2025, the current published edition, the ten categories run from Broken Access Control at the top down to a newly added category on mishandling exceptional conditions. Each category names a class of weakness, not a single bug, so testing it means picking a technique that would surface that class on a modern cloud-hosted application, paired with the mitigation that closes the class rather than just the one instance you found.
Structured elaboration
| # | Category | One-sentence description | A testing technique | A mitigation |
|---|---|---|---|---|
| A01 | Broken Access Control | Users can act outside their intended permissions or reach another user's data, and this edition folds in server-side request forgery as a form of the server itself being tricked into acting outside its intended bounds | Swap object identifiers between two authenticated sessions of different privilege to see if the server enforces ownership, and test any server-initiated "fetch this URL" feature for attacker-controlled targets | Enforce authorization server-side on every request by default, never trust a client-supplied role or ID, and allowlist destinations for server-initiated outbound requests |
| A02 | Security Misconfiguration | Insecure defaults, verbose error output, or unnecessarily exposed services and features | Review response headers for missing security headers, and check for default credentials or exposed debug and admin panels | Maintain a hardened baseline configuration as code, with automated checks for configuration drift |
| A03 | Software Supply Chain Failures | Risk introduced through third-party dependencies, build pipelines, or compromised packages, broader than just an outdated library | Review a software bill of materials (SBOM) for known-vulnerable dependency versions, and check whether the build pipeline verifies artifact integrity | Maintain an SBOM, pin dependencies with verified checksums or signatures, and keep a real patch cadence |
| A04 | Cryptographic Failures | Sensitive data exposed due to missing or weak encryption in transit or at rest | Check TLS configuration and cipher strength, and look for sensitive data appearing in plaintext in logs, URLs, or storage | Enforce modern TLS, and use vetted algorithms for password hashing and stored-data encryption |
| A05 | Injection | Untrusted input is interpreted as part of a command or query rather than as plain data | Insert a syntax-breaking character into an input and observe whether the application's behavior changes in a way that suggests it reached an interpreter unescaped | Use parameterized queries or prepared statements everywhere (the query text is fixed and the user's input is sent to the database separately as data, so it can never be read as part of the SQL command), and run with a least-privilege database account |
| A06 | Insecure Design | A missing or inadequate security control baked into the architecture itself, not something a code patch alone fixes | Review the business logic of sensitive workflows, such as account recovery or payment, for a missing control a scanner would never flag, since there's no broken code, just an absent rule | Threat-model sensitive workflows during design, and use established secure design patterns |
| A07 | Authentication Failures | Weaknesses in login, session, or credential handling | Check for missing rate limiting or lockout on login, and test session token predictability | Require multi-factor authentication, use cryptographically random session tokens, and rate-limit authentication attempts |
| A08 | Software and Data Integrity Failures | Code, infrastructure, or data trusted without verifying its integrity, such as insecure deserialization (deserialization is turning stored or received bytes back into a live in-memory object, so deserializing untrusted input lets an attacker smuggle in a hostile object) or an unsigned auto-update mechanism | Check for endpoints that deserialize untrusted input, and check whether update mechanisms verify a signature before applying | Sign and verify updates and critical artifacts, and avoid deserializing untrusted data |
| A09 | Security Logging and Alerting Failures | Insufficient logging or alerting means an attack goes undetected | Attempt a detectable action, such as several failed logins, and check whether it's logged and whether anything actually alerts on it | Centralize logging with tamper resistance, and wire active alerting into a SIEM (Security Information and Event Management) system for suspicious patterns |
| A10 | Mishandling of Exceptional Conditions | Improper error handling or logic that fails open (on an error the system defaults to allowing the action, whereas failing closed defaults to denying it), or leaks internals, when something unexpected happens | Send malformed input or trigger a resource limit and observe whether the application fails open or leaks a stack trace | Fail closed by default, return generic error messages to users, and centralize exception handling with secure defaults |
Worked example
For A05, Injection, on a search endpoint that passes a query parameter straight into a database query: appending a single quote character to the parameter and observing an unhandled database error in the response, rather than a clean "no results" page, is a low-risk, high-signal way to confirm the input reaches an interpreter without proper escaping, without needing to actually extract any data to prove the class of vulnerability is present.
Trade-offs and pitfalls
- Citing a bare category number without naming the edition is misleading, since category numbers and even category names have moved between the 2021 and 2025 editions; Cryptographic Failures, for instance, was A02 in 2021 and is A04 here.
- The list ranks classes by aggregated prevalence and impact data, not by how severe any single instance you find will be; a lower-ranked category can still produce a critical finding on a specific application.
- Treating the list as a checklist to run through mechanically misses its actual purpose, a shared vocabulary for classes of risk, not a substitute for understanding the specific application in front of you.
Explain manual techniques to identify and exploit SQL injection when an application produces no error messages (blind SQLi). Cover detection, boolean-based and time-based payloads, payload tuning for different DBMS types, and safe guidance for demonstrating impact to developers without destructive actions.
Sample Answer
Direct answer
Blind SQL injection (SQLi) is exploited without ever seeing a database error or the query's raw output; instead you infer the truth or falsity of an injected condition, or measure a deliberately induced time delay, purely from how the application's own normal behavior changes. Because there is no error text to reveal anything about the backend, both the injection syntax and the specific payload have to be tuned to the target database management system's (DBMS) dialect, which makes blind techniques far more syntax-sensitive than error-based injection.
Structured elaboration
- Detection: start with any parameter that plausibly flows into a query where the application shows no distinguishing error for a malformed input, returning a generic page or an identical response whether the underlying condition is true or false. That absence of a visible error is itself the first signal you might be dealing with a blind case rather than an error-based or in-band one.
- Boolean-based detection: inject a pair of conditions, one that should evaluate true and one that should evaluate false, and compare the application's responses. The canonical minimal probe is appending
' OR '1'='1(always true) and, separately,' OR '1'='2(always false) to a parameter, then diffing the two responses by content length, specific text present or absent, or HTTP status. If the application behaves like a normal valid request when the condition is true and differently, for example a "no results" state, when it is false, the parameter is unsanitized and reaches a boolean-evaluated position in the query. - Time-based detection, used when the response content shows no visible difference at all, for example a background or logging query with no direct output: inject a conditional time delay and measure latency, comparing a true-condition payload against a false-condition equivalent, and confirm a consistent multi-second delay only on the true branch, repeated more than once to rule out ordinary network jitter.
- DBMS-specific payload tuning: this is where blind SQLi differs most from error-based injection, since there is no error message to reveal the backend. MySQL uses a
SLEEP(n)function; PostgreSQL usespg_sleep(n); Microsoft SQL Server usesWAITFOR DELAY '0:0:n'; Oracle has no direct sleep function, so testers typically substitute a deliberately slow, CPU-intensive computation as a proxy time-based signal. Comment syntax also differs across dialects, and getting it wrong is a common reason a real injection point silently produces no signal, which is easy to misread as "not vulnerable" when it is actually "wrong dialect assumed." - Safe demonstration without destructive actions: never use a payload that writes, deletes, or drops data to prove the finding, since a purely read-only boolean or time-based signal is entirely sufficient to prove exploitability. If a client wants to see concrete impact rather than a timing difference, demonstrate it with a narrowly scoped, single-value extraction of something already non-sensitive, such as the database version string pulled one character at a time, which proves arbitrary read access without ever touching real customer data.
Worked example
A search parameter ?id=5 behaves identically whether the id exists or not, always rendering the same custom "no results" page, with no SQL error ever surfacing. Testing ?id=5' OR '1'='1 returns the full normal result set (the true branch), while ?id=5' OR '1'='2 returns the "no results" page (the false branch): that difference is the boolean-based confirmation. To fingerprint the specific DBMS rather than guessing, you would test each dialect's sleep syntax in turn, for example ?id=5' OR SLEEP(5)-- - causing a roughly five-second delay while the PostgreSQL equivalent ?id=5' OR pg_sleep(5)-- - returns immediately, confirming a MySQL backend since only the matching dialect's function actually executes rather than silently failing. From there, blind read capability can be demonstrated safely with something like ?id=5' AND SUBSTRING(@@version,1,1)='5 (a true/false test of the first character of the version string), repeated across positions to reconstruct the full version, using a value that is already public and non-sensitive rather than any real customer data.
Trade-offs and pitfalls
- Time-based signals are noisy. A single multi-second delay could be network jitter or backend load rather than proof, so always confirm with a control run of the false-condition payload responding immediately, and repeat the true-condition test more than once before treating it as confirmed.
- Never escalate a proven blind SQLi finding into a full manual or automated data-extraction run during a live engagement without explicit client agreement on scope. Proving that unsanitized input reaches a SQL query and is fully exploitable only needs a handful of controlled probes; further automated extraction should follow the same read-only, minimal-footprint discipline agreed with the client, not default to a full data dump.
- Blind techniques are inherently slow, since character-by-character extraction requires many requests, and this pattern is exactly the kind of traffic likely to trip a rate limiter or web application firewall, so pace testing deliberately and coordinate timing with the client if the application is fragile.
Describe how you would structure a 20-minute executive briefing after a penetration test. Include slide topics and time allocation (e.g., summary, top risks, remediation roadmap, cost/impact), what material to present verbally versus in appendices, and an approach for handling difficult executive questions about legal exposure or remediation cost estimates.
Sample Answer
I treat 20 minutes as a fixed budget spent almost entirely on the three to five findings that actually change a business decision, push technical depth into an appendix, and rehearse the two or three hard questions I already know are coming, so I'm not improvising the organization's risk posture live in the room.
Time allocation
| Segment | Time | Content |
|---|---|---|
| Opening and context | 1-2 min | Scope, test window, objective, one-line risk posture |
| Executive summary | 3-4 min | Top 3-5 findings by business impact, a likelihood-times-impact framing, headline risk rating |
| Top risks, deep dive | 6-7 min | One slide per top finding: plain-language description, attack path in one sentence, business consequence |
| Remediation roadmap and cost/impact | 4-5 min | Quick wins versus structural fixes, rough effort tiers, expected risk reduction |
| Compliance and legal note (if relevant) | 1-2 min | Factual regulatory exposure only, no legal interpretation |
| Q&A and next steps | 2-3 min | Owners assigned per finding, retest date |
What's verbal versus in the appendix
The exploit-chain narrative and the business-impact framing get said out loud, in plain language. Raw evidence, full CVSS breakdowns, request/response dumps, and detailed remediation task lists live in an appendix the executives can hand straight to their engineering leads afterward, rather than being read aloud to a room that doesn't need that level of detail.
Handling difficult executive questions
- Legal exposure: stay strictly factual ("we demonstrated X was possible under Y conditions"), and explicitly defer any interpretation of liability to the client's counsel rather than characterizing exposure yourself.
- Remediation cost: give effort tiers (small, medium, large) grounded in what the finding actually requires to fix, rather than fabricating a precise dollar figure, and offer to loop in engineering leadership for a firmer estimate after the meeting.
Worked example
For an unauthenticated admin endpoint, the verbal line is something like: "This lets anyone on the internet who finds this URL take over the admin account without a password; it took my team about ten minutes to find." The appendix carries the full request/response evidence and the exact severity breakdown, but only the plain-language version ever reaches the slide.
Trade-offs and pitfalls
A common pitfall is spending the first ten minutes walking through methodology and tooling that executives neither need nor remember, leaving too little time for the findings that actually matter to them. Another is putting a full technical severity score or vector string directly on an executive slide; it reads as noise to that audience and belongs in the appendix instead.
You suspect a business logic flaw in a payments microservice allows creating transactions without sufficient balance checks. Outline a safe red-team style plan to test the hypothesis in a production-like environment: how to limit blast radius with test accounts, telemetry to enable, steps to construct a responsible PoC, and how to escalate findings to engineering and finance while preserving evidence.
Sample Answer
Direct answer
Testing a suspected balance-check bypass in a live payments service safely means proving the hypothesis with the smallest, most reversible transaction possible against a dedicated test account, with enhanced telemetry turned on beforehand, and escalating to engineering and finance the moment it is confirmed, not waiting for the final report.
Structured approach
Limiting blast radius. Never test against a real customer account or real money movement. Use dedicated test accounts with synthetic funding, and if the environment is genuinely production rather than staging, coordinate with engineering to flag those specific test account IDs so they are excluded from real settlement and downstream reconciliation before touching them, since a "successful" proof of concept that actually posts to the real ledger and gets picked up by nightly reconciliation is now a real incident, not a demonstrated finding. Cap the transaction amounts attempted at the smallest value that still proves the hypothesis, for example a one-dollar overdraft-style transaction rather than draining a large synthetic balance, to minimize consequence even if something behaves unexpectedly.
Telemetry to enable beforehand. Ask the client to temporarily elevate logging and distributed tracing on the payments service and any downstream ledger or settlement service, scoped specifically to the test account IDs to avoid capturing unrelated customer data. Confirm whether a monitoring dashboard or alert exists that would have caught this pattern in production, since part of the value of the exercise is finding out whether existing detection would have noticed a real attacker doing this, not only whether the bug exists.
Constructing a responsible proof of concept. Reproduce the suspected flaw with the smallest, most reversible action first: set a test account's balance to zero or near-zero and attempt a single low-value transaction to see whether the balance check is actually enforced server-side, rather than only in the client UI. Use Burp to intercept and directly manipulate the transaction request, bypassing whatever client-side balance check exists, since the point is testing the server's authority, not the UI's. Stop at the first successful confirmation; do not escalate the amount or repeat the transaction to see "how far it goes," since each repetition compounds real-world risk with no additional evidentiary value once the flaw is proven once.
Escalating while preserving evidence. Capture the full request and response pair, the account balance before and after (via screenshots or an authoritative API read), and the timestamp and any correlation ID. Notify engineering immediately, since a live balance-check bypass in a payments service needs an emergency fix path started now, not a wait for the final report. Notify finance and accounting as well, both so a single test transaction is properly flagged and excluded or reversed in reconciliation, and because finance can independently scan transaction history for the same anomalous pattern to check whether a real customer has already found and exploited the same flaw. Write the finding up immediately as an interim, out-of-band critical-severity alert with clear reproduction steps, then continue the engagement rather than sitting on a live, exploitable payments bug until the scheduled report date.
Worked example
A test account is set to a zero balance. Intercepting the "create transaction" request in Burp and submitting a one-dollar purchase confirms the server processes it and the account is now negative, with no server-side rejection, even though the client UI would have blocked the same action. The before/after balance reads, the raw request/response pair, and the timestamp are captured immediately, and an interim critical finding is sent to engineering and finance within the hour, well ahead of the final written report.
Trade-offs & pitfalls
The instinct to prove the flaw is "really bad" by pushing a larger transaction or repeating it several times is exactly the wrong instinct in a live payments system: it adds real financial and reconciliation risk without adding any evidentiary value beyond the first confirmed instance. The other common mistake is treating this like a normal finding that can wait for the scheduled report; a confirmed, exploitable financial-integrity bug in a live system is time-sensitive and should be escalated the same day it is confirmed.
Explain the purpose and typical contents of 'rules of engagement' (RoE) for an authorized penetration test. Provide concrete examples of restrictions organizations commonly impose (for example: test time windows, systems to avoid, thresholds for failed logins, or third-party-supplied systems), how to implement a kill-switch or emergency stop, and how to respond if an unexpected production outage occurs during testing.
Sample Answer
Direct answer
Rules of engagement (RoE) are the operational rulebook layered on top of the legal authorization: they spell out exactly what a tester may and may not do, when, and what happens if something goes wrong, so both sides have a shared, written answer before anything unexpected happens mid-test. Beyond the broad scope document, RoE typically nail down concrete restrictions, an explicit kill-switch process, and a pre-agreed response if testing itself causes an outage.
Structured elaboration
Common concrete restrictions organizations impose:
- Testing time windows, for example only during business hours, or conversely only after-hours, to limit collateral impact on real users.
- Systems to avoid entirely, legacy systems known to be fragile, anything already flagged as unstable, or third-party-supplied systems the client doesn't have authority to authorize testing against.
- Thresholds for automated actions, for example a maximum number of failed login attempts per account before automated authentication testing must stop, to avoid mass account lockouts.
- Third-party-supplied systems, a payment processor's hosted checkout or a SaaS vendor's admin console, anything requiring separate authorization from a party who isn't part of the client's own authorization letter.
Kill-switch and emergency stop implementation:
- A named, reachable point of contact on both sides for the full duration of the testing window, not just office hours.
- A pre-agreed communication channel, commonly a shared chat channel or a dedicated phone line, monitored throughout testing.
- A pre-agreed stop signal, a specific phrase or command, that immediately halts all active testing the moment either side invokes it, with no negotiation required in the moment.
- A written expectation that invoking the kill-switch is never treated as an accusation or a failure; it's simply the safety mechanism doing its job.
Response if an unexpected production outage occurs during testing:
- Stop all active testing immediately via the kill-switch, before investigating cause.
- Notify the pre-agreed emergency contact right away, even before root cause is confirmed, since minutes matter for a real outage.
- Preserve everything you were doing at the moment of the outage, the exact request, timestamp, and tool state, so the client's team can correlate it against their own monitoring, whether or not your testing turns out to be the actual cause.
- Only resume testing after the client explicitly confirms it's safe to do so, and consider adjusting technique, slower timing or avoiding the specific action that coincided with the outage, before resuming.
- Document the incident and the response in the final report regardless of whether testing was the actual cause, since it's a real event that happened during an authorized engagement.
Worked example
An RoE document for an internal network engagement might read: testing permitted weekday business hours only; do not target the legacy inventory management server flagged as unstable by the client; automated login testing must not exceed five failed attempts per account per hour; emergency contact reachable for the full testing window; kill-switch phrase honored immediately by all testers upon receipt in the shared incident channel. If the client's monitoring flags a service disruption during the window and the tester's own log shows a scan running against an in-scope host at that exact time, testing stops immediately via the kill-switch phrase, the emergency contact is notified before the cause is even confirmed, and the specific scan configuration is preserved for the client's own investigation.
Trade-offs and pitfalls
- An RoE with vague restrictions, such as "avoid causing problems," gives no actual guidance in the moment something looks risky; restrictions need to be concrete enough to act on without a judgment call under pressure.
- A kill-switch that only one side can invoke, or that requires approval before taking effect, isn't actually a kill-switch; it needs to work the instant either party says stop.
- Treating an outage during testing as automatically the tester's fault, or automatically not, before investigating wastes the client relationship either way; the right first move is always to stop and preserve evidence, not to assign blame.
Unlock Full Question Bank
Get access to all 36 Penetration Testing Methodology and Execution interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.