Exploitation, Post-Exploitation, and Red Team Operations Questions
The hands-on offensive tradecraft of compromising, pivoting through, and persisting in systems while evading defenses. Covers exploit development, privilege escalation, Active Directory and Windows exploitation, lateral movement, persistence, command-and-control, and attack chaining, extending into adversary-emulation campaigns: red-team engagement planning and objectives, multi-stage attack planning, operational security for offensive operators, and detection and defense evasion including web application firewall detection and bypass. The advanced offensive-operations layer executed against real targets, where staying undetected is itself an objective, distinct from the methodical scoped-assessment workflow of a penetration test.
Given the sudoers entry 'deploy ALL=(ALL) NOPASSWD: /usr/bin/vim', explain step-by-step how the user 'deploy' could escalate to root using vim features. Then propose at least two practical mitigations (configuration, engineering, or process) and describe how you would safely prove the issue to a customer without spawning an interactive root shell on production.
Sample Answer
This sudoers entry grants the deploy user the ability to run vim as root with no password prompt, and the mechanism that turns that into full root is simple: vim can spawn a subshell, and any subshell it spawns inherits whatever privilege the vim process itself is running with, which here is root.
Why it works
sudo doesn't sandbox what a permitted binary can do internally; it only decides that the binary may run as root at all. Vim, like many editors, has a documented way to escape to an interactive shell from within the editor (its shell-out functionality). Once that shell exists, it is a root shell, not a deploy-privileged one, because the process it was spawned from (vim) was already running as root. The NOPASSWD clause matters too: it means this escalation requires no further authentication step, so it works silently and can be scripted, not just used interactively.
Two mitigations
- Remove NOPASSWD, or better, remove the grant entirely and replace the underlying task with a narrowly scoped wrapper script that does only what
deployactually needs (for example, editing one specific configuration file) rather than granting an entire general-purpose editor with root privileges. Before granting any binary via sudo, check whether it has a documented shell-escape capability; a surprising number of "safe-looking" tools do, and this class of misconfiguration is well cataloged. - Use sudoers' restricted-execution options (such as
NOEXEC) to prevent the permitted binary from spawning child processes at all, which blocks this specific escalation path, though it isn't a universal fix since some tools have alternative ways to achieve code execution that don't go through a traditional child process.
Proving it safely, without a live root shell on production
Rather than actually spawning an interactive root shell on the client's real host, the recommended approach is: run sudo -l as the deploy user to show, in writing, the exact grant that exists; document that this binary is a documented shell-escalation risk (an established, cataloged class of misconfiguration, not a novel exploit); and, where a live demonstration is genuinely required, either use a cloned non-production instance, or have the sudo-elevated command perform a single benign, root-only action, such as reading a marker file placed there specifically for the test, rather than opening an interactive shell. A screenshot of that benign root-only read is strong enough evidence for the finding without ever putting an actual root TTY on a production system.
Trade-offs and pitfalls
A common mistake is treating "no password required" as the vulnerability, when the real vulnerability is granting an unrestricted general-purpose tool via sudo at all; removing NOPASSWD alone still leaves a working, if slightly less convenient, escalation path once the user re-authenticates. The mitigation has to address the scope of what's granted, not just the friction of using it.
Plan a red team engagement for an industrial control systems (ICS/OT) environment where process disruption is unacceptable. Define permitted activities, required safety review checkpoints, simulation strategies (controller emulation vs live commands), validation criteria, and explicit rollback procedures.
Sample Answer
Direct answer
For an ICS/OT (industrial control systems / operational technology) engagement, the plan itself must treat "do not disrupt the process" as a hard constraint, not a nice-to-have. That means: passive-only reconnaissance and IT-side techniques by default, anything with a physical side effect gets proven against an emulated or offline copy of the control system first, live production controllers are only touched after an explicit, named-authority safety sign-off inside a scheduled maintenance window, and every planned action has a rehearsed rollback and an abort trigger defined before testing starts. The plan looks less like a standard penetration test scope and more like a joint IT-security-and-plant-safety change-management process with a red team objective layered on top.
Structured elaboration
1. Scoping and permitted activities (rules of engagement)
Split authorization by the Purdue Model, the standard reference architecture for segmenting an ICS network from the corporate business layer down to the physical process:
| Layer | Typical assets | Default authorization |
|---|---|---|
| Level 4-5 (enterprise IT) | Corporate network, email, engineering workstations | Treat like a normal red team: phishing, Active Directory attacks, lateral movement are in scope |
| Level 3 (site operations) | SCADA (supervisory control and data acquisition) servers, historians, engineering workstations | Active testing generally in scope, but any state-changing action needs per-technique sign-off |
| Level 2 (control) | HMIs (human-machine interfaces), engineering stations used to program controllers | Passive observation and credential-access testing only; no direct manipulation |
| Level 0-1 (process, control) | PLCs (programmable logic controllers), RTUs (remote terminal units), sensors, actuators, safety instrumented systems (SIS) | Off-limits to live write or command traffic by default; anything here is proven on an emulated or offline twin, not live |
A named asset type appearing in more than one row above is not itself a scoping decision: "engineering workstation" by name does not determine the authorization level a given machine gets, and treating it as if it did is exactly the ambiguity that could get a control-system-connected machine attacked under the mistaken belief it is ordinary corporate IT. What determines the level is what the machine can actually do. A general corporate laptop that happens to have some engineering software installed, with no live connection to a controller, belongs in the Level 4-5 row. A site engineering or historian workstation on the operations network, used for monitoring and configuration but without direct write access to a PLC or HMI, belongs in the Level 3 row. Any workstation, regardless of which network segment it physically sits on or what a plant asset inventory happens to call it, that has direct programming or write access to a controller or HMI must be scoped and authorized as a Level 2 asset, no exceptions. Confirming which of these three categories a given "engineering workstation" actually falls into is a required output of the pre-contract kickoff review below, never an assumption made from the asset name alone.
Most ICS/OT red team engagements scope "assumed breach" starting at the IT/OT boundary rather than full black-box, since re-proving a phishing-to-domain-admin path adds little OT-specific insight. Budget and time should weight toward what happens after that foothold.
2. Safety review checkpoints
This is the part with no IT red-team equivalent, and it is the actual project-management backbone of the engagement:
- Pre-contract kickoff: plant safety engineer, control systems engineer, site operations lead, and IT security jointly review the process, the safety instrumented system architecture, and a first-pass blast radius for every technique the red team wants to test.
- Per-phase HAZOP-style walkthrough (Hazard and Operability Study, the standard process-safety technique for asking "what happens if this fails or behaves unexpectedly"): before any phase reaching Level 2 or below, walk the specific planned technique past the process engineer and ask what a malformed, delayed, or unexpected version of that command does to the physical process, not just to the network.
- Live-fire gate: a written, named sign-off before any packet reaches a production Level 0-2 device, naming the person with abort authority (usually the shift supervisor or control room operator on duty) who can halt the test immediately, for any reason, with no justification required.
- Maintenance window: intrusive or state-changing tests are scheduled inside existing planned downtime whenever one exists, so a mistake degrades to "the planned outage ran a little differently" rather than an unplanned plant state.
3. Simulation strategy: controller emulation vs. live commands
Default to emulation for anything beyond passive listening:
- Vendor simulators (for example, Siemens's PLCSIM for its SIMATIC controller family, or Rockwell's Logix Emulate) reproduce the exact firmware behavior of that vendor's controllers, the highest-fidelity option when the target vendor is known.
- Open testbeds (OpenPLC-based rigs, academic ICS security research testbeds) are cheaper and faster to stand up but may not reproduce vendor-specific firmware bugs, so a finding that only reproduces on a generic testbed needs that caveat in the report.
- Bench twin: a spare, identical physical controller wired to a simulated process, not the live one, gives the highest fidelity for anything vendor-firmware-specific, at the cost of procurement lead time, so this needs to be scoped and budgeted before testing starts.
The decision rule: prove the technique achieves its intended effect with no unintended side effect (crash, undefined state, unexpected restart) on the emulated target first. Escalate to a live production system only for actions that are read-only (observing state, not changing it) or a narrowly-scoped write already validated harmless on the twin and separately approved through the live-fire gate above.
4. Validation criteria
Define upfront what counts as "demonstrated" without requiring live impact: a captured, replayable command sequence the emulated controller accepts and acts on as predicted, a real credential obtained that the process engineer independently confirms would grant historian or engineering-station access, or a config review plus emulated proof-of-concept the process engineer signs off as workable against production. Treating "we ran it live and it worked" as the only acceptable validation is itself a planning failure, since it forces unnecessary live risk to satisfy a reporting preference.
5. Rollback procedures
Before touching any device or segment, even in a test expected to be non-destructive:
- Capture and store a baseline: current running configuration, ladder logic, and firmware version for anything the technique touches.
- Write the exact revert steps for that specific technique before running it, not after something goes wrong.
- Keep the OT engineer with change authority on standby (ideally physically present or on a live call) for the duration of any live-environment action.
- Define abort triggers up front: unexpected controller reboot, loss of communication to the safety instrumented system, alarm floods, or any physical process indicator the plant safety engineer names as a stop condition. Any one of these halts the test and triggers rollback immediately, with the abort authority from the live-fire gate empowered to make that call unilaterally.
Worked example
Take a water treatment plant engagement where the objective is to demonstrate whether an attacker who compromises the corporate network could reach and manipulate a chlorine dosing PLC, without actually manipulating it. A realistic plan:
- Week 1: assumed-breach start on the corporate IT network (a foothold is handed to the team rather than re-proving phishing), passive OT network discovery only, capturing traffic at the IT/OT boundary switch to map protocol traffic and identify the historian and engineering workstation, since active scanning against embedded ICS devices can crash them. The kickoff safety review happens before this week starts.
- Week 2: lateral movement and credential harvesting is tested live within the IT and site-operations layers; any technique that would reach the chlorine dosing PLC's own segment is instead run against an offline bench twin (an identical PLC model wired to a simulated dosing loop), where the team confirms a captured HMI credential would grant write access and that a crafted setpoint-change command is accepted by the twin.
- Week 3: the live-fire gate is presented to the plant. Since the dosing-change technique is a write action with real physical consequences (over-dosing is a safety event), it is not escalated to the live PLC. The finding is instead reported as "validated on identical hardware, would have succeeded against production" with the twin evidence attached, and remediation focuses on network segmentation and HMI credential hygiene rather than on the exploit primitive itself.
This keeps the demonstrated risk real and specific (a credentialed path from corporate email to a safety-relevant PLC) while the chlorine dosing setpoint itself is never touched on live equipment, exactly the trade the "process disruption is unacceptable" constraint is asking for.
flowchart TD
A[Scoping and rules of engagement] --> B[Joint kickoff: IT, OT engineering, safety, plant ops]
B --> C[Passive-only recon: boundary packet capture, asset inventory]
C --> D{Safety review gate 1}
D -->|approved| E[Emulated testing: PLC simulator or bench twin]
D -->|not approved| A
E --> F{Safety review gate 2}
F -->|technique is read-only or pre-validated harmless| G[Limited live validation in maintenance window]
F -->|residual risk too high| H[Report as theoretical finding, no live test]
G --> I[Rollback and credential rotation]
H --> J[Reporting and remediation]
I --> J
Trade-offs and pitfalls
- Being overly conservative (never touching the OT network, testing only IT-side lateral movement) under-tests the actual risk the engagement exists to measure; the OT-specific misconfigurations are exactly what a purely IT-style assumed-breach test misses.
- Being overly aggressive (running standard IT scanning tools, like unthrottled connection scans, against PLCs and RTUs) can crash fragile embedded devices that were never designed to handle unexpected traffic, turning a security test into an actual outage. Passive-first tooling has to be the default against Level 0-2 devices, with active probing reserved for segments explicitly cleared for it.
- Schedule pressure to compress the engagement is the most common way safety checkpoints get skipped in practice. The plant's own change-management calendar, not the red team's preferred timeline, is usually the real scheduling constraint, so build it into the plan from the start rather than negotiating it under time pressure later.
- Emulation fidelity is a real budget trade-off: a generic testbed is fast to stand up but a finding that only reproduces there needs a caveat, while a vendor simulator or physical bench twin costs more and takes longer to source, so fidelity requirements need to be decided before pricing the engagement.
- Frameworks worth citing when writing the plan: NIST SP 800-82 Rev. 3 ("Guide to Operational Technology Security") and the IEC 62443 series give the baseline segmentation and risk-assessment language (zones and conduits, the Purdue levels used above) that most plant safety teams already recognize, which makes the safety conversation faster than introducing red-team-specific vocabulary from scratch.
A web application accepts JSON payloads and is protected by a web application firewall (WAF). During authorized, in-scope security testing you find that many injection test payloads are blocked. Explain conceptually WHY a WAF misses some payloads (encoding and normalization differences, parser differentials between the WAF and the application, content-type handling), how a tester ensures test coverage while staying within scope and behaving ethically, and what this teaches defenders about WAF tuning and defense-in-depth. Keep it at the level of understanding WAF limitations, not a catalogue of specific bypass strings.
Sample Answer
A web application firewall (WAF) inspects a request with its own parser and its own decoding rules, then passes the same raw bytes on to the application, which parses that request a second time with a different parser. A parser differential is any point where the two parsers disagree about what a byte sequence actually means. Most missed injection payloads during authorized testing come down to exactly the three gaps the question names: how each side decodes and normalizes encoded characters, how strictly each side follows the JSON grammar, and how each side reacts when the declared content type does not match the real body. The WAF only has to be wrong once for its rule set to miss what the application still executes.
Structured elaboration
Encoding and normalization differences. Encoding is representing characters as bytes or escape sequences (percent-encoding, JSON's own backslash-u escape for a character); normalization is collapsing different representations of the same content into one canonical form before comparing it against a signature. A WAF typically runs one fixed decode-and-normalize pass before checking the result. If the application's parser performs an extra decode the WAF did not, such as unescaping backslash-u sequences or normalizing Unicode differently, a payload that looks inert to the WAF's single pass can become live application syntax once the app finishes its own decoding.
Parser differentials between the WAF and the application. Even without encoding tricks, two JSON parsers rarely agree on every grammar edge case: duplicate keys in one object, malformed UTF-8, or how strictly the grammar is enforced. A common example is a body containing the same key twice with two different values: many inspection engines evaluate the first occurrence, while many application-side parsers keep the last one assigned. If the WAF inspects the first, benign value and the application acts on the second, the WAF has correctly evaluated a value the application never actually uses.
Content-type handling. A WAF's JSON-aware rules usually only fully parse a body when Content-Type says application/json. Frameworks are often more lenient: some still parse a body as JSON when the header is missing, mislabeled as text/plain, or carries an unexpected charset, favoring compatibility over strictness. If the WAF skips deep inspection on a content type it does not recognize as JSON while the application parses the body anyway, the request passes through uninspected.
Ensuring coverage while staying in scope and behaving ethically. Deliberately varying encodings to find what a control misses is itself a form of testing that control, not just the application, so it has to be negotiated up front rather than assumed. Get written authorization in the rules of engagement that explicitly covers WAF-evasion testing, and where possible have the client allowlist your test source or issue a WAF-bypass header so you can separate two questions: does the underlying vulnerability exist, and does the WAF actually catch it. Keep the variation structured and low-volume, mirroring documented parser edge cases rather than brute-forcing every transform, coordinate timing with the client's security operations team so a spike in blocked requests is not read as a real attack, and log every variant so findings are reproducible.
What this teaches defenders about WAF tuning and defense-in-depth. A WAF is a compensating control, not a fix: if the underlying code still concatenates untrusted input into a query, a rule change on either side can reopen the same gap. Defenders reduce parser differentials by validating input using the same parsing logic the application itself uses rather than trusting a separate box to fully understand the app's grammar, by enforcing that the declared content type matches the real body instead of parsing leniently, and by logging requests the WAF allowed through so later findings can be checked retroactively. The durable fix for a parser differential lives in secure coding at the application layer (parameterized queries, output encoding, strict schema validation) at least as much as in WAF rule tuning, because the WAF is one layer of defense-in-depth, not the whole defense.
Worked example
Testing a /search endpoint that accepts {"query": "<value>"} behind a WAF rule that blocks a known injection test token when it appears literally in the body illustrates the mechanism without needing a reusable payload string in the write-up:
- Send the test token as plain ASCII inside the JSON value. The WAF blocks it, confirming its signature matches against the literal, already-decoded bytes of the string.
- Send the identical characters, but express part of the token using JSON's own backslash-u escape syntax so the bytes on the wire no longer contain the literal token. If the response is now a normal application reply instead of a block page, and the application's behavior (a database error, a changed result count, a slower response) shows the token still reached the same code path as the first test, that proves the application's parser unescapes those sequences before evaluating the string while the WAF's engine matches before that unescaping happens. That is the parser differential named above, demonstrated and reportable as "the WAF inspects before JSON unescaping; the application evaluates after it," not as a copy-pasteable bypass string.
Trade-offs & pitfalls
A block-page count measures the WAF's behavior, not the application's real exposure: "most payloads got blocked" can coexist with a fully exploitable flaw sitting behind the one variant that got through, so severity should track the underlying vulnerability, not how rarely the WAF missed it. Reaching for a scanner's built-in evasion features before securing encoding-variation authorization risks generating traffic that reads as an unauthorized bypass attempt and can trip defenses (rate limiting, an on-call page) outside the agreed test window. On the defense side, the common wrong turn is treating every discovered miss as "add one more signature": that chases the encoding found this month and leaves the underlying parser differential, and the unvalidated code behind it, untouched for the next one.
Describe the communication and escalation procedures you would set up before starting a red team exercise. Include the emergency contact list structure, agreed safety triggers or 'kill switches', notification blackout windows, and how to verify receipt of critical messages.
Sample Answer
Before day one, lock down four things in a signed communications plan: a tiered emergency contact list, pre-agreed "kill switch" triggers, notification blackout windows, and a receipt-confirmation protocol for anything urgent. All four exist because a red team is operating against live production, and the biggest operational risk isn't getting caught, it's a real incident happening while nobody trusted can reach anybody trusted.
Tiered contact list
Structure it in tiers, not one flat list:
- Tier 1 (operational): team lead plus a named "trusted agent" on the client side, the one person read into the fact the engagement is live, contacted for anything test-related.
- Tier 2 (executive): the CISO (Chief Information Security Officer) or security leadership, contacted for scope questions or a real incident.
- Tier 3 (leadership/legal): contacted only for a major incident or an engagement-ending event.
Publish this list inside the signed rules of engagement (ROE, the document governing what's allowed and how), not a side email, so it survives personnel turnover mid-engagement. Name at least two people per tier; a single point of contact going on leave has stalled real engagements.
Kill switches / safety triggers
These need to be objective conditions agreed in advance, not judgment calls made mid-incident: a defined production error-rate threshold on a monitored dashboard, the trusted agent invoking a pre-agreed safe word over the emergency channel, or an automatic pause if contact with the trusted agent is lost for more than an agreed window (for example 4 hours) during active testing. Authority to pull the switch should sit with a small named group on each side, and it should default to stop under any ambiguity.
Notification blackout windows
Negotiate specific dates and times, independent of the kill switch, when testing must pause or scale back: a quarterly earnings close, a major product launch, a scheduled maintenance window when the client's own on-call team is already stretched thin. Document exact timestamps and time zone in the ROE. Ambiguity here has caused real production incidents to be misattributed to the red team simply because the two sides read a blackout window differently.
Verifying receipt of critical messages
Any message that could stop or escalate the engagement needs a positive acknowledgment loop, not a fire-and-forget notification: a required response service-level agreement (SLA) plus a named fallback contact if the primary doesn't confirm receipt in time.
Worked example: severity ladder
| Severity | Trigger | Notified | Response SLA | If no acknowledgment |
|---|---|---|---|---|
| SEV-1 | Production outage or kill switch invoked | Trusted agent + team lead | 15 minutes | Escalate to CISO, halt all testing regardless |
| SEV-2 | Scope question or unexpected finding | Trusted agent | 2 hours | Escalate to team lead, pause affected workstream |
| SEV-3 | Routine daily status | Trusted agent | End-of-day digest | No escalation required |
Trade-offs and pitfalls
Routing everything through one named person with no backup is the single most common failure. Using a channel the client's own security operations center (SOC) already monitors as the emergency channel is a self-inflicted operational security (opsec) risk if that channel's logs get discovered or subpoenaed. And a kill switch phrased as "stop if things look bad" is useless under stress; only crisp, binary triggers hold up in the moment.
Create a scoring and metrics framework to measure red team success across technical outcomes (e.g., foothold achieved, data exfiltrated) and organizational outcomes (e.g., MTTD reduction). Specify categories, weighting, an example rubric with scores, and how you would present results differently for technical teams and executives.
Sample Answer
Build the rubric around two axes scored separately, technical outcomes (did the team actually achieve access or impact) and organizational outcomes (did the client detect and respond well), weight them explicitly so no single flashy technical result dominates the story, then translate the same underlying scores into two different presentations for two different audiences.
Why split it this way
Pure technical success (foothold, privilege escalation, data access) measures the red team's competence but says nothing about the client's defenses. Pure detection metrics, mean time to detect (MTTD) and mean time to respond (MTTR), measure the client's defenses but say nothing if the red team never got far enough to trigger them. A useful framework needs both, so an engagement that fails to get a foothold but forces several detections still scores as valuable, and one that reaches the crown jewels completely undetected registers as the serious finding it is.
Example category weighting (illustrative; tune per engagement objective)
- Objective / crown-jewel access achieved: 30%
- Detection and response effectiveness (MTTD/MTTR against the kill chain): 25%
- Operational stealth maintained (how far the team got relative to how long it stayed undetected): 15%
- Proof-of-impact quality (clean, non-destructive evidence versus a vague claim): 15%
- Remediation-readiness of the final report (how actionable and prioritized the findings are): 15%
Example rubric with scores
Score each category 0-5 against a published anchor definition (0 = not attempted, 5 = fully achieved with strong evidence), then convert to a weighted percentage. Suppose one engagement scores: objective access 4/5, detection and response 2/5, stealth 5/5, proof-of-impact 3/5, remediation-readiness 4/5:
Weighted score=54(30)+52(25)+55(15)+53(15)+54(15)=24+10+15+9+12=70That reads as: strong access, weak detection and response, excellent stealth, decent evidence, solid remediation guidance, an overall 70 out of 100. The individual category scores matter more than the single number; the weighted total is a headline, not the finding.
Presenting it to different audiences
Technical teams get the full rubric breakdown, the underlying kill-chain narrative mapped to a framework such as MITRE ATT&CK (a widely used catalog of adversary tactics and techniques), and raw MTTD/MTTR timestamps so they can validate their own detection engineering against what actually happened. Executives get the single weighted score, a one-page narrative of what an attacker could have done to the business in terms of dollars, customers, or downtime, a trend line against the previous engagement's score, and no more than three headline recommendations; the underlying category breakdown stays available on request but isn't the leading artifact.
Trade-offs and pitfalls
A framework that only measures technical success incentivizes chasing the flashiest compromise rather than the most business-relevant one; a framework that only measures detection incentivizes the team to go loud just to "test the SOC." Keep both axes weighted even when one stakeholder cares more about the other. Weighting stealth too heavily punishes a team for succeeding fast, and weighting detection too heavily punishes a client whose security operations center (SOC) simply never got the chance to see anything, so recalibrate the weights to the specific objective the sponsor cares about rather than reusing one fixed rubric on every engagement. A rubric with no published anchor definitions just relocates the subjectivity into "what does a 4 mean," so publish what each score level looks like before scoring, not after.
Unlock Full Question Bank
Get access to all Exploitation, Post-Exploitation, and Red Team Operations interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.