Secure Software Delivery: DevSecOps, Pipeline, and Supply Chain Security Questions
Embedding security into how software is built, assembled from dependencies, and shipped. Covers shift-left and secure-SDLC practices, infrastructure-as-code security, CI/CD pipeline and secrets management, integrating security scanning into build and deploy, and configuration and secret management across environments, together with software supply chain security: software composition analysis (SCA), dependency and open-source vulnerability management, build-provenance and artifact integrity, and mitigating supply-chain attack vectors. The 'secure the delivery pipeline and everything it pulls in' discipline, distinct from vendor-risk governance.
Describe how to safely integrate DAST scans into CI/CD for services that rely on third-party APIs and internal-only endpoints. Include strategies to avoid flaky results from external partners, protect credentials used by DAST tools, and ensure DAST tests do not cause harmful side effects in production.
Sample Answer
Integrating DAST safely into CI/CD for services that depend on third-party APIs and internal-only endpoints means the scan target itself introduces risk beyond the usual 'does this take too long' concern DAST placement discussions usually focus on.
Avoiding flaky results from external partners
Run DAST against a staging environment configured to use STUBBED or mocked responses for any third-party API dependency, rather than the actual live third-party service, since a live third-party dependency's availability, rate limits, or unrelated changes would make DAST results non-deterministic and unreliable as a gating signal; a stub that returns realistic, controlled responses keeps the scan's pass/fail outcome tied to your own application's behavior, not to a partner's uptime on any given day.
Protecting credentials used by DAST tools
The DAST tool itself typically needs authenticated access to properly test authenticated parts of the application, meaning it holds a real (or realistic test) credential; that credential should be scoped to the staging environment specifically, rotated on the same cadence as any other pipeline secret, and never be a credential that also has access to the production third-party account, since a DAST tool is itself effectively an automated, credentialed client making requests, and its credential deserves the same handling rigor as any other pipeline secret discussed throughout this topic.
Ensuring DAST doesn't cause harmful side effects in production
DAST must never run against production directly for a service where its probing (submitting forms, testing injection payloads, triggering workflows) could cause a real side effect (an actual email sent, an actual charge processed, actual data mutated); this is why DAST runs against a staging environment with fully isolated infrastructure, including isolated internal-only endpoints, rather than staging environments that happen to share a database or downstream integration with production. For internal-only endpoints specifically, make sure the staging environment's network topology actually mirrors production's internal segmentation, since testing an internal endpoint in an environment where it's accidentally reachable in a way production wouldn't allow gives a false sense of what's actually exposed.
Trade-offs
Stubbing third-party dependencies for DAST makes the scan reliable and fast but means DAST won't catch an integration-specific vulnerability that only manifests when interacting with the REAL third-party service's actual behavior; that gap is an accepted trade-off for CI-integrated DAST specifically, with a periodic, separate, less-frequent scan against a fully-live staging environment (accepting its slower, less-deterministic nature) as a supplementary check for exactly the integration-specific risk the stubbed version can't catch.
A popular third-party GitHub Action used across your org requests 'secrets' access. Evaluate the security risks of allowing third-party actions access to organization secrets and propose at least five mitigations or alternatives to reduce risk while maintaining developer productivity.
Sample Answer
A popular third-party GitHub Action requesting secrets access is exactly the kind of dependency risk that's easy to underweight, because it doesn't look like a typical software dependency: it's a workflow-level integration, but it can read anything the workflow's permission scope exposes to it.
Evaluating the risk
The core question is: what could this Action's code (or a future, compromised version of it) do with the secrets it's requesting access to, given that once granted, the Action runs with the SAME access the rest of the workflow step has. Check the Action's maintenance history (is it actively maintained by a reputable source, or a single-maintainer project with irregular updates), and check exactly which secrets it's requesting versus which secrets it actually needs for its stated function, since an Action requesting broader access than its function requires is itself a signal worth investigating.
Mitigations
- Pin the Action to an immutable commit SHA, not a mutable version tag, so an upstream compromise (a malicious update pushed to the same tag you trust) can't silently affect your workflows without you explicitly updating the pin.
- Scope the secrets available to the specific workflow step running this Action as narrowly as possible, using a job-level or step-level permission scope rather than exposing every organization secret to every step in the workflow by default.
- Run the Action in a workflow with restricted network egress, if your CI platform supports it, so even if the Action's code is malicious or compromised, its ability to exfiltrate whatever it reads is constrained.
- Prefer a first-party or well-audited alternative if one exists that accomplishes the same function without needing broad secrets access at all.
- Monitor the Action's actual behavior post-adoption (what network destinations does it reach, does its resource usage or runtime pattern change unexpectedly after an update) rather than treating the initial adoption review as a one-time, permanent clearance.
Balancing risk against developer productivity
The honest tension here is that a popular Action is popular because it saves real engineering time; banning all third-party Actions outright trades away that productivity for a security posture stronger than most organizations actually need. The mitigations above (pinning, scoping, monitoring) aim to preserve most of the productivity benefit while closing the specific, well-documented attack surface (a compromised or malicious Action silently exfiltrating whatever secrets it can reach) rather than treating 'no third-party Actions' as the only safe answer.
Trade-offs
Pinning to a commit SHA specifically trades away the convenience of automatically picking up the Action's latest release, requiring a deliberate, periodic review to bump the pin; that overhead is the direct, proportionate cost of closing the exact vulnerability class (a moving tag silently repointing to malicious code) that this whole evaluation exists to guard against.
Design an algorithm or pseudocode for scanning build artifacts for likely secrets using a combination of entropy analysis and regex patterns. Describe how you would minimize false positives (for example by whitelisting) and automatically trigger a revocation workflow for confirmed leaks while avoiding noisy rotations.
Sample Answer
Detecting a likely secret in a build artifact combines two independent signals, pattern matching for known secret SHAPES and statistical entropy for high-randomness strings that don't match any known pattern, since relying on either alone misses what the other catches.
import math
import re
from collections import Counter
REGEX_PATTERNS = {
"aws_access_key_id": re.compile(r"AKIA[0-9A-Z]{16}"),
"generic_api_key_assignment": re.compile(
r"(?i)(api[_-]?key|secret|token)\s*[:=]\s*['\"]([A-Za-z0-9_\-/+=]{16,})['\"]"
),
"private_key_header": re.compile(r"-----BEGIN ([A-Z]+ )?PRIVATE KEY-----"),
}
WHITELIST_TOKENS = {"AKIAIOSFODNN7EXAMPLE"} # AWS's own published example key
def shannon_entropy(s):
if not s:
return 0.0
counts = Counter(s)
length = len(s)
return -sum((c / length) * math.log2(c / length) for c in counts.values())
def find_candidate_tokens(line, min_len=20):
for token in re.findall(r"[A-Za-z0-9_\-/+=]{%d,}" % min_len, line):
if token in WHITELIST_TOKENS:
continue
has_upper = any(c.isupper() for c in token)
has_lower = any(c.islower() for c in token)
has_digit = any(c.isdigit() for c in token)
if sum([has_upper, has_lower, has_digit]) < 2:
continue
yield token
def scan_text(text, entropy_threshold=4.3):
findings = []
for line_no, line in enumerate(text.splitlines(), start=1):
for name, pattern in REGEX_PATTERNS.items():
for m in pattern.finditer(line):
if m.group(0) in WHITELIST_TOKENS:
continue
findings.append({"line": line_no, "rule": name, "match": m.group(0), "method": "regex"})
for token in find_candidate_tokens(line):
e = shannon_entropy(token)
if e >= entropy_threshold:
findings.append({"line": line_no, "rule": "high_entropy_string", "match": token,
"entropy": round(e, 2), "method": "entropy"})
return findings
Minimizing false positives
A whitelist covers known-public example values (AWS's own documented example access key appears in countless SDKs and tutorials and would otherwise fire on every scan of any codebase that includes a code sample). The character-diversity check on candidate tokens (requiring at least two of uppercase, lowercase, and digit) filters out long, low-diversity strings like repeated-character padding or a long all-lowercase identifier, which have low real entropy despite their length. Confirmed regex matches are treated as high-confidence; entropy-only matches with no corroborating regex pattern are lower-confidence and routed to human triage rather than immediately triggering an automatic revocation, since a random-looking but non-secret string (a generated test fixture ID, for instance) can still trigger the entropy threshold alone.
Avoiding noisy automatic rotation
An automatic revocation workflow should trigger only on a HIGH-confidence finding (a regex match on a known secret shape, ideally confirmed against an overlapping high-entropy string too), never on an entropy-only signal alone, since automatically rotating a credential based on a false positive creates real operational disruption for no security benefit; entropy-only findings route to a review queue instead.
Verified
Executed a seven-case test suite covering all three regex rules plus the entropy path: confirmed detection of a real AWS-shaped key; confirmed detection of a generic secret assignment; confirmed the whitelisted AWS example key is correctly suppressed; confirmed clean text produces no findings; confirmed a low-diversity repeated-character string has near-zero entropy while a genuinely random-looking token has high entropy; confirmed a regex-plus-entropy overlap correctly produces both a regex and an entropy finding on the same underlying secret; and confirmed private-key PEM header detection across all four real-world header variants (-----BEGIN RSA PRIVATE KEY-----, -----BEGIN EC PRIVATE KEY-----, -----BEGIN OPENSSH PRIVATE KEY-----, and the bare PKCS8 -----BEGIN PRIVATE KEY-----). All seven passed. (An earlier version of this pattern, (RSA|EC|OPENSSH|PRIVATE) KEY-----, only matched the bare PKCS8 header and silently missed the three traditional-format headers; the corrected pattern makes the type prefix optional so it matches all four.)
Trade-offs
The character-diversity filter and the whitelist both trade a small amount of detection sensitivity (a genuine secret that happens to be low-diversity or that happens to match a whitelisted pattern by coincidence would be missed) for a meaningfully lower false-positive rate; that trade is the right one specifically because a scanner with too many false positives gets its findings ignored entirely, which is a worse outcome than occasionally missing an unusual-shaped secret.
Write an OPA/Rego policy that denies creation or modification of object storage buckets that do not have server-side encryption enabled or that allow public access. Explain how you would integrate this policy into CI (pre-commit hooks and pipeline checks) and into runtime enforcement (admission controller or cloud governance). Describe unit and integration tests you would write to validate the policy.
Sample Answer
This policy denies two distinct storage-bucket misconfigurations: missing server-side encryption, and public ACL access, checked as separate rules so each produces its own specific, actionable message.
package storage.security
import rego.v1
deny contains msg if {
some resource in input.resource_changes
resource.type == "aws_s3_bucket"
not has_encryption_sibling(resource.address)
msg := sprintf("bucket %q has no associated server-side encryption configuration", [resource.address])
}
has_encryption_sibling(bucket_address) if {
some r in input.resource_changes
r.type == "aws_s3_bucket_server_side_encryption_configuration"
count(r.change.after.rule) > 0
startswith(r.address, sprintf("aws_s3_bucket_server_side_encryption_configuration.%s", [split(bucket_address, ".")[1]]))
}
deny contains msg if {
some resource in input.resource_changes
resource.type == "aws_s3_bucket_public_access_block"
resource.change.after.block_public_acls == false
msg := sprintf("bucket %q allows public ACLs (block_public_acls=false)", [resource.address])
}
Why encryption is checked via a sibling resource
In Terraform's AWS provider, server-side encryption for an S3 bucket is configured as a SEPARATE resource (aws_s3_bucket_server_side_encryption_configuration) linked to the bucket, not an inline attribute on the bucket resource itself; the policy has to look for that sibling resource's existence and content rather than checking a field directly on the bucket resource, which correctly models how the actual infrastructure is described and would silently miss the check entirely if it only inspected the bucket resource in isolation.
Integration into CI: pre-commit hooks and pipeline checks
Two distinct CI integration points matter here, and they catch different things. A pre-commit hook runs this same policy against the LOCAL Terraform plan (or a quick terraform validate plus a locally-generated plan JSON) before the developer even pushes, giving the fastest possible feedback loop, entirely on the developer's own machine; because it runs pre-push, it cannot be relied on as the actual gate, since a developer can skip or misconfigure a local hook. The pipeline check is what actually enforces the policy: it evaluates against terraform plan's JSON output on every pull request touching storage resources, as a pre-merge gate that cannot be bypassed the way a local hook can. The pre-commit hook exists purely to shift feedback left and save the round-trip to CI for an obvious violation; the pipeline check is the one that's actually authoritative.
Runtime enforcement
As a runtime backstop (since a resource created outside this specific pipeline, through the console or a different automation path, wouldn't be caught by a plan-time check at all), the same logic should also run as a cloud-governance policy (an AWS Config rule, or an admission-style check for infrastructure-as-code applied through a different path), catching drift or out-of-band changes the CI gate never saw.
Tests
A table-driven test suite should cover: a bucket with no encryption sibling at all (should deny), a bucket with an encryption sibling but an empty rule list (should deny, since a resource existing with no actual rule configured is not the same as being encrypted), a bucket with a properly configured encryption rule (should pass), and separately, a public-access-block resource with block_public_acls set to both true and false (should pass and deny respectively).
Verified
Evaluated with opa eval against five fixtures: a bucket with no encryption sibling (denied, as expected), a bucket whose encryption sibling has an empty rule list (denied), a bucket with a properly configured encryption rule (passed, no deny), a public-access-block resource with block_public_acls=false (denied) and with block_public_acls=true (passed). A combined fixture (unencrypted bucket plus public ACLs allowed) produced both deny messages together, and a combined fixture with both settings correct produced an empty result.
Trade-offs
Checking for the encryption sibling by matching on a naming convention (startswith against the bucket's resource name) assumes a consistent Terraform module naming pattern across the organization; a codebase with inconsistent naming between a bucket and its encryption configuration would need a more robust matching approach, such as checking the sibling resource's bucket reference attribute directly rather than inferring the relationship from resource address naming.
A pipeline runner was compromised and an attacker inserted a malicious build step that pushed a backdoored image to production. Describe detection methods to identify the compromise, containment steps across CI and production (including registry and K8s), evidence collection for forensics, remediation actions (revocation, rebuild, redeploy), and long-term controls to prevent recurrence.
Sample Answer
A compromised runner that pushed a backdoored image to production is a scenario where the response has to move fast on containment while preserving enough evidence to actually understand what happened, since those two goals can pull in opposite directions if handled carelessly.
Detection
The telltale signs are usually: a build that ran at an unexpected time or from an unexpected trigger, a build step that doesn't match what the pipeline definition in the repository actually specifies (indicating the runner's execution diverged from the declared pipeline, not just that the pipeline itself was malicious), or a signed image whose signing identity doesn't match the expected pipeline. Correlating the timeline of the suspicious build against the runner's own system logs (if the runner logs its own activity outside the build job itself) often reveals the injected step.
Containment, across CI and production
Immediately revoke the compromised runner's credentials and any secrets it had access to, and quarantine (do not terminate yet, if forensics are needed) the runner instance itself rather than the whole runner pool, to limit disruption to unrelated jobs. In the registry, prevent the backdoored image from being pulled further (mark it untrusted or remove the tag pointing to it) without necessarily deleting it yet, since the image itself is evidence. In production/Kubernetes, identify every workload currently running that backdoored image (matched by digest, not just tag, since a tag can be reused) and begin rolling those workloads back to the last known-good, signed image.
Evidence collection
Preserve the runner's build logs, the exact backdoored image and its layers, and the signing metadata (which identity signed it, when) before any remediation step that would overwrite or delete this evidence; a forensic snapshot of the compromised runner's filesystem, if feasible, before it's destroyed, is valuable for understanding exactly how the injection happened.
Remediation
Revoke the signing key or certificate that was used to sign the backdoored image if there's any indication the signing process itself, not just the runner, was compromised; rebuild a clean image from a verified-good commit on a runner confirmed uncompromised, sign it fresh, and redeploy. Rotate any secret the compromised runner had access to, on the assumption it may have been exfiltrated even if there's no direct evidence it was.
Long-term controls
Move to ephemeral, single-use runners so no future compromise can persist across jobs, add build-step verification that compares the runner's actual executed steps against the pipeline definition declared in source control (catching exactly this kind of divergence automatically going forward), and require signature verification at the Kubernetes admission layer so an unsigned or unexpectedly-signed image can never reach production even if a similar compromise recurs.
Trade-offs
Quarantining rather than immediately terminating the compromised runner preserves forensic value but leaves a compromised system technically still present in the environment for a short window; that trade is worth it specifically when understanding the root cause (to prevent recurrence) matters as much as stopping the immediate damage, though a runner posing an active, ongoing threat (still actively exfiltrating data) would tip the balance toward immediate termination instead.
Unlock Full Question Bank
Get access to all Secure Software Delivery: DevSecOps, Pipeline, and Supply Chain Security interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.