Secure Software Delivery: DevSecOps, Pipeline, and Supply Chain Security Questions
Embedding security into how software is built, assembled from dependencies, and shipped. Covers shift-left and secure-SDLC practices, infrastructure-as-code security, CI/CD pipeline and secrets management, integrating security scanning into build and deploy, and configuration and secret management across environments, together with software supply chain security: software composition analysis (SCA), dependency and open-source vulnerability management, build-provenance and artifact integrity, and mitigating supply-chain attack vectors. The 'secure the delivery pipeline and everything it pulls in' discipline, distinct from vendor-risk governance.
Provide a sample CI/CD workflow (YAML or pseudocode) that enforces separation of duties: developers can build and push artifacts but cannot promote to production; release and deployment require an independent approver and only signed artifacts are promoted. Include artifact signing and verification and least-privilege runner identities.
Sample Answer
Enforcing separation of duties in a pipeline means the same identity that can build and push an artifact structurally cannot also be the one that approves and executes its promotion to production; this has to be enforced by permissions and explicit configuration, not by a policy document asking people not to do both.
Workflow design
name: build-and-promote
on:
push:
branches: [main]
jobs:
build:
runs-on: ubuntu-latest
permissions:
contents: read
packages: write
id-token: write
steps:
- uses: actions/checkout@v4
- name: Build and push (unsigned, dev-only tag)
run: docker build -t "$IMAGE:dev-${{ github.sha }}" . && docker push "$IMAGE:dev-${{ github.sha }}"
- name: Sign image
run: cosign sign --yes "$IMAGE:dev-${{ github.sha }}"
promote-to-production:
needs: build
runs-on: ubuntu-latest
environment:
name: production # requires configured required reviewers AND "Prevent self-review" enabled
permissions:
contents: read
packages: write
id-token: write
steps:
- name: Verify signature before promoting
run: cosign verify --certificate-identity-regexp ".*" --certificate-oidc-issuer https://token.actions.githubusercontent.com "$IMAGE:dev-${{ github.sha }}"
- name: Retag and push production tag
run: |
docker pull "$IMAGE:dev-${{ github.sha }}"
docker tag "$IMAGE:dev-${{ github.sha }}" "$IMAGE:prod-${{ github.sha }}"
docker push "$IMAGE:prod-${{ github.sha }}"
How separation of duties is actually enforced
The build job runs under the identity of whoever pushed the commit, with permissions scoped to push a dev-tagged image and sign it, but with NO permission to push a prod- tag directly. The promote-to-production job is gated by GitHub's own environment protection rule, which requires a configured human approver before the job runs at all. That alone is not enough to guarantee separation of duties: by default GitHub does not stop the person who triggered the workflow from also being the required reviewer who approves their own promotion. The environment's protection settings must explicitly enable "Prevent self-review" for the approver to be guaranteed distinct from whoever pushed the change; without that setting turned on, a developer with reviewer permissions could approve their own promotion, which would defeat the entire point of this control. With it enabled, "developers can build and push but cannot promote" becomes an actually-enforced permission boundary rather than a documented convention someone could bypass.
Signature verification and least-privilege runner identities
The promotion job re-verifies the image's signature before retagging it for production, so promotion depends on the signature actually being valid at promotion time, not merely on trusting that the build job signed it correctly earlier. For keyless verification, cosign verify requires both a certificate-identity match and a certificate-OIDC-issuer match (here, GitHub Actions' own OIDC issuer); a verify command that supplies only the identity regex and omits the issuer will fail with a missing-required-flag error rather than silently passing, so both flags have to be present for the check to run at all. Each job uses its own scoped permissions block (least privilege per job, following GitHub Actions' job-level permission model) rather than one broad permission set shared across the whole workflow.
Trade-offs
Requiring a re-verification of the signature at promotion time, rather than trusting the build job's own signing step, is a small amount of redundant work but is exactly what prevents a compromised or buggy build job from being able to push an unsigned or improperly-signed image straight to the production tag by skipping its own signing step. Similarly, explicitly enabling "Prevent self-review" costs nothing beyond a one-time configuration change, but skipping it leaves the entire separation-of-duties guarantee resting on an assumption about GitHub's default behavior that does not actually hold; the promotion job's independent signature check plus the self-review-blocked environment rule together are what actually enforce the separation, not just the presence of a signing step and an approval gate somewhere in the pipeline.
Design an algorithm or pseudocode for scanning build artifacts for likely secrets using a combination of entropy analysis and regex patterns. Describe how you would minimize false positives (for example by whitelisting) and automatically trigger a revocation workflow for confirmed leaks while avoiding noisy rotations.
Sample Answer
Detecting a likely secret in a build artifact combines two independent signals, pattern matching for known secret SHAPES and statistical entropy for high-randomness strings that don't match any known pattern, since relying on either alone misses what the other catches.
import math
import re
from collections import Counter
REGEX_PATTERNS = {
"aws_access_key_id": re.compile(r"AKIA[0-9A-Z]{16}"),
"generic_api_key_assignment": re.compile(
r"(?i)(api[_-]?key|secret|token)\s*[:=]\s*['\"]([A-Za-z0-9_\-/+=]{16,})['\"]"
),
"private_key_header": re.compile(r"-----BEGIN ([A-Z]+ )?PRIVATE KEY-----"),
}
WHITELIST_TOKENS = {"AKIAIOSFODNN7EXAMPLE"} # AWS's own published example key
def shannon_entropy(s):
if not s:
return 0.0
counts = Counter(s)
length = len(s)
return -sum((c / length) * math.log2(c / length) for c in counts.values())
def find_candidate_tokens(line, min_len=20):
for token in re.findall(r"[A-Za-z0-9_\-/+=]{%d,}" % min_len, line):
if token in WHITELIST_TOKENS:
continue
has_upper = any(c.isupper() for c in token)
has_lower = any(c.islower() for c in token)
has_digit = any(c.isdigit() for c in token)
if sum([has_upper, has_lower, has_digit]) < 2:
continue
yield token
def scan_text(text, entropy_threshold=4.3):
findings = []
for line_no, line in enumerate(text.splitlines(), start=1):
for name, pattern in REGEX_PATTERNS.items():
for m in pattern.finditer(line):
if m.group(0) in WHITELIST_TOKENS:
continue
findings.append({"line": line_no, "rule": name, "match": m.group(0), "method": "regex"})
for token in find_candidate_tokens(line):
e = shannon_entropy(token)
if e >= entropy_threshold:
findings.append({"line": line_no, "rule": "high_entropy_string", "match": token,
"entropy": round(e, 2), "method": "entropy"})
return findings
Minimizing false positives
A whitelist covers known-public example values (AWS's own documented example access key appears in countless SDKs and tutorials and would otherwise fire on every scan of any codebase that includes a code sample). The character-diversity check on candidate tokens (requiring at least two of uppercase, lowercase, and digit) filters out long, low-diversity strings like repeated-character padding or a long all-lowercase identifier, which have low real entropy despite their length. Confirmed regex matches are treated as high-confidence; entropy-only matches with no corroborating regex pattern are lower-confidence and routed to human triage rather than immediately triggering an automatic revocation, since a random-looking but non-secret string (a generated test fixture ID, for instance) can still trigger the entropy threshold alone.
Avoiding noisy automatic rotation
An automatic revocation workflow should trigger only on a HIGH-confidence finding (a regex match on a known secret shape, ideally confirmed against an overlapping high-entropy string too), never on an entropy-only signal alone, since automatically rotating a credential based on a false positive creates real operational disruption for no security benefit; entropy-only findings route to a review queue instead.
Verified
Executed a seven-case test suite covering all three regex rules plus the entropy path: confirmed detection of a real AWS-shaped key; confirmed detection of a generic secret assignment; confirmed the whitelisted AWS example key is correctly suppressed; confirmed clean text produces no findings; confirmed a low-diversity repeated-character string has near-zero entropy while a genuinely random-looking token has high entropy; confirmed a regex-plus-entropy overlap correctly produces both a regex and an entropy finding on the same underlying secret; and confirmed private-key PEM header detection across all four real-world header variants (-----BEGIN RSA PRIVATE KEY-----, -----BEGIN EC PRIVATE KEY-----, -----BEGIN OPENSSH PRIVATE KEY-----, and the bare PKCS8 -----BEGIN PRIVATE KEY-----). All seven passed. (An earlier version of this pattern, (RSA|EC|OPENSSH|PRIVATE) KEY-----, only matched the bare PKCS8 header and silently missed the three traditional-format headers; the corrected pattern makes the type prefix optional so it matches all four.)
Trade-offs
The character-diversity filter and the whitelist both trade a small amount of detection sensitivity (a genuine secret that happens to be low-diversity or that happens to match a whitelisted pattern by coincidence would be missed) for a meaningfully lower false-positive rate; that trade is the right one specifically because a scanner with too many false positives gets its findings ignored entirely, which is a worse outcome than occasionally missing an unusual-shaped secret.
Design a CI/CD pipeline for a microservices web application showing where and when to run SAST, SCA, DAST, unit tests, and integration tests. Define security gates (which findings block progression), fail criteria, and a rollback strategy for DAST findings that are discovered post-merge. Discuss latency considerations and how to keep developer feedback fast.
Sample Answer
The placement decision comes down to matching each check's speed and blast radius to the point in the pipeline where a false result is cheapest to act on.
Where each check runs
flowchart LR
PR[Pull Request] -->|SAST + SCA, seconds-minutes| Merge[Merge to main]
Merge --> Build[Build + Unit Tests]
Build --> Deploy[Deploy to Staging]
Deploy -->|DAST, minutes-hours, async| Prod{Promote to Production?}
Prod -->|post-merge finding blocks HERE, not the PR| Prod
- SAST and SCA run synchronously at PR time, gating the merge itself. Both are fast (seconds to a few minutes on incremental changes) and need nothing but the source and dependency manifest, so blocking the merge on a CRITICAL finding does not meaningfully slow the team down.
- Unit and integration tests run in the standard build stage after merge, same as any other CI job.
- DAST runs against staging, asynchronously, after deploy, because it needs a live, running target and can take minutes to hours for a thorough scan; making a developer wait on that synchronously at PR time would make the pipeline unusable.
Gating criteria and the DAST rollback wrinkle
SAST and SCA findings above a defined severity threshold (typically CRITICAL, with a documented exception path for anything lower) block the merge outright, since the fix is cheap at this stage and the developer has full context. DAST findings are discovered after the code is already merged and often already promoted toward production, so the response is different in kind: rather than blocking a merge that already happened, the pipeline should automatically halt further promotion (do not let this build reach production) and, if it already reached a canary or production slice, trigger the standard rollback path rather than inventing a bespoke one. This is why the security gate design has to include an explicit answer to 'what happens when the finding shows up after the code has already moved past the point that finding would normally block' rather than assuming every check fires at the same stage.
Keeping developer feedback fast
Two practical techniques keep this from feeling slow: run SAST and SCA incrementally (only re-scan changed files and their dependency deltas, not the whole repository on every push), and surface findings as inline PR comments on the exact line rather than a link to an external dashboard, so the fix loop stays inside the tool the developer is already using. On a monolith versus a microservices layout the same principles apply, but a monolith often needs more aggressive incremental scanning (changed-file analysis) to keep PR-time SAST fast, since the whole-repository scan surface is larger.
Assess the pros and cons of automatically blocking PR merges for vulnerabilities above a CVSS threshold vs allowing merges and auto-creating prioritized remediation tickets. Consider developer productivity, attack window, context-aware exploitability, and false positives. Recommend a policy that balances security and velocity and describe an exception process.
Sample Answer
Auto-blocking on a raw CVSS (Common Vulnerability Scoring System) threshold and auto-creating a ticket both fail in the same direction if applied blindly: CVSS alone doesn't distinguish a theoretical, unreachable finding from an actively-exploited one, so a policy built purely on that score either blocks too aggressively or lets real risk slide through as 'just a ticket'.
Pros and cons of each
Auto-blocking on CVSS threshold: strong guarantee (nothing above the threshold ever ships without explicit action), but high false-positive cost, since CVSS alone doesn't account for whether the vulnerable code path is actually reachable in this specific application; blocking on raw CVSS routinely stops merges for findings that pose no practical risk in context, which is exactly the pattern that erodes developer trust in the gate.
Allow-merge-plus-ticket: preserves developer velocity and avoids blocking on unreachable findings, but risks a genuinely dangerous, actively-exploited vulnerability shipping to production while its ticket sits in a backlog with no enforced urgency, especially if the ticket has no real deadline or escalation path.
Considering developer productivity, attack window, and context-aware exploitability
The honest resolution isn't picking one policy over the other uniformly, but making the blocking decision a function of BOTH severity and reachability/exploitability together: a CRITICAL, actively-exploited, reachable finding blocks the merge outright regardless of any velocity cost, since the attack-window cost of shipping it clearly outweighs the delay cost of fixing it now; a HIGH-CVSS but unreachable or theoretical finding becomes a tracked, SLA-bound ticket rather than a blocker, since blocking here buys little real risk reduction for a real velocity cost.
A balanced policy
Block automatically only on the combination of high severity AND either confirmed reachability or known active exploitation; everything else routes to a ticket with a severity-tiered SLA (days for the highest tier, weeks for lower tiers) and an escalation path if the SLA is missed, so 'just a ticket' still carries real enforced urgency rather than becoming an indefinitely-deferred backlog item.
An exception process
Even a rule this well-calibrated will occasionally need a genuine, documented exception (a finding that's technically reachable but mitigated by a compensating control elsewhere the automated policy can't see); the exception needs an owner, a stated reason, a re-review date, and visibility to the security team, following the same pattern used for every other exception mechanism discussed in this topic.
Trade-offs
This combined-criteria policy is more sophisticated to implement (it requires reachability analysis, not just a CVSS lookup) than either pure-blocking or pure-ticketing alone, but it's the calibration that actually distinguishes real risk from noise, which is the entire point of moving past CVSS-only gating in the first place.
Your organization detects unauthorized use of an HSM root key. Describe the forensic investigation steps, how to assess the scope and impact of the compromise on CI/CD pipelines and signing processes, and define a recovery and key-rotation strategy that preserves trust where possible.
Sample Answer
Unauthorized use of an HSM (hardware security module) root key is one of the most severe possible findings in a signing pipeline, since the root key is typically the trust anchor everything else in the signing chain ultimately derives from; the response has to assume the worst about scope until evidence narrows it.
Forensic investigation
Start with the HSM's own access and operation logs (most HSMs log every cryptographic operation performed, including which key, what operation, and from which authenticated client), correlating the timeline of unauthorized use against known-legitimate signing operations to identify exactly which operations were NOT initiated by an expected, authorized pipeline. Cross-reference against network logs and authentication logs for the systems that have legitimate access to the HSM, looking for an unexpected authentication source or an authentication pattern (time of day, request volume) inconsistent with normal pipeline behavior.
Assessing scope and impact
Every artifact signed using the root key (or a key derived from it) during the window of unauthorized access has to be treated as potentially untrustworthy, not just the specific artifact that first drew attention; this means enumerating every signature produced during that window against the artifact registry and treating each one as needing re-verification or re-signing. If the root key signs intermediate keys rather than artifacts directly (a common PKI pattern), the scope assessment has to extend to everything trusted transitively through any intermediate key the root key issued or could have issued during the compromise window.
Recovery and key rotation, preserving trust where possible
The root key itself must be revoked and replaced; because it's a root of trust, this cascades: every intermediate certificate it issued needs to be re-issued from the new root, and every previously-signed artifact that relied on the old root's trust chain needs re-signing or an explicit, published transition plan customers and downstream consumers can follow (a documented key-rotation event, with the old root's revocation and the new root's public key published through the same trusted channel customers already use to verify your signatures). Where feasible, maintain the OLD root as revoked-but-documented (rather than silently disappearing) so downstream systems that cached the old root can be updated deliberately rather than suddenly failing verification with no explanation.
Trade-offs
Treating every signature from the compromise window as suspect, rather than trying to selectively determine which specific signings were the attacker's versus legitimate, is the conservative and correct choice here, even though it means re-signing artifacts that may well have been signed legitimately during that same window; the alternative, trying to cherry-pick which signings to trust, risks leaving a genuinely attacker-signed artifact in circulation because it was mistakenly judged legitimate.
Unlock Full Question Bank
Get access to all Secure Software Delivery: DevSecOps, Pipeline, and Supply Chain Security interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.