Identity, Authentication, and Access Management Questions
Designing and operating identity and access control systems. Covers authentication protocols and standards (OAuth, SAML, OIDC, MFA), authorization models (RBAC, ABAC), identity lifecycle and privilege management, IAM architecture and automation, and access control across cloud and on-premises environments. The 'who can do what' control plane, distinct from cryptographic key management.
Propose a defense-in-depth architecture to prevent broken authentication logic. Include recommendations for centralizing authentication and authorization, canonicalizing inputs, using nonces/CSRF tokens, consistent error handling, secure defaults, and CI/testing gates to catch regressions.
Sample Answer
Direct answer
Broken authentication logic is almost never one big bug; it is usually a dozen small, scattered checks (an ad-hoc session check here, a slightly different permission check there, a login endpoint with its own bespoke error handling) that individually look reasonable and collectively leave gaps an attacker can find by testing enough edge cases. A defense-in-depth architecture against this treats authentication and authorization as a single, centralized, reusable service rather than logic re-implemented per endpoint, canonicalizes every identity-bearing input before it is compared or matched, uses nonces and CSRF tokens to make replay and forgery structurally harder, standardizes error handling so failures never leak which specific check failed, chooses secure defaults so a missing configuration fails closed rather than open, and backs all of it with CI and testing gates that catch a regression before it reaches production rather than relying on manual review to notice it.
Structured elaboration
Centralizing authentication and authorization. The single highest-leverage architectural decision is making authentication (who is this) and authorization (what can they do) shared, mandatory library or middleware calls that every endpoint goes through, rather than logic each team reimplements. When every service independently writes its own "check if the user is logged in and has permission" code, inconsistencies are inevitable: one endpoint checks a session flag, another checks a JWT claim, a third forgets to check anything at all on an internal-facing route that later gets exposed publicly by a routing change. A centralized authorization service or middleware, ideally enforced at a single choke point (a gateway, a shared decorator/filter applied by policy rather than by convention, or a policy engine like a PDP, policy decision point, queried by a PEP, policy enforcement point, at the edge of every service) means a fix or a new check applies everywhere at once, and a new endpoint is secure by construction rather than by the author remembering to add a check.
Canonicalizing inputs. Any identity-bearing value used in a comparison, a username, an email address, a redirect URL used in an OAuth flow, a resource identifier used in an authorization check, must be normalized to one canonical form before comparison, or an attacker can exploit the gap between two different-looking representations of the same logical value. Classic examples: case-insensitive email comparison done inconsistently (registering Admin@example.com when admin@example.com already exists), path values with different encodings or trailing slashes bypassing a string-match-based authorization check, or Unicode normalization differences making two visually identical usernames compare as distinct. Centralizing this normalization in the same shared layer as authentication and authorization (rather than trusting each endpoint to normalize consistently) closes an entire class of comparison-bypass bugs at once.
Nonces and CSRF tokens. A nonce (a number used once) inside login, password reset, or state-changing requests prevents replay: a captured, valid request cannot simply be resubmitted, because the server tracks which nonces have already been consumed. CSRF tokens serve the related but distinct purpose of ensuring a state-changing request actually originated from the application's own page, not a forged cross-site submission riding on the user's ambient session cookie. Both are cheap, well-understood primitives that close specific, well-known gaps, and their absence is one of the most common findings in a broken-authentication audit precisely because they are easy to forget on a hand-rolled endpoint that bypasses the centralized flow.
Consistent error handling. Authentication and authorization failures must return the same response regardless of the actual reason for failure: "invalid username or password" for both a nonexistent account and a wrong password (never revealing which), and a generic "not authorized" or "not found" for a resource the user cannot access (never distinguishing "you don't have permission" from "this doesn't exist," which itself leaks whether the resource exists to someone probing IDs they should not be able to see). Inconsistent error messages, or worse, inconsistent response timing between different failure paths, are one of the most common ways scattered, per-endpoint authentication logic quietly becomes an information-disclosure or account-enumeration vector even when no single check is individually broken.
Secure defaults. Every configuration surface, a new route registered without an explicit authorization policy, a feature flag rolled out mid-migration, a fallback path taken when a downstream permission service times out, must default to deny, not allow. A system where "no policy configured" means "anyone can access this" turns every future omission into a live vulnerability; a system where the same omission means "nobody can access this until a policy is explicitly granted" turns an omission into a support ticket instead of a breach. This is the architectural expression of least privilege applied to the authorization layer's own failure modes, not just to the permissions it grants.
CI and testing gates to catch regressions. Centralizing logic and choosing secure defaults only stays true over time if the pipeline actively verifies it: automated tests that assert an unauthenticated request to every registered route is rejected (a "deny by default" regression test that fails the build the moment a new route is added without going through the shared middleware), static analysis or a linter rule that flags any endpoint bypassing the centralized authorization decorator, and periodic authenticated-vs-unauthenticated fuzzing of the route table as part of the deployment pipeline, not a manual security review that happens quarterly. This is what turns a defense-in-depth design from a one-time audit finding into a durable property of the codebase.
Worked example
A concrete regression this architecture is built to catch: a team adds a new internal reporting endpoint, GET /internal/reports/{id}, intending it to be reachable only from the internal network, and skips the standard authorization middleware because "it's internal, the network boundary handles it." Six months later, a routing change during a migration to a shared API gateway exposes /internal/* publicly by accident, and this endpoint, having never gone through the centralized authorization check, is now reachable by anyone who can guess or enumerate report IDs, no login required at all.
Trace how each layer of the proposed architecture would have caught or prevented this:
- Centralization would have made the endpoint's authorization non-optional: if every route must be registered through the shared middleware to be routable at all (rather than authorization being an opt-in decorator a developer can forget), there is no code path that reaches the handler without a permission check running first.
- Secure defaults mean that even if the endpoint were technically registered without an explicit policy, the default behavior is deny, so the accidental public exposure would return "not authorized" rather than the report data.
- A CI regression gate (an automated test asserting every route in the route table requires authentication unless explicitly allowlisted) would fail the build the moment this endpoint was added without the middleware, catching the gap before the six-month gap between introduction and exploitation ever opened.
- Consistent error handling means that even during the window before the gate exists, an attacker probing report IDs sees a uniform "not authorized" response for both existing reports they cannot access and nonexistent report IDs, rather than a distinguishable 404 vs 403 that would let them enumerate which IDs are real.
Trade-offs and pitfalls
- Centralization has a real engineering cost: a shared authentication/authorization layer becomes a critical-path dependency for every request, and a bug or outage in that layer now affects the entire system at once rather than one endpoint; this is a deliberate trade of blast radius concentration for consistency, and it needs its own reliability investment (caching, graceful degradation that still fails closed, not open) to be worth making.
- "Fail closed" defaults can create availability incidents if the permission-checking dependency itself becomes unreliable; a downstream policy service outage that causes every request to be denied is a real operational cost of choosing secure-by-default over available-by-default, and needs to be an explicit, accepted trade-off, not a surprise the first time it happens.
- CSRF tokens and nonces are frequently added to the main login flow but forgotten on secondary flows (password reset, account recovery, admin impersonation/support tooling), which is exactly the scattered-logic problem this architecture exists to prevent; a layered security review has to explicitly enumerate every state-changing entry point, not just the primary one.
- CI gates that check for the presence of a decorator or middleware call are a proxy, not a guarantee, of correct authorization; a route can technically call the shared middleware and still pass the wrong resource identifier or scope into it, so testing gates should assert actual behavior (an unauthenticated or wrongly-scoped request against the route returns a denial) rather than only checking that some authorization code path was invoked.
- The common failure mode this whole design targets is not one dramatic vulnerability but attrition: any one endpoint that quietly bypasses the shared pattern, for a deadline, for a "just internal" assumption, for a legacy integration, reintroduces the exact scattered-logic risk the architecture is meant to eliminate, which is why the CI gate matters as much as the initial design.
Design roles and granular permissions for an HR application so that no single user can both create employees and approve payroll (separation of duties). Describe role templates, the atomic permissions set you would model, how to represent SoD constraints in the policy engine and UI, and how to detect and remediate SoD violations during access reviews.
Sample Answer
Direct answer
Build the HR application's permission model from small, single-purpose atomic permissions rather than a handful of broad roles, compose role templates from those atomic permissions, and encode "no identity may hold both employee:create and payroll:approve" as an explicit separation of duties (SoD, the rule that certain permission combinations must never be held by the same identity, because either half alone is safe but the combination lets one person both create a fraudulent record and approve payment against it) constraint enforced in the policy engine, not only in the UI. The subtle part senior candidates get right is that the dangerous combination is usually assembled gradually across two separate role grants over time, not created by one obviously-risky role, so detection has to check an identity's accumulated permission set, not each role grant in isolation.
Structured elaboration
Atomic permission set. Decompose the domain into single-purpose permissions rather than task-shaped roles: employee:create, employee:read, employee:update, employee:terminate, payroll:submit, payroll:approve, payroll:read, payroll:reconcile, audit:read, access:review. Each permission maps to exactly one action on exactly one resource type, so a conflict rule can name the two specific permissions that must never co-occur instead of trying to reason about two broad roles that each happen to bundle many actions.
Role templates. Compose templates from those atomic permissions, keeping the conflicting halves in disjoint templates:
| Role template | Permissions granted |
|---|---|
| HR Data Entry | employee:create, employee:read, employee:update |
| HR Manager | employee:read, employee:update, payroll:submit |
| Payroll Processor | payroll:submit, payroll:read, payroll:reconcile |
| Payroll Approver | payroll:approve, payroll:read, audit:read |
| Compliance Auditor | audit:read, access:review |
No single template contains both employee:create and payroll:approve. That is necessary but not sufficient: nothing in the template design stops an administrator from later assigning both HR Data Entry and Payroll Approver to the same person, which is exactly the accumulated-combination risk called out above.
Representing SoD constraints in the policy engine. Encode the conflict as an explicit deny rule evaluated against an identity's full resolved permission set (the union across every role currently assigned to them), not against one role assignment at a time, expressed as policy-as-code so the rule lives in version control and is auditable like any other code change:
# policy-as-code SoD rule (illustrative; Rego-style deny-overrides)
deny["SoD violation: employee:create + payroll:approve"] {
input.subject.effective_permissions[_] == "employee:create"
input.subject.effective_permissions[_] == "payroll:approve"
}
This rule must run at two points: at grant time (block the assignment before it takes effect) and continuously against the current state (catch a conflict that arises from two separately-approved, individually-innocuous grants).
Representing SoD constraints in the UI. The assignment UI calls the same policy-engine check before allowing a save, and shows the conflicting existing grant inline ("this user already holds Payroll Approver, which conflicts with employee:create") rather than a generic error. A "what-if" preview lets an administrator simulate a role combination before committing it. Critically, the UI check is a convenience, not the enforcement boundary: anyone with direct API or admin-console access must hit the same policy-engine deny rule, or the control is bypassable by construction.
Detecting and remediating SoD violations during access reviews. Detection is a periodic query that computes each identity's union of permissions across all current role assignments and checks it against the conflict matrix, explicitly including violations that were never granted in one action, for example a query joining a role-assignment table to itself to find any subject with rows granting both employee:create and payroll:approve regardless of which two role assignments produced each half, and regardless of how far apart in time they were granted. Remediation is staged: flag and open a ticket, require the resource owner to justify or revoke one half within an SLA, and for a rare genuine business need to hold both temporarily (common in a very small team), require a compensating control, such as mandatory independent review of every payroll batch touched by that identity, with an explicit expiry date on the exception so "temporary" cannot quietly become permanent.
Worked example
A quarterly access review query surfaces this: user jsmith was granted HR Data Entry on 2024-03-01 to help with a hiring surge, and separately granted Payroll Approver on 2024-11-15 after moving teams; nobody re-ran the SoD check at the second grant because it was approved as an unrelated request. The union of their current permissions includes both employee:create and payroll:approve, a live violation that neither grant alone would have triggered. Remediation: the reviewer revokes employee:create (the role that no longer matches their current job function) rather than the more recently and deliberately granted payroll:approve, closes the ticket, and the policy engine's continuous check confirms the violation no longer exists.
For auditor evidence, rather than asserting "we reviewed the org chart and found no conflicts," a stronger artifact combines three things: the versioned conflict-matrix definition itself (so the auditor can see exactly what was checked), the access-review certification record showing when this specific combination was checked and by whom, and privileged access management (PAM, the system brokering and logging privileged sessions) session logs showing that even during any period where a conflicting pair of permissions existed on paper, no single identity's logged session both created an employee record and approved a payroll run against that same record. That transaction-level proof is materially stronger than a policy-compliance statement alone, because it demonstrates the control held even during the gap before the violation was caught.
Trade-offs and pitfalls
Decomposing permissions too finely creates a combinatorial explosion of role templates and makes the conflict matrix itself hard to maintain; decomposing too coarsely (bundling payroll:submit and payroll:approve into one "Payroll" permission, say) makes the SoD rule impossible to express at all, because the very actions that must be separated no longer exist as separate grants. The atomic set above is sized to the specific conflict being enforced, not to some abstract ideal of granularity.
The single biggest pitfall is checking SoD only at the moment of a single role grant instead of against the accumulated permission set: two individually-approved, individually-reasonable role assignments can combine into a violation that neither approver saw, which is exactly what the worked example shows. A second pitfall is enforcing the rule only in the UI: any direct database or admin-API path around the UI silently defeats the entire control unless the policy engine itself is the enforcement point. A third is granting a "temporary" compensating-control exception with no expiry: without a forced re-review date, the exception becomes the new permanent state and the compensating control (the extra review step) usually erodes in practice long before anyone notices the underlying grant was never actually temporary.
Create a PowerShell solution (outline or code) to collect the local 'Administrators' group membership from every domain-joined computer in an OU, identify non-approved users, and produce a CSV report with computer name, account, SID, and whether the account is a domain or local account. Describe remoting and permission requirements.
Sample Answer
Direct answer
Enumerate the computers in the target organizational unit (OU) from Active Directory, remotely query each machine's local "Administrators" group membership, resolve each member to a security identifier (SID, the unique identifier Windows uses to represent an account, independent of its display name) so renamed or ambiguous accounts cannot slip past a name-only comparison, classify each member as a domain or local account, and compare against an explicit allowlist before writing the CSV. The part that is easy to get wrong is the comparison itself: checking only by name misses an account that has been renamed, so the allowlist check has to match on SID as well as name.
Structured elaboration
Approach. For each computer, remotely enumerate local group membership (in production, via Invoke-Command calling the WinNT ADSI provider, Get-LocalGroupMember on modern Windows, or the Win32_Group/Win32_GroupUser WMI classes), translate each member's account name to a SID, classify domain versus local by whether the account's prefix matches the local computer name or BUILTIN, and flag any member whose name and SID both fail to match the allowlist. Offline or unreachable computers are caught and reported as an explicit error row rather than silently omitted from the CSV, so a gap in the report is visible instead of looking like a clean result.
Key points.
- Compare against the allowlist by both name and SID, not name alone: a domain account that has been renamed keeps the same SID, so a name-only check can both miss a renamed unapproved account and false-flag a renamed approved one.
BUILTIN\Administratorand any account prefixed with the target computer's own name are local accounts; anything else is a domain account, which matters for triage (a local account added directly on one machine bypasses group-policy-managed domain group membership entirely).- An unreachable computer must produce a visible error row, not a silent gap, since "the report has fewer rows than expected because three machines were offline" and "three machines are clean" look identical unless the report says otherwise.
$AllowList = @("CONTOSO\CorpAdmin", "S-1-5-32-544", "BUILTIN\Administrator", "S-1-5-32-544-500")
# Fixed SID lookup table standing in for NTAccount.Translate([SecurityIdentifier]);
# real deployments resolve these from the domain/local SAM, not a literal map.
$SidTable = @{
"CONTOSO\CorpAdmin" = "S-1-5-21-111-222-333-1001"
"BUILTIN\Administrator" = "S-1-5-32-544-500"
"CONTOSO\jdoe" = "S-1-5-21-111-222-333-1042"
"CONTOSO\svc-backup" = "S-1-5-21-111-222-333-2099"
}
# In production: $computers = Get-ADComputer -SearchBase $OU -Filter * | Select -Expand Name
# then Invoke-Command -ComputerName $c -ScriptBlock { (Get-LocalGroupMember Administrators) }
# Below, MockFleet models exactly that per-computer result for three representative machines,
# including one offline host to exercise the error-handling branch.
$MockFleet = @(
[pscustomobject]@{ Computer="WKS01"; Reachable=$true; Members=@("CONTOSO\CorpAdmin","BUILTIN\Administrator","CONTOSO\jdoe") }
[pscustomobject]@{ Computer="WKS02"; Reachable=$true; Members=@("CONTOSO\CorpAdmin","BUILTIN\Administrator","CONTOSO\svc-backup") }
[pscustomobject]@{ Computer="WKS03"; Reachable=$false; Members=@() }
)
function Get-AccountType {
param([string]$Account, [string]$ComputerName)
if ($Account -like "$ComputerName\*") { return "Local" }
if ($Account -like "BUILTIN\*") { return "Local" }
return "Domain"
}
$Results = foreach ($host_ in $MockFleet) {
if (-not $host_.Reachable) {
[pscustomobject]@{ Computer=$host_.Computer; Account="ERROR"; SID=""
AccountType=""; Approved="Error: WinRM connect failed (host unreachable)" }
continue
}
foreach ($m in $host_.Members) {
$sid = if ($SidTable.ContainsKey($m)) { $SidTable[$m] } else { $null }
$type = Get-AccountType -Account $m -ComputerName $host_.Computer
$approved = (($AllowList -contains $m) -or ($sid -and ($AllowList -contains $sid)))
[pscustomobject]@{ Computer=$host_.Computer; Account=$m; SID=$sid
AccountType=$type; Approved=if ($approved) { "Yes" } else { "No" } }
}
}
$csvPath = "./LocalAdminsReport.csv"
$Results | Export-Csv -Path $csvPath -NoTypeInformation
Get-Content $csvPath
$flagged = $Results | Where-Object { $_.Approved -eq "No" }
Write-Output "Flagged (non-approved) accounts: $($flagged.Count)"
$flagged | ForEach-Object { Write-Output (" {0}\{1}" -f $_.Computer, $_.Account) }
Worked example
Run unmodified with pwsh, this produces:
"Computer","Account","SID","AccountType","Approved"
"WKS01","CONTOSO\CorpAdmin","S-1-5-21-111-222-333-1001","Domain","Yes"
"WKS01","BUILTIN\Administrator","S-1-5-32-544-500","Local","Yes"
"WKS01","CONTOSO\jdoe","S-1-5-21-111-222-333-1042","Domain","No"
"WKS02","CONTOSO\CorpAdmin","S-1-5-21-111-222-333-1001","Domain","Yes"
"WKS02","BUILTIN\Administrator","S-1-5-32-544-500","Local","Yes"
"WKS02","CONTOSO\svc-backup","S-1-5-21-111-222-333-2099","Domain","No"
"WKS03","ERROR","","","Error: WinRM connect failed (host unreachable)"
Flagged (non-approved) accounts: 2
WKS01\CONTOSO\jdoe
WKS02\CONTOSO\svc-backup
WKS01\CONTOSO\jdoe and WKS02\CONTOSO\svc-backup are correctly flagged as non-approved, a domain account added to local admins on one machine without going through the approved-group process; CONTOSO\CorpAdmin and BUILTIN\Administrator correctly pass on both machines that report in; and WKS03, which is offline, produces an explicit ERROR row rather than silently vanishing from the CSV, exactly the visible-gap behavior the "key points" section calls for. Note that if the SID lookup for CONTOSO\jdoe had, purely hypothetically, matched an entry in $AllowList by SID even though the name did not, the account would correctly show Approved=Yes, which is precisely why the comparison checks both name and SID rather than either alone.
Complexity. For c computers with an average of m local admin members each, the remote enumeration is O(c) round-trips (one per computer, each returning its own membership in one call), and the allowlist comparison is O(c⋅m) simple lookups against a hash-backed allowlist, which is cheap even for a large OU; the actual bottleneck in a real environment is network round-trip latency across potentially thousands of computers, not the comparison logic.
Edge cases. An unreachable computer must be caught and reported, not allowed to throw and abort the whole run; a member whose SID cannot be resolved (a deleted domain account still listed in a local group) should still appear in the report with a blank or best-effort SID rather than being silently skipped, since an orphaned SID in a privileged local group is itself a finding worth surfacing; and a renamed approved account must still resolve as approved via its SID even though its current display name no longer matches the allowlist's name entry.
Trade-offs and pitfalls
Remoting to every computer in an OU serially does not scale past a few hundred machines in a reasonable window; production runs typically parallelize with a bounded throttle (Invoke-Command -ThrottleLimit) or route through an existing management channel (SCCM, Intune, or a similar endpoint-management tool) rather than opening a fresh WinRM session to every host from a single script.
Remoting and permission requirements. PowerShell Remoting (WinRM) must be enabled and reachable on every target computer, and the account running the script needs rights to query local group membership on each target, typically local administrator rights or a delegated equivalent. Watch for the Kerberos "double hop" problem: if the remote command itself needs to reach back out to Active Directory (for example, to resolve a SID against the domain rather than a local cache), the credential used to connect does not automatically carry forward to that second hop unless you use CredSSP, resource-based constrained delegation, or a CIM session with an explicitly provided credential.
A common pitfall is trusting name-only comparison against the allowlist, which both of the flagged accounts above would still be correctly caught by, but which would wrongly clear a renamed unapproved account or wrongly flag a renamed approved one; matching on SID as well is what makes the comparison correct under renames. A second pitfall is letting an unreachable computer fail silently: a script that throws and stops on the first offline host, or one that simply omits offline hosts from the output, produces a report that looks complete but is not, which is worse than an explicit error row because nobody investigating the gap knows to look for it.
Design detection and alerting rules (for auditd + SIEM like Splunk/ELK) to detect suspicious user-administration activity: additions to /etc/sudoers or /etc/sudoers.d, changes to group 'sudo' membership, new SSH keys written to home directories, and creation of UID 0 users. Provide example auditd rules or file watches and high-level SIEM query patterns and thresholds to reduce false positives.
Sample Answer
Direct answer
auditd (the Linux kernel's auditing subsystem daemon) can watch a fixed set of file paths for reads, writes, and attribute changes and tag each event with a searchable key, which covers three of the four things this question asks for directly: sudoers edits, group-file changes, and passwd/UID-0 creation all show up as writes to a small number of well-known files. The fourth, new SSH keys, is the one that needs a design decision rather than a single rule, because auditd watches paths, not filename patterns, and an authorized_keys file can live inside any user's home directory.
Structured elaboration
Additions to /etc/sudoers or /etc/sudoers.d. A direct write-watch on both catches any edit, whether made through visudo or a direct file write:
-w /etc/sudoers -p wa -k sudoers_changes
-w /etc/sudoers.d/ -p wa -k sudoers_changes
-p wa watches for writes and attribute changes (not reads), and -k attaches a key string so the SIEM can query on sudoers_changes directly instead of matching on the raw path.
Changes to sudo group membership. Local group membership lives in /etc/group (and the shadowed password half in /etc/gshadow), so watching both catches a direct edit to the sudo group's member list:
-w /etc/group -p wa -k group_changes
-w /etc/gshadow -p wa -k group_changes
This only catches membership stored locally. If the environment resolves group membership through a directory service (an LDAP- or Active-Directory-integrated NSS module), the authoritative change happens on the directory side, not in a local file, and this rule sees nothing; that gap has to be closed by the directory's own audit logging, not by auditd on the Linux host.
New SSH keys written to home directories. auditd watches a fixed path, not a glob, so there is no single rule for "any authorized_keys file under any home directory." The practical options are a broad directory watch on the home-directory root, accepting that it fires on any write anywhere under it and needs a downstream filename filter, or enumerating each known home directory into its own watch, accepting the maintenance burden of updating rules as accounts are added:
-w /home -p wa -k home_dir_writes
The SIEM-side query then narrows the resulting stream to events whose path ends in .ssh/authorized_keys or .ssh/authorized_keys2, since that filtering has to happen after the event is captured, not in the auditd rule itself.
Creation of UID 0 users. auditd records that /etc/passwd was written, not what changed inside it, so detecting a new UID 0 account needs two complementary signals: a write-watch on the file itself, and a syscall-level rule on the command that creates or modifies accounts, so the SIEM can inspect the actual arguments:
-w /etc/passwd -p wa -k passwd_changes
-a always,exit -F arch=b64 -S execve -F exe=/usr/sbin/useradd -k user_admin_exec
-a always,exit -F arch=b64 -S execve -F exe=/usr/sbin/usermod -k user_admin_exec
The SIEM then correlates a user_admin_exec event whose captured arguments include -u 0 (or -o -u 0, since -o allows a duplicate/non-unique UID) with the following passwd_changes write, and treats a resulting UID field of 0 for any account other than the pre-existing root line as a near-certain finding, since a legitimate second UID-0 account is created vanishingly rarely in ordinary operations.
High-level SIEM query patterns and thresholds. For sudoers_changes and group_changes, correlate the event with whether it happened through the organization's change-management tooling (a known automation service account, during a recorded maintenance window) versus an interactive human session outside one; alert immediately on the latter and suppress or batch the former, since routine configuration-management-driven edits are common and a threshold based on volume alone would either miss a single unauthorized edit or drown the analyst in expected noise. For home_dir_writes filtered to authorized_keys paths, apply the same actor-based filter: a key rotation performed by the known configuration-management identity is routine, while the same file written by an interactively logged-in session is comparatively rare and worth a low-volume, near-zero-tolerance alert threshold. For the UID-0 correlation, the threshold is effectively "greater than zero, ever," since there is essentially no legitimate volume of new UID-0 account creation to distinguish from noise.
Worked example
Consider a host where the configuration-management agent rotates SSH keys under every user's ~/.ssh/authorized_keys every night as part of routine key hygiene, generating dozens of home_dir_writes events tagged to that agent's own service-account identity. One evening, a home_dir_writes event fires for /home/deploy/.ssh/authorized_keys attributed to an interactive session under a human user's login, not the configuration-management agent. That single event, filtered by the actor-identity rule above, is exactly the rare, high-signal case the design is built to surface: the routine nightly rotations from the known automation identity are suppressed from paging anyone, while this one write from an unexpected actor triggers an immediate alert, because "someone logged in interactively and wrote to a deploy account's authorized_keys file" has essentially no benign explanation in an environment where key management is otherwise fully automated.
Trade-offs and pitfalls
The biggest limitation to be explicit about is that auditd observes local file-system events, not directory-service events; in any environment where account and group management is centralized (LDAP, Active Directory, or a configuration-management system that is itself the source of truth), the local watches above catch only changes made directly on that one host, which is a real but partial view, and the directory service's own change auditing has to cover the rest.
A second pitfall is inode-based watch fragility: a file watch in auditd is tied to the underlying inode at the time the rule was loaded, so a file that gets deleted and recreated (rather than edited in place) can silently stop being watched until the rules are reloaded; periodically re-applying the rule set (or watching the parent directory instead of the file itself where that trade-off is acceptable) avoids a watch quietly going stale without anyone noticing.
A third pitfall is threshold design that ignores actor identity: alerting on raw event volume for home_dir_writes would either miss a single malicious key addition buried among routine automated rotations, or generate so many alerts from the automation itself that analysts learn to ignore the queue; filtering by which identity performed the write, as in the worked example, is what actually makes the threshold usable rather than just present.
Perform a threat modeling exercise for an enterprise IAM platform. Identify top attack vectors (token theft, account takeover, IdP compromise, provisioning abuse, privileged escalation, lateral movement) and propose concrete mitigations, detection strategies, and compensating controls for each vector.
Sample Answer
Direct answer
A threat model for an enterprise identity and access management (IAM) platform should walk each stage where trust is established or extended, credential issuance, token use, account elevation, and inter-system access, and ask what an attacker gains at each stage and what specific control catches or blocks it. The six vectors named here span three parts of that lifecycle: the integrity of tokens and the identity provider (IdP) that issues them, the moment an identity is created or elevated, and what an attacker does after gaining an initial foothold.
Structured elaboration
This applies STRIDE-style reasoning (Spoofing, Tampering, Repudiation, Information disclosure, Denial of service, Elevation of privilege, the standard threat-categorization lens) directly to the IAM platform rather than teaching the methodology itself. For each vector: what the attacker actually does, the primary preventive mitigation, how you would detect it, and a compensating control that limits damage if the primary mitigation is absent or fails.
| Vector | Attacker action | Mitigation | Detection strategy | Compensating control |
|---|---|---|---|---|
| Token theft | Steals a valid, unexpired token via cross-site scripting (XSS), insecure client storage, or a malicious browser extension | Short token lifetimes; sender-constrained tokens (mutual TLS or DPoP, Demonstrating Proof-of-Possession, so a stolen token cannot be replayed from a different client); store tokens in httpOnly cookies, not scriptable storage | Same token used from two different IP addresses or user agents in a short window; impossible-travel pattern between two token uses | Fast revocation via a token-introspection endpoint or short-lived-token expiry, plus step-up authentication required for sensitive actions even inside an already-authenticated session |
| Account takeover | Gains control of a user's identity via a phished password, phished push-based multi-factor approval, or SIM-swap-based SMS interception | Phishing-resistant authentication (FIDO2/WebAuthn hardware-bound passkeys) preferred over SMS or push-based multi-factor authentication (MFA); MFA required on every account | New-device or new-location login alerting; an unusual action sequence immediately after login, such as a bulk data export or an MFA-method change | Risk-based step-up authentication on sensitive actions regardless of how the session began, and session-level anomaly monitoring able to force mid-session re-authentication |
| IdP compromise | Compromises the identity provider itself: its signing key, its admin console, or a federation trust configuration; the highest blast-radius vector, since it can mint a valid token for any identity | Hardware security module (HSM)-backed signing keys so private key material is never directly exposed even to IdP administrators; the IdP's own admin accounts get the strongest privileged access management (PAM) and MFA treatment of any account in the environment | Monitoring the IdP's own admin audit log for configuration changes (a new federation trust added, a signing key exported or rotated unexpectedly); anomaly detection on token-issuance volume | Short-lived tokens bound the maximum damage window even if a signing key is compromised, paired with a rehearsed emergency key-rollover runbook so the actual rollover takes minutes, not days |
| Provisioning abuse | A malicious or coerced actor abuses the account-creation or entitlement-granting workflow itself, for example through SCIM (System for Cross-domain Identity Management, the standard protocol many IdPs use to auto-provision downstream apps), rather than compromising an existing account | Dual-control approval on any provisioning action granting elevated entitlements, so no single actor can both request and approve; the provisioning system itself is treated as a privileged system | Alerting on provisioning events without a matching change ticket; periodic reconciliation between the HR system of record and actual granted entitlements | Periodic access review and attestation, a named owner actively re-certifying who has access on a fixed cadence, catches an abusively-provisioned account even if the initial detection missed it |
| Privileged escalation | A foothold in a lower-privileged account or system is used to reach a higher-privileged one, via excessive standing permissions or a flaw in authorization logic | Least privilege by default plus just-in-time (JIT) elevation instead of standing privileged access, so there is no permanently-elevated credential sitting around to escalate into | Alerting on the elevation event itself (a JIT request, an addition to a privileged group), correlated against whether the requesting identity's recent behavior looks anomalous | Session recording and brokering through the privileged access management layer, so a successful escalation is fully observed and time-boxed rather than open-ended |
| Lateral movement | Uses one compromised identity's access to reach additional systems, most dangerous when one credential or broadly-trusted identity is valid everywhere | Segmented workload and service identities: short-lived, narrowly-scoped credentials per system rather than one shared service account reused across many systems | Correlating a single identity's access pattern across multiple systems in a short window against its historical baseline | Distinct credentials and scopes per trust boundary mean reaching one system with a stolen identity does not automatically grant reachability to the next |
Worked example
A realistic chained attack shows why treating these six vectors in isolation understates the real risk. An attacker phishes a push-based MFA approval from a standard user (account takeover). From that lower-privileged foothold, they discover a service account with excessive standing permissions, including access to the IdP's admin console, and use it to escalate (privileged escalation, enabled by the absence of just-in-time elevation). With admin access to the IdP, they attempt to add a new federation trust so their own external identity provider is accepted as authoritative (an IdP compromise attempt). Reading this chain against the table above: the account-takeover step should have been caught by new-device login alerting; if it was not, the privileged-escalation step should have been caught by alerting on the elevation event itself, since a standard user reaching admin-console access is a clear deviation from baseline; if that was also missed, the IdP's own admin audit log monitoring for a newly-added federation trust is the last line before the attacker has durable, org-wide token-minting capability. No single control in the table is expected to be perfect, the chain is stopped by whichever layer actually catches it, which is the point of listing detection strategies at every stage rather than only at the first one.
Trade-offs and pitfalls
- Treating each vector as independent understates chained risk. As the worked example shows, a weak mitigation at one stage (no JIT elevation, so standing over-permissioned service accounts exist) turns a low-severity account takeover into a high-severity IdP compromise attempt. A mature threat model reviews chains across vectors, not just each row of the table in isolation.
- Detection-only coverage for the IdP-compromise vector is not enough given its blast radius. Because a compromised IdP can mint tokens for any identity, this is the one vector where the compensating control (short token lifetime plus a rehearsed rollover runbook) matters as much as the primary mitigation; relying purely on detecting the compromise after the fact leaves too large a window of full-organization exposure.
- Just-in-time elevation without session recording only half-solves privileged escalation. JIT reduces the window an elevated credential exists, but without session recording and brokering, a successful escalation inside that window is still unobserved; the two controls are complementary, not substitutes.
- A common wrong turn is treating provisioning abuse as purely a technical control problem. Dual-control approval workflows help, but the compensating control that actually catches a determined insider or a coerced approver is the human process of periodic access review, a technical gate alone does not substitute for someone actively re-certifying access on a cadence.
Unlock Full Question Bank
Get access to all 22 Identity, Authentication, and Access Management interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.