Microsoft Systems Administrator (Mid-Level) Interview Preparation Guide
Microsoft's Systems Administrator interview process typically consists of a recruiter screening call, followed by technical phone screening, and onsite interviews covering systems knowledge, infrastructure design, troubleshooting, and cultural fit. The process emphasizes hands-on infrastructure management experience, problem-solving under pressure, and alignment with enterprise IT best practices.
Interview Rounds
Recruiter Screening
What to Expect
Initial 20-30 minute call with a recruiter to discuss your background, career goals, and fit for the Systems Administrator role. This round covers your experience with infrastructure management, reasons for considering Microsoft, and logistics of the interview process. Not technical—focuses on motivation and baseline qualifications.
Tips & Advice
Be clear about your systems administration experience and specific achievements in managing infrastructure. Explain why you're interested in working at Microsoft and how this role aligns with your career growth. Ask thoughtful questions about the team and role. Have your resume and key projects documented clearly. Show enthusiasm for learning enterprise-scale infrastructure practices.
Focus Topics
Motivation for Systems Administrator Role at Microsoft
Why you're interested in this specific role, what attracts you to Microsoft's infrastructure operations, and how it fits your career trajectory.
Practice Interview
Study Questions
Key Infrastructure Projects & Achievements
Specific examples of infrastructure projects you've owned or contributed to, challenges overcome, and measurable outcomes.
Practice Interview
Study Questions
Career Background & Infrastructure Management Experience
Your professional journey as a systems administrator, types of infrastructure managed, scale of environments, and progression in the field.
Practice Interview
Study Questions
Technical Phone Screen
What to Expect
45-60 minute technical conversation with a systems engineer or senior administrator covering infrastructure fundamentals, Windows Server knowledge, Active Directory concepts, networking basics, and troubleshooting approach. Expect practical scenario-based questions about managing systems at scale.
Tips & Advice
Speak confidently about your hands-on experience with Windows Server administration and Active Directory. Use specific terminology (Group Policy, DNS, DHCP, domain controllers) but explain concepts clearly. When given a scenario, think out loud and ask clarifying questions. Demonstrate troubleshooting methodology: gather information, identify root cause, implement solution, verify. Be prepared to discuss your experience with system monitoring tools, patch management, and backup strategies. Show familiarity with PowerShell as an automation tool.
Focus Topics
Backup, Recovery & Business Continuity
Backup strategies (full, incremental, differential), recovery time objective (RTO) and recovery point objective (RPO), disaster recovery planning, and testing recovery procedures.
Practice Interview
Study Questions
Networking Fundamentals for Infrastructure
TCP/IP basics, DNS configuration and troubleshooting, DHCP scope management, network segmentation, IPv4 and IPv6 concepts, and how networking relates to system connectivity.
Practice Interview
Study Questions
Active Directory & Group Policy Management
Domain controller architecture, organizational unit structure, user account lifecycle, group policy creation and application, replication, FSMO roles, and troubleshooting domain issues.
Practice Interview
Study Questions
Infrastructure Scenario-Based Problem Solving
Responding to real-world scenarios: user lockouts, domain replication issues, service failures, performance degradation, permission problems, and network connectivity issues.
Practice Interview
Study Questions
Windows Server Administration & Core Concepts
Hands-on knowledge of Windows Server editions, server roles, feature installation, user and group management, local security policies, and remote administration.
Practice Interview
Study Questions
System Monitoring, Performance Tuning & Troubleshooting
Using Event Viewer, Performance Monitor, Task Manager, and third-party tools to diagnose issues; understanding metrics like CPU, memory, disk I/O; common bottlenecks; and systematic troubleshooting approach.
Practice Interview
Study Questions
Onsite Round 1: Systems & Infrastructure Knowledge
What to Expect
60-90 minute technical interview focused on deep understanding of Windows systems, server operations, and infrastructure components. Expect detailed questions about system architecture, configuration, security hardening, and hands-on scenarios. May include whiteboarding infrastructure diagrams or discussing architectural decisions.
Tips & Advice
This round evaluates your depth of systems knowledge at a mid-level. Go beyond surface-level knowledge; explain why certain configurations are preferred and trade-offs involved. Discuss your experience with system hardening, security best practices, and compliance requirements. Be prepared to design or critique infrastructure solutions. Walk through your thought process when analyzing complex system problems. Show awareness of cost-efficiency and performance optimization. Mention specific tools you've used and scripting experience with PowerShell.
Focus Topics
Storage, RAID & Disk Management
RAID levels and use cases, disk partitioning, volume management, Storage Spaces, iSCSI configuration, and data protection strategies.
Practice Interview
Study Questions
Virtual Machine Management (Hyper-V or similar)
Virtual machine creation and configuration, resource allocation, virtual networking, snapshot management, live migration, and capacity planning for virtualized environments.
Practice Interview
Study Questions
Infrastructure Architecture & Design Principles
Designing scalable, reliable, and maintainable infrastructure; high availability concepts; load balancing; redundancy; disaster recovery planning; and cost-benefit analysis of architectural decisions.
Practice Interview
Study Questions
System Security Hardening & Compliance
Security baselines, disabling unnecessary services, applying security updates and patches, firewall configuration, user rights assignment, audit policies, and compliance frameworks (e.g., CIS benchmarks).
Practice Interview
Study Questions
Server Roles, Features & Installation
Understanding different Windows Server roles (File Services, Print Services, Remote Desktop Services, DNS, DHCP, Hyper-V), feature dependencies, and appropriate deployment scenarios.
Practice Interview
Study Questions
Onsite Round 2: Active Directory, Group Policy & Identity Management
What to Expect
60-90 minute interview focused on Advanced Directory Services, identity management, and group policy implementation. Discuss real-world scenarios involving user lifecycle management, permission delegation, policy application challenges, and enterprise directory design.
Tips & Advice
This round goes deep into Active Directory expertise. Discuss complex scenarios you've managed: multi-forest environments, trust relationships, organizational unit design, group policy inheritance and conflicts, permission delegation, troubleshooting replication issues. Show understanding of user lifecycle (creation, modification, deletion) in enterprise contexts. Discuss security groups versus distribution groups, and when to use each. Talk about group policy testing and troubleshooting tools (gpupdate, gpresult, Group Policy Editor, Active Directory Users and Computers). Mention experience with Exchange integration if applicable. Be ready to design AD structure for hypothetical organizations.
Focus Topics
Identity & Access Management Best Practices
Privilege account management, least privilege principle, auditing user access, managing service accounts, and integration with cloud identity services (Azure AD concepts).
Practice Interview
Study Questions
Active Directory Replication & Health Monitoring
Inter-site replication configuration, connection objects, knowledge consistency checker (KCC), replication monitoring, troubleshooting replication conflicts, and domain controller health checks.
Practice Interview
Study Questions
Group Policy Implementation & Troubleshooting
Creating and linking GPOs, policy precedence and inheritance, user versus computer policies, security group filtering, loopback processing, Group Policy Preferences, and troubleshooting application failures.
Practice Interview
Study Questions
Active Directory Architecture & Domain Design
Forest and domain structure, organizational unit design, trust relationships, domain controller placement, FSMO role distribution, and replication topology.
Practice Interview
Study Questions
User & Permission Management at Scale
User account lifecycle management, delegation of control, group membership management, permission assignment strategies, and automated account provisioning/deprovisioning.
Practice Interview
Study Questions
Onsite Round 3: Networking, Infrastructure Operations & Troubleshooting
What to Expect
60-90 minute interview covering enterprise networking, connectivity issues, infrastructure operations, and real-world troubleshooting scenarios. Expect complex troubleshooting cases, network design discussions, and how systems communicate across infrastructure.
Tips & Advice
This round tests your troubleshooting depth and networking knowledge. Walk through complex real-world scenarios methodically. Discuss DNS and DHCP architectures in enterprise settings, including redundancy and failover scenarios. Talk about network segmentation and how it relates to system security. Be prepared to diagnose complex issues involving multiple components (network, systems, services). Discuss monitoring and alerting strategies. Show understanding of performance baselines and how to identify anomalies. If you've used tools like network analyzers, packet capture, or advanced monitoring solutions, mention them. Discuss lessons learned from past incidents and preventive measures taken.
Focus Topics
System Performance Optimization & Capacity Planning
Analyzing performance metrics, identifying bottlenecks (CPU, memory, disk, network), capacity planning for growth, and optimization strategies without major infrastructure changes.
Practice Interview
Study Questions
High Availability & Load Balancing Solutions
Failover clustering, Network Load Balancing, redundancy design, failover testing, and ensuring services remain available during maintenance or failure.
Practice Interview
Study Questions
Infrastructure Incident Response & Change Management
Responding to critical incidents, communication during outages, change control procedures, testing changes in non-production, and post-incident reviews.
Practice Interview
Study Questions
Complex Troubleshooting & Root Cause Analysis
Methodical troubleshooting approach, using diagnostic tools (Event Viewer, Performance Monitor, network analyzers), analyzing logs, identifying root causes in multi-component failures, and preventing recurrence.
Practice Interview
Study Questions
Enterprise Networking for Infrastructure Operations
Network segmentation, VLAN configuration, network access control, DNS and DHCP in enterprise contexts, failover scenarios, and network monitoring.
Practice Interview
Study Questions
Onsite Round 4: Behavioral, Collaboration & Cultural Alignment
What to Expect
45-60 minute behavioral interview assessing how you work with teams, handle pressure and conflicts, contribute to team success, and align with Microsoft's culture. Expect situational questions about past experiences, mentoring others, owning projects, and cross-functional collaboration.
Tips & Advice
Prepare specific examples from your infrastructure experience using the STAR method (Situation, Task, Action, Result). Discuss times you've mentored junior team members or worked with peers to resolve issues. Talk about a time you had to learn a new technology quickly and how you approached it. Discuss how you've contributed to team knowledge sharing (documentation, training sessions, etc.). Be ready to discuss how you've managed pressure during incidents or critical deployments. Show examples of collaborating across teams (networking, security, database, development). Discuss your approach to continuous learning and staying current with infrastructure trends. Show genuine interest in Microsoft's technology and culture.
Focus Topics
Continuous Learning & Technology Awareness
How you stay current with infrastructure trends, certifications pursued, communities involved in, and awareness of emerging technologies relevant to the role.
Practice Interview
Study Questions
Mentoring & Knowledge Transfer
Your experience mentoring junior team members, documenting processes, conducting training, and helping others grow in infrastructure skills.
Practice Interview
Study Questions
Handling Pressure, Conflicts & Learning Agility
Examples of high-pressure situations (critical incidents, tight deadlines), how you stayed composed and solved problems, conflicts with colleagues and resolution, and quickly learning new technologies.
Practice Interview
Study Questions
Team Collaboration & Cross-Functional Work
Working effectively with other infrastructure teams, networking teams, security teams, developers, and stakeholders; communication during incidents; sharing knowledge.
Practice Interview
Study Questions
Project Ownership & Initiative
Examples of infrastructure projects you've owned end-to-end, challenges faced, how you drove them to completion, and impact achieved.
Practice Interview
Study Questions
Frequently Asked Systems Administrator Interview Questions
You need to roll out an update to a service behind an L7 load balancer that still uses sticky sessions. Compare blue-green, canary, and rolling-update approaches: for each, explain how you would drain connections, how you would migrate or preserve session state, and how you would validate success before committing.
Sample Answer
Direct Answer
All three deployment strategies need the same underlying primitive when sticky sessions are involved: keep serving already-affinitized sessions from their old destination while steering new sessions to the new version, and validate with real traffic before fully committing. What differs is blast radius and rollback speed. Canary gives the smallest blast radius and fastest safe rollback since only a slice of traffic ever touches the new version. Blue-green gives the fastest full rollback, a single traffic flip back to the untouched old fleet. Rolling update is the middle ground: it uses less infrastructure than blue-green but exposes both versions to production traffic for the longest window.
Comparing the Three Approaches
| Approach | Connection draining | Session state handling | Validation before commit | Rollback |
|---|---|---|---|---|
| Blue-green | Old (blue) fleet stays fully up and serving until cutover; drain blue only after traffic is flipped | Best served by an externalized session store so sessions survive the fleet swap cleanly; if sessions are cookie-pinned to instances, the cutover needs a session-mirroring step | Gradual traffic shift (e.g., 10% then 100%) while watching error rate, latency, and session-continuity checks before fully committing | Flip the router back to blue; near-instant since blue was never torn down |
| Canary | New (canary) instances take only new sessions; canary is scaled down and drained gracefully if promoted or rejected | Same principle: externalized state lets a session move between canary and baseline transparently; if using cookie affinity, route by cookie so a canaried session stays on a canary instance able to handle it | Small, tight-SLO evaluation window on a small slice of real traffic, with automated gates (error rate, latency, business metrics) before expanding | Set canary traffic weight to zero and terminate; blast radius was already small, so rollback impact is minimal |
| Rolling update | Each instance is marked out of rotation, allowed to finish in-flight work (or hand off state), then replaced, in small batches | Requires session state to be externalized or explicitly handed off before an instance is replaced, since there's no single "old fleet" to fall back to | Monitor per-batch metrics with a pause-and-check gate between batches, rather than one global before/after comparison | Halt the rollout and redeploy the previous version to the already-replaced instances; slower than blue-green because it's also incremental |
Worked Example
If a canary receives 5% of instance capacity and that version has a latent bug, the fraction of live traffic that can be affected during the validation window is bounded above by that same 5%, by construction, since the routing weight is what determines exposure. A rolling update, by contrast, exposes a growing share of the fleet as each batch completes; if you roll in 10 batches of 10% each and something is only caught at the fourth batch, roughly 40% of the fleet has already been exposed by the time you halt, an order of magnitude more exposure than the canary case for the same underlying bug. This is the concrete reason canary is preferred for higher-risk changes even though it takes more total upfront tooling to run the automated gates.
Trade-offs and Pitfalls
- Blue-green doubles infrastructure cost for the duration of the cutover window, since two full fleets run simultaneously; this is the price of the fastest possible full rollback.
- Rolling update's extended window with two versions live simultaneously creates a real risk of version-skew bugs: if a sticky client reconnects mid-rollout and lands on a different version than its previous request, and the two versions disagree on session schema or API contract, that's a bug class that blue-green and canary largely avoid by keeping version boundaries sharper.
- Canary's small sample size is a double-edged sword: it bounds blast radius but also means a rare bug (one that only manifests on 1 in 1000 requests) may not surface at all before the canary is judged "clean" and promoted.
- Whichever strategy is used, connection draining timeouts need to be sized for the actual protocol involved; a WebSocket-heavy service needs a much longer drain window than a typical request-response API, and a drain window that's too short converts a graceful strategy into a de facto hard cut for anything still in flight.
A bug only appears under heavy load and cannot be reproduced locally. Describe how to build a deterministic experiment or test harness to reproduce the issue: synthetic traffic generators, seeding state, concurrency controls, time manipulation, and deterministic schedulers. Explain how to minimize noise and prove causality.
Sample Answer
A bug that only appears under heavy load needs a test harness that can manufacture load deterministically, since waiting for production traffic to happen to trigger it again is not a repeatable investigation.
Building the harness
- Synthetic traffic generation at a controlled, repeatable concurrency and request mix (a load-testing tool driving realistic request shapes, not just raw throughput).
- Seed shared state deterministically (fixed starting data, fixed random seeds where the system uses randomness) so the only varying factor between runs is scheduling/timing, not also data variance.
- Concurrency controls and deterministic schedulers where available, to bias the interleaving toward the suspected contention point rather than relying purely on load volume to eventually hit it.
- Time manipulation (accelerating or controlling clock-dependent logic) when the bug involves timers, TTLs, or scheduled work, so you don't need to wait real-world hours to observe a time-triggered condition.
- Minimize noise, prove causality: run the same load profile with and without the suspected contributing factor (a specific code path enabled/disabled, a specific config toggled) and compare failure rates statistically rather than trusting a single run either way.
A concrete run
Suppose the suspected bug is a cache-eviction race that only shows up under heavy concurrent writes. Running the harness at a fixed seed (seed=42) and concurrency=20 reproduces the race 0 times in 50 runs; the same harness at concurrency=200 reproduces it in 12 of 50 runs, a clear, statistically meaningful signal that concurrency level, not chance, drives the failure, and a concrete number, 12 of 50 at 200 versus 0 of 50 at 20, a reviewer can rerun and check rather than take on faith.
Applying this without dedicated load-test infrastructure
A single server that fails intermittently under heavy load, or that can't reproduce in staging, benefits from the cheaper version of the same idea: increase load safely on that one host (careful, bounded synthetic load) while capturing fine-grained logs/traces, rather than waiting for the next natural production spike.
Trade-offs and pitfalls
Synthetic load rarely matches production traffic shape exactly (arrival patterns, payload variety, cache warmth); the harness's value is in reliably reproducing the class of failure for iteration, not necessarily reproducing the exact production timeline, so a fix validated only against synthetic load still needs confirmation against real traffic before being called durable.
Tell me about a time you mentored someone. What were they starting from, what did you actually do, and how do you know they grew because of it?
Sample Answer
Direct answer
The strongest mentoring story names a concrete starting point (not "they were new," but what specifically they didn't yet know or couldn't yet do), describes what you actually did differently because of that starting point, and points to a real change in what the person could do independently afterward as the evidence of growth, not just that time passed or that they were nice about it.
Structured elaboration
What "starting from" should actually specify
Vague ("they were junior") is weak. Specific ("they could write correct code but always needed help scoping the actual problem before writing it") is strong, because it sets up a real before and after.
What "what you did" should show
The interesting part isn't a list of activities (pairing, reviews, 1:1s); it's the judgment behind them: why you chose that particular intervention for that particular gap, and what you adjusted when the first approach didn't fully work.
What "how you know they grew" should show
This is the part candidates under-answer. Two things separate a senior answer here:
- Independence as the real signal, not sentiment. The strongest evidence isn't "they thanked me," it's a concrete example of them handling something on their own that they previously couldn't, ideally something you didn't have to prompt.
- Reframing your own impact as leverage, not personal output. A senior candidate can articulate that developing someone else who can now independently do the work is a multiplier on team capacity, arguably more valuable than the same hours spent on your own individual output, because it compounds. That's a different, and stronger, claim than "I helped someone and it felt good."
The real tension: mentoring time vs. delivery
Mentoring genuinely competes with your own delivery time, especially early in a relationship when the payoff hasn't materialized yet. A senior answer is honest about this rather than pretending mentoring is free: it names a moment where mentoring time actually cost something (a deadline got tighter, you did more of the work yourself that cycle) and explains the judgment call for when it's right to deliberately scale mentoring back temporarily to protect a real deadline, versus when protecting the mentoring time is the higher-leverage call even under pressure.
Worked example
Situation
I mentored someone who was technically capable but consistently needed help before they'd start: given an ambiguous problem, they'd wait for someone to scope it into clear steps rather than attempting that themselves.
Action
Instead of continuing to scope tasks for them, I deliberately started handing over problems one level more ambiguous than they were comfortable with, then worked through their proposed scoping with them afterward rather than before, so the struggle happened on their side first. Early on this slowed things down, and I redid some of their scoping myself before it went further, which cost real time on a couple of deadlines.
Result
Over time the gap between their first attempt at scoping and a workable plan narrowed, until they were handling genuinely ambiguous problems without needing that step from me at all. The clearest evidence wasn't a compliment, it was a specific instance of them independently scoping and delivering something ambiguous while I was out, without anyone asking them to check with me first.
The trade-off moment
Partway through, we had a hard deadline where I made a deliberate call to scope their next task myself rather than continuing the hands-off approach, because the team couldn't absorb the risk of a slower first pass that cycle. I was explicit with them about why, so it didn't read as a loss of confidence in them, just a temporary trade-off.
Trade-offs & pitfalls
- Confusing activity with growth. Listing pairing sessions and 1:1s isn't evidence of anything; a senior answer points to a specific, observable change in independent capability.
- Never naming the cost. A story where mentoring never competed with anything else usually isn't a very real story. Naming a moment you scaled it back, and why, is more credible than claiming it was free.
- Missing the leverage framing entirely. Describing mentoring purely as "helping a nice person" misses the stronger claim: that growing someone else's independent capability is a real multiplier on what the team can deliver.
When designing audit trails for identity and access management, which access-related events must be logged (authentication, authorization failures, role changes, provisioning events, token issuance/revocation), what metadata should each event include (who, what, when, where, why, correlation IDs), how long should logs be retained, what measures ensure log integrity/tamper-resistance, and how to make logs actionable inside a SIEM or for compliance requests?
Sample Answer
Direct answer
An identity and access management (IAM) audit trail needs to log every state change that affects who can act as whom and what they can then do: authentication attempts (success and failure), authorization failures, role and permission changes, provisioning events, and token issuance and revocation. Each event needs enough metadata (who, what, when, where, why, and a correlation ID tying related events together) to answer an investigator's question without a second lookup, has to be retained long enough to matter for both incident response and whichever compliance regime applies, and has to be tamper-evident, because a log an attacker (or a rogue insider) can quietly edit after the fact is not actually evidence.
Structured elaboration
What to log. Five event families cover the IAM surface: authentication events (both successes and failures, since a run of failures followed by a success is itself a signal); authorization failures (a request denied by policy, which is often the earliest visible trace of privilege escalation or a misconfigured client); role and permission changes (a grant, a revoke, or a role assignment change, regardless of whether it was made through the normal workflow or an emergency path); provisioning events (account creation, deprovisioning, and the HR-triggered lifecycle events that drive them); and token issuance and revocation (every access, refresh, or session token minted or explicitly invalidated, since a revoked-but-still-accepted token is exactly the failure mode a security review needs to be able to rule out from the log alone).
Metadata per event. Every event should carry: who (the identity or service principal that took or attempted the action), what (the specific action and target resource), when (a precise, consistently-timezoned timestamp), where (source IP, device, or originating service), why (the business or technical context, such as which workflow or API call triggered it), and a correlation ID (a value shared across every event in one logical operation, such as a single login session or one approval-to-provisioning chain) so an investigator can pull the entire related sequence with one query instead of reconstructing it from timestamps alone.
Retention. How long to keep logs is driven by whichever compliance regime, contract, or internal policy applies to the organization, not a single universal number, but a common shape is a shorter "hot," searchable tier (commonly on the order of 90 days to a year) for active investigation and alerting, backed by a longer, cheaper "cold" archive tier (commonly multiple years) that satisfies audit and legal-hold requirements without needing to stay in an expensive, fully-indexed store the whole time.
Log integrity and tamper resistance. Write logs to an append-only store (object-lock storage, or a dedicated log database that has no update or delete API exposed to normal operators) so that even a compromised application account cannot alter history. Strengthen this with cryptographic signing, either per-event signatures or a periodically-published hash chain (a Merkle-tree-style root computed over a batch of events and published somewhere outside the log store itself), so that a tampering attempt is mathematically detectable rather than merely against policy. Just as importantly, the identities that can administer the logging pipeline should be different from the identities that administer the systems being logged, so a single compromised account cannot both take the action and erase its own trace.
Making logs actionable. A consistent, structured event schema (the same "who/what/when/where/why/correlation ID" shape across every event family) is what makes both real-time alerting and after-the-fact compliance response possible from the same data: a security information and event management (SIEM) platform can write correlation rules once against a stable schema instead of per-source-system parsers, and a compliance request ("show every access change for this employee over the last year") becomes a structured query rather than a manual log archaeology exercise.
Worked example
flowchart LR
AUTH[Authentication and authorization events] --> COL[Collector or log shipper]
PROV[Provisioning and token issuance or revocation events] --> COL
ADCH[On-prem: AD privileged group changes and odd-hour logons] --> COL
COL --> SIGN[Append-only store with per-event signing]
SIGN --> SIEM[SIEM ingestion and correlation rules]
SIGN --> COMP[Compliance export or attestation reports]
SIEM --> ALERT[Alerting and playbooks]
An on-premises worked example fills in the "role and permission change" family concretely: integrating Active Directory (AD) account auditing means capturing, at minimum, additions to privileged groups (someone added to Domain Admins) and logons occurring at unusual hours relative to that account's normal pattern. A single event record might read: who = CONTOSO\jdoe, what = "added to Domain Admins," when = 2026-03-14T02:11:00Z, where = "administrative workstation ADM-07," why = "change ticket CHG-4471" (or blank, which is itself a finding), correlation ID = the same ID shared with the change-ticket-approval event that authorized it. The odd-hour signal here (02:11 local time, outside this account's typical 9-to-6 pattern) and the privileged-group-change signal reinforce each other: either alone might be routine, but logged together with a shared correlation ID they are exactly the kind of combination a SIEM correlation rule should escalate rather than a human having to notice by cross-referencing two separate reports.
The same design also has to satisfy a stricter compliance-attestation requirement in some domains: consider a campaign-operations platform (a system tracking access to sensitive, high-stakes operational data) that must produce not just logs but generate attestations for compliance audits, a signed statement that a given set of access events is complete and unaltered for a specific period. That requirement is exactly what the append-only, cryptographically-signed store above is built to support: the compliance export is a query against the same immutable log, with the per-event or batch signatures serving as the proof the attestation can point to, rather than a separate manual sign-off process bolted on afterward.
Trade-offs and pitfalls
The main trade-off is volume versus usefulness: logging every authentication success as well as failure, every token refresh, and every authorization check at high granularity produces a very large volume of low-information events, which raises both storage cost and the noise a SIEM correlation rule has to filter through. The fix is not to log less of the IAM-relevant surface, but to route high-volume, low-signal events (routine successful token refreshes, for instance) to the cheaper cold tier immediately while keeping the genuinely decision-relevant events (failures, role changes, provisioning, revocations) in the actively-searched hot tier.
A common pitfall is logging the event but not the "why": a role-change record that shows who was granted what and when, but not which request or ticket authorized it, cannot answer the question an auditor actually asks, which is whether the change was authorized, not merely whether it happened. A second pitfall is treating the logging pipeline's own access control as an afterthought: if the same administrators who manage the identity systems being audited also have unrestricted write access to the audit store, the tamper-resistance guarantee is only as strong as that overlap, no matter how good the cryptographic signing scheme is on paper.
A configuration review of a Windows file server finds services running that nobody can justify. How do you decide what to disable, and which policy and endpoint controls keep it that way?
Sample Answer
Direct answer
Treat every running service as unjustified until someone can name the business function, the dependency, or the product feature that needs it. Inventory what is running and what is listening, classify each service as required by the file-server role, required by a dependency, or unowned, and disable the unowned ones in stages (stop, wait one full business cycle, then lock the start type through Group Policy). Then keep it that way with three layers: Group Policy sets the start type, the host firewall closes the ports the service opened, and endpoint controls plus auditing catch a service that is re-enabled or newly installed.
1. Build the evidence (inventory, then ownership)
Collect, for each server, what is installed and what is actually reachable:
Get-CimInstance -ClassName Win32_Service |
Select-Object Name, DisplayName, State, StartMode, StartName, PathName |
Sort-Object StartMode, Name
Get-NetTCPConnection -State Listen |
Select-Object LocalAddress, LocalPort, OwningProcess |
Sort-Object LocalPort
Illustrative output (a hand-written sample for a hypothetical server, not captured from a real one), trimmed to a few columns:
Name State StartMode StartName PathName
---- ----- --------- --------- --------
LanmanServer Running Auto LocalSystem C:\Windows\system32\svchost.exe -k netsvcs -p
Spooler Running Auto LocalSystem C:\Windows\System32\spoolsv.exe
VendorSync Running Auto LocalSystem D:\Tools\VendorSync\sync.exe
LocalAddress LocalPort OwningProcess
------------ --------- -------------
:: 135 1180
0.0.0.0 445 4
0.0.0.0 8443 4412
Reading a row: each line of the first table is one service, whether it is running now, how it starts, the account it runs as and the program behind it. In the second table each line is one listening port and the process ID that owns it. The third service stands out because its program sits under D:\Tools, outside the Windows and Program Files folders, and it runs as LocalSystem. StartMode is Auto, Manual or Disabled (also Boot and System for drivers). StartName is the account the service runs as: anything running as LocalSystem (a built-in account with full control of the local machine, and the computer's own identity on the network) or as a domain account (an account that works across the domain, so a stolen one reaches beyond this server) deserves extra scrutiny. PathName shows the binary: a service whose binary sits outside the Windows or Program Files folders is a question in itself. Match a listening port to a service through the process ID (OwningProcess against the service's ProcessId). One way to join the two lists, which I ran in PowerShell 7 against sample objects shaped like the real ones (illustrative values):
$byPid = @{}
Get-CimInstance Win32_Service | Where-Object ProcessId -gt 0 |
ForEach-Object { $byPid[[int]$_.ProcessId] += @($_.Name) }
Get-NetTCPConnection -State Listen | Sort-Object LocalPort |
Select-Object LocalAddress, LocalPort, OwningProcess,
@{ Name = 'Services'; Expression = { $byPid[[int]$_.OwningProcess] -join ',' } }
LocalAddress LocalPort OwningProcess Services
------------ --------- ------------- --------
:: 135 1180 RpcEptMapper,RpcSs
0.0.0.0 445 4
0.0.0.0 8443 4412 VendorSync
The first command builds a lookup from process ID to the names of the services living in that process (several services can share one process, so a row can list more than one: on a real server port 135 is the RPC endpoint mapper, which is why the sample shows the RPC services there; all service names in the sample rows are illustrative). The second adds a Services column to each listening port by looking up its owner. Port 8443 belongs to the unowned VendorSync agent, so that is a service worth asking about. A blank cell, as on port 445, means no service process owns the port; port 445 normally shows owner process ID 4, the Windows kernel's own System process, so blank is expected there. Before disabling anything, check what depends on it with Get-Service -Name <name> -DependentServices and what it requires with -RequiredServices; disabling a service that others depend on takes those down too.
Ownership: send the list of unexplained services to the application owner, the backup team, and the security team. "Nobody can justify it" is the starting position, but the answer has to be findable: a vendor agent installed by a team that left is still a service somebody owns.
2. Decide what to disable
Use a short decision rule, applied per service:
- Is it part of the file-server role or a dependency of something that is? Keep (document why).
- Is it a feature this server could use but does not (printing, remote registry access, a web server)? Candidate for disable. Fewer running services means a smaller attack surface (the set of code an attacker can reach), and some of these, such as the Print Spooler, have repeatedly been remote-exploit targets.
- Is it third-party and unowned? Stop it on a pilot server, and escalate to the owner; do not delete it first.
Worked example, one file server:
| Service | Evidence | Decision | How you verify it afterwards |
|---|---|---|---|
Server (LanmanServer) | Hosts the SMB (Server Message Block) file shares; the role depends on it | Keep | Get-Service LanmanServer is Running; shares reachable from a client |
Print Spooler (Spooler) | No printers or print queues on this server | Disable | Get-CimInstance Win32_Service -Filter "Name='Spooler'" shows StartMode Disabled and State Stopped |
Remote Registry (RemoteRegistry) | Nobody can name a tool that reads this server's registry remotely | Disable after asking the monitoring and backup owners | Same Win32_Service query shows Disabled; monitoring still green after a day |
Windows Search (WSearch) | Users search the shares through it | Keep | Get-Service WSearch is Running; a test search returns results |
| Vendor sync agent (unowned) | Runs as LocalSystem from a path outside the Windows and Program Files folders; no owner found | Stop on this server first, escalate to owners | Get-CimInstance Win32_Service -Filter "Name='<agent>'" shows Stopped; ask the owner to confirm in writing before removal |
Staged rollout: stop the service and set the start type to Disabled on one pilot server, watch for a full business cycle (including month-end jobs and the backup window, because a monthly job is exactly what a 24-hour test misses), and only then push the setting to the rest of the file servers. Keep the rollback to one command: Set-Service -Name Spooler -StartupType Manual followed by Start-Service Spooler.
3. Keep it that way: policy and endpoint controls
| Layer | Control | What it prevents or catches |
|---|---|---|
| Configuration | A Group Policy Object linked to the file-server OU, under Computer Configuration, Policies, Windows Settings, Security Settings, System Services, with each approved-off service set to Disabled | Drift: the start type is reapplied when Group Policy next processes the setting, so a local admin change is reverted on a later refresh rather than instantly |
| Network | Host firewall: inbound blocked by default, then explicit allow rules for only SMB and the management ports you use (on a locally managed computer -DefaultInboundAction already defaults to Block, so the point is to enforce it through Group Policy and keep the allow list short; Set-NetFirewallProfile -All -DefaultInboundAction Block is the local equivalent) | A service that is running but should not be reachable stays unreachable |
| Execution | Application allow-listing with App Control for Business or AppLocker, the two application control technologies Windows includes | With App Control for Business, a newly dropped service binary that no rule allows does not run (AppLocker has limits, described below) |
| Detection | Audit service installation: event 4697 "A service was installed in the system" in the Security log (subcategory Audit Security System Extension) | An unexpected new service raises an alert on a high-value server |
| Verification | A scheduled compare of Win32_Service against the approved list, with a report of anything not on it | Services that appear outside the change process |
Which of the two application control technologies to start with: Microsoft's guidance is to use App Control for Business (the newer feature, which decides which programs and drivers may run on the whole machine) wherever you can, because AppLocker still receives security fixes but no new features. AppLocker fits when you have older Windows versions in the mix or need different rules for different users on a shared computer.
Three limits of AppLocker matter for a services review. By default AppLocker policy only applies to code launched in a user's context, so a service running as SYSTEM is outside it unless you turn on its services enforcement (supported on Windows Server 2016 and later). Microsoft describes AppLocker as a defense-in-depth feature rather than a defensible security boundary, and recommends App Control for Business when the aim is robust protection. And AppLocker is not supported on Server Core installations, so check which installation option your file servers use before relying on it.
Two caveats about the audit event: it records the service name, binary path, start type and account at install time, so a later change to the binary path is not logged and has to be caught by process-creation auditing; and the event only exists if the audit subcategory (Audit Security System Extension, one switch in the Advanced Audit Policy settings that covers system-level changes such as installing a service) is enabled in policy.
Why a layered approach and not just Group Policy: the policy only controls the start type. It does not stop a service that is installed fresh, it does not close a port that something else opened, and an administrator or malware with admin rights can stop the policy from applying. The firewall, the allow-list and the audit event each cover a different gap.
Trade-offs and pitfalls
- Disabling is reversible, uninstalling is not. Prefer Disabled plus a recorded decision first, and remove the software later.
- Do not disable by name list copied from a hardening guide without checking this server's role: the same service that is a risk on one server is a dependency on another.
- Triggered services (services set to start only when an event happens, such as a device arriving, and to stop again afterwards) may look idle in a snapshot. Look at the start type, not only State.
- Record each decision (service, owner, date, evidence) so the next review does not start from zero, and review the approved list on a schedule.
How do you factor a critical vendor's own continuity posture into your DR and business continuity planning? Cover how you'd assess whether they can actually meet the recovery commitments you're relying on, and what you do when a single vendor's downtime would take down something you promised to keep running.
Sample Answer
Direct answer
I treat a critical vendor as an extension of the organization's own criticality tiering, not as someone else's problem once a contract is signed. That means verifying the vendor's own continuity posture with evidence rather than a sales claim, classifying how much concentration risk it represents (how many of our functions have no alternative path if it fails), negotiating contractual continuity obligations up front, and deciding, before anything breaks, what the business needs to happen when the vendor is down. That last decision belongs to the business-process owner who depends on the vendor, not to whichever engineer is on call when the outage starts.
Structured elaboration
Verifying the vendor's own continuity posture. Ask for evidence, not assurance: the date and outcome of their most recent continuity or DR test, a relevant attestation (SOC 2 Type II, ISO 22301 or 27001 certification), their own sub-vendor dependencies (a vendor who is itself entirely dependent on one cloud region is a risk you've now inherited), and their actual incident history rather than their marketing page's uptime claim.
Classifying concentration risk, "single point of failure" in the supplier sense. This means a vendor whose outage stops multiple business functions or customers because there is no alternative path, the same idea as a technical single point of failure, just applied to a supplier relationship rather than a system component. Tier it the way a business impact analysis (BIA), the exercise that quantifies what an outage costs each business function and sets its recovery targets, tiers internal functions:
| Tier | Definition | Example posture |
|---|---|---|
| Tier 1 | No alternative path exists; outage stops a revenue-critical or safety-critical function immediately | Requires the deepest due diligence, a tested manual workaround, and often a qualified secondary vendor |
| Tier 2 | Alternative exists but is degraded or manual | Documented workaround required, secondary vendor optional depending on cost |
| Tier 3 | Redundant paths already exist, or the function can tolerate extended downtime | Standard contract terms, lighter-touch monitoring |
Contractual levers. Recovery commitments (the vendor's own stated recovery time objective (RTO, the target time to get a function working again) and recovery point objective (RPO, how much recent data the function can afford to lose, measured in time) as contract terms), a notification SLA for when the vendor itself has an incident, audit or right-to-test clauses, and exit or data-portability terms so switching away from the vendor is actually possible rather than theoretical. SLA credits are worth including but shouldn't be mistaken for protection: a financial credit compensates the contract, it doesn't recover a lost transaction or a missed regulatory deadline.
The business-level fallback decision. For each Tier 1 or Tier 2 vendor, the business-process owner decides, in advance, one of: accept the risk (documented, with a name and a date attached to that decision so it's revisited, not forgotten), qualify and periodically re-test a secondary vendor, stand up a documented manual or degraded-mode workaround with a stated capacity limit, or insource the function. What the technology needs to deliver follows from that decision; the decision itself is a business call informed by, not delegated to, engineering.
Worked example
A mid-size payments company depends on a single external identity-verification vendor to open new customer accounts. The concentration risk is total: if new-account signups drive roughly $80,000/day in first-year revenue and identity verification is the only path to opening an account, then every hour of vendor downtime puts a proportional share of that revenue at risk:
$80,000/24≈$3,333 per hour at riskGiven that, the business-process owner (not engineering) sets the fallback: a manual review process run by the compliance team, using existing documentation to verify identity for a capped number of applications per hour, explicitly rated to absorb up to 4 hours of vendor downtime before the queue backs up faster than compliance can clear it. That 4-hour figure is the business's own RTO for this dependency, arrived at the same way a BIA sets any other recovery target, and it's what tells engineering what the fallback integration actually needs to support, not the other way around.
Trade-offs and pitfalls
SLA credits create a false sense of protection: a contractual credit worth a few hundred dollars does nothing to offset a day with $3,333/hour of exposure sitting behind it. Auditing every vendor to the same depth is unrealistic, which is exactly why the concentration-risk tiering has to gate how much due-diligence effort a given vendor gets. A "qualified backup vendor" that was tested once at onboarding and never revisited is a paper mitigation, not a real one, since vendor capabilities and your own dependency on them both drift over time. The trap specific to a technical audience: letting engineering quietly build a technical workaround to a second vendor without the contractual, compliance, and business-tolerance work behind it creates a shadow dependency that nobody has actually assessed or approved.
Design an archival policy that moves older data from a warm or hot storage tier into a cheaper cold or archival tier, without breaking jobs that occasionally still need to read snapshots of that older data. Cover the lifecycle transition rules you would set, how you would catalog or index archived data so it can still be found, the restore workflow and its expected latency, the trade-off between retrieval cost and access speed, and how you would avoid unexpected restore failures or surprise costs when older data actually gets requested back.
Sample Answer
Direct answer
Build the policy around three separate pieces that people usually collapse into one: (1) transition rules that move data between tiers automatically based on age or access pattern, (2) a catalog that always knows which tier and which retrieval class a given piece of data currently lives in, independent of where it physically sits, and (3) a restore path whose retrieval-speed class is chosen up front, at tiering time, to match the worst-case restore-time SLA (service-level agreement: a promised or contractually required target, here the maximum acceptable time to get requested data back) that data might ever need, not improvised after someone requests it back. Verify the whole thing works with real, scheduled restore drills rather than trusting the design on paper.
Lifecycle transition rules
- Trigger type: age-based (simplest: "move to cold after 365 days," configured as a native bucket/table lifecycle policy, e.g. S3 Lifecycle rules or Azure Blob lifecycle management, so the storage service itself enforces it rather than relying on an application job someone eventually forgets to run) or access-frequency-based (more precise: track actual read frequency and demote data that has genuinely gone cold, e.g. S3 Intelligent-Tiering's automatic monitoring). Age-based is the right default; add frequency-based only where access patterns don't correlate cleanly with age.
- Event-based overrides: some data should move immediately on a business event (a case closes, a record is finalized) rather than waiting for an age threshold; model this as an explicit catalog write that fires the transition, not a special case bolted onto the age rule.
- Prefer declarative, storage-service-native lifecycle rules over custom migration code wherever the storage system supports it: a misconfigured lifecycle rule fails loudly (data doesn't move, and you can see that in the catalog's per-tier byte counts); a custom migration job can fail silently and nobody notices until a restore request comes in for data that was never actually archived.
Cataloging and indexing archived data
The moment data leaves the hot, natively-queryable store, you need a durable catalog that maps a logical identifier (a file path, a partition key, a record range) to where it currently lives: tier, storage/retrieval class, the physical object key, its size, a checksum recorded at archive time, and when it becomes eligible for expiry. Without this, "find it" degenerates into enumerating every tier by hand.
- Implement it as a small, always-hot, highly available table (a Postgres table, a DynamoDB table, or, if the archived data is itself tabular, an open-table-format manifest such as Apache Iceberg or Delta Lake: a structured metadata file format used in data-lake/analytics stacks that tracks a table's underlying data files, schema, and partitions), never as something that itself gets archived: the catalog is on the critical path of every restore, so it must always be instantly readable.
- Update the catalog as part of the same operation that performs the tier transition (or reconcile it against the transition job reliably), not as an unrelated best-effort side effect; a catalog that is out of sync with reality is worse than no catalog, because it answers confidently and wrong at exactly the moment (an audit, an incident) when correctness matters most.
Restore workflow and retrieval-speed class
Cold-tier restores aren't uniform: choose among near-instant, expedited, or bulk retrieval classes based on the restore-time SLA the specific data actually needs, decided when the data is tiered, not guessed at restore time.
- Near-instant (e.g. S3 Glacier Instant Retrieval): millisecond-latency GETs, priced closer to a warm tier. Use it for cold data that is rarely read but must never make a caller wait.
- Expedited (e.g. S3 Glacier Flexible Retrieval's Expedited option): typically 1 to 5 minutes, at a premium per-request and per-GB retrieval price. Use it for data bound by a tight, hard restore-time SLA that near-instant pricing can't justify for the whole dataset.
- Bulk: the cheapest per-GB retrieval option, typically hours (and up to about two days on the deepest archival classes). This is the right default for the vast majority of archived data, which is rarely if ever requested back and has no hard restore-time bound.
The mistake this guards against: assuming any cold tier can be rehydrated "fast enough" and only discovering the real restore-time bound during an actual audit or incident, when it's too late to move the data into a faster-retrieval class.
After a restore completes, verify it, don't just trust that the read call returned success: recompute a checksum on the restored object and compare it against the checksum recorded in the catalog at archive time.
Avoiding restore failures and surprise costs
- Restore-drill verification: schedule periodic real test restores, not synthetic health checks. Pick a sample of archived partitions, run the actual restore workflow end to end, verify row counts and checksums against the catalog's recorded values, and log the observed restore duration as evidence the assigned retrieval class is genuinely meeting its SLA target. This is what catches a wrong retrieval-class assignment before a real, time-boxed request does.
- Automated compliance reporting: a recurring report or dashboard that reads the catalog and surfaces per-tier byte counts (so data that silently failed to migrate shows up as an unexpected size anomaly in the wrong tier), upcoming retention expirations (so purges happen exactly on schedule, neither late, which is a compliance finding, nor early, which is evidence destruction), and legal-hold status per record.
- Surprise-cost guardrail: expedited and near-instant retrieval are priced at a premium per GB and per request; a misconfigured or accidental bulk job that restores a whole cold partition at expedited speed instead of bulk can spend a meaningful fraction of a month's storage budget in one run. Enforce a retrieval-class allowlist per caller (only the narrow incident-response path is allowed to request expedited; scheduled batch jobs are pinned to bulk) and alert on retrieval-request volume.
- Legal hold and lifecycle expiry are two independent gates on deletion, and both must block it: a lifecycle rule alone will happily delete data that is under active litigation hold unless the hold is checked first.
Worked example: 7-year audit-log retention with a 4-hour rehydrate requirement
A system ingests 200 GB/day of audit logs, subject to a 7-year regulatory retention requirement, and an audit process that can demand a named partition rehydrated within 4 hours.
gb_per_day, years, hot_days = 200, 7, 30
warm_days = 365 - hot_days
total_gb = gb_per_day * 365 * years
hot_gb = gb_per_day * hot_days
warm_gb = gb_per_day * warm_days
cold_gb = gb_per_day * 365 * (years - 1)
# Illustrative US East list prices, fetched live from aws.amazon.com/s3/pricing/
p_std, p_ia, p_deep, p_flex = 0.023, 0.0125, 0.00099, 0.0036
sla_frac = 0.10 # share of the cold tier subject to the 4h rehydrate SLA
cold_sla_gb = cold_gb * sla_frac
cold_bulk_gb = cold_gb * (1 - sla_frac)
cost_split_cold = cold_sla_gb * p_flex + cold_bulk_gb * p_deep
total_monthly = hot_gb * p_std + warm_gb * p_ia + cost_split_cold
baseline_all_std = total_gb * p_std
print(f"total 7yr volume: {total_gb:,} GB")
print(f"tiers: hot={hot_gb:,} GB, warm={warm_gb:,} GB, cold={cold_gb:,} GB "
f"(of which {cold_sla_gb:,.0f} GB is SLA-bound)")
print(f"monthly cost, tiered with the SLA split: ${total_monthly:,.0f}")
print(f"monthly cost, all-Standard baseline: ${baseline_all_std:,.0f}")
print(f"savings: {(1 - total_monthly/baseline_all_std):.0%}")
Output:
total 7yr volume: 511,000 GB
tiers: hot=6,000 GB, warm=67,000 GB, cold=438,000 GB (of which 43,800 GB is SLA-bound)
monthly cost, tiered with the SLA split: $1,523
monthly cost, all-Standard baseline: $11,753
savings: 87%
The concrete decision the SLA forces: S3 Glacier Deep Archive's retrieval options are Standard (9-12 hours) and Bulk (up to 48 hours); neither meets a 4-hour bound. The 10% of cold data actually subject to that bound has to live in a faster-retrieval class (here, Glacier Flexible Retrieval, whose Expedited option is typically 1-5 minutes) instead, at roughly 3.6x Deep Archive's per-GB price, an explicit, budgeted trade for meeting the SLA rather than an assumption that gets discovered wrong during a real audit. The restore drill for this slice runs monthly: pick one random SLA-bound partition, run an Expedited restore, record the wall-clock time to completion, verify its row count and checksum against the catalog, and post the result to the compliance dashboard, so drift toward missing the 4-hour bound (for instance during a regional capacity crunch) is caught by the drill instead of by an auditor.
Trade-offs and pitfalls
- Defaulting everything to near-instant or expedited retrieval defeats the purpose of tiering: most archived data is never requested back, and the savings tiering exists to capture come from most of it staying in the cheapest, slowest-to-restore class. Reserve the premium retrieval class narrowly, for the specific data actually bound by a hard SLA, the way the worked example does with its 10% split.
- A catalog is only as trustworthy as its consistency with the actual data movement; treat catalog updates as part of the transition operation, not an afterthought.
- "The restore call returned success" and "the restored data is correct" are different claims; skip the checksum/row-count verification and a silently truncated or corrupted restore is indistinguishable from a good one until someone reads the data and notices.
- Don't wait for a real audit to discover a retrieval-class mismatch. A restore drill that only checks "did the API call succeed" without timing it against the SLA target isn't actually validating the thing that matters.
Your organisation is adopting Microsoft 365 and needs sign-in with on-premises AD while keeping password policies. Compare the ways to connect the two directories in terms of sign-in behaviour, security, operational overhead and resilience.
Sample Answer
Direct answer
Default to password hash synchronization (PHS) plus Seamless single sign-on (SSO), which signs in domain-joined PCs without a password prompt. It has the least infrastructure, keeps working when on-premises servers are down, and still applies your on-premises password complexity and history rules at change time. Choose pass-through authentication (PTA, where an on-premises agent checks the password against a domain controller) only if you must enforce on-premises lockout, disabled, expired and sign-in-hours state at every single sign-in. Choose federation with Active Directory Federation Services (AD FS) only for needs Microsoft Entra ID cannot meet natively, such as third-party multifactor authentication (MFA) or sign-in with a sAMAccountName. Enable PHS even if you choose PTA or federation, as the fallback.
What each method does at sign-in
A few terms first. A tenant is your organization's own dedicated instance of Microsoft Entra ID. Microsoft Entra Connect is the on-premises software that copies users from AD to the tenant. Federation means the tenant hands the password check to a separate on-premises sign-in service (AD FS), which vouches for the user afterwards.
- Password hash sync (PHS): Connect copies a processed form of each password hash to the cloud, and the cloud itself checks the password. Nothing on-premises is involved at sign-in.
- Pass-through authentication (PTA): the cloud passes the typed password to an on-premises agent, which asks a domain controller whether it is right and returns yes or no.
- Federation (AD FS): the cloud redirects the user to the on-premises AD FS farm, which checks the password and issues a signed token back to the cloud.
For choosing a method, what matters is where the password is checked, what happens when on-premises is down, and how quickly a disabled account is refused; the intervals and agent counts in the table are operating details that follow from that choice.
The options compared
| Password hash sync (PHS) | Pass-through authentication (PTA) | Federation (AD FS) | |
|---|---|---|---|
| Where the password is checked | In the cloud, against a hash of a hash synchronized from AD | On-premises, by an agent that validates against a DC | On-premises, by the federation farm |
| On-premises policy behaviour | Complexity and history apply when the password changes on-premises; expiry in the cloud is "never" by default for synced users; disabled accounts can lag up to 30 minutes; locked-out and expired states are not synced | Disabled, locked out, account expired, password expired and sign-in hours are enforced at each sign-in | Same state checks as PTA, enforced on-premises |
| Seamless SSO | Yes | Yes | No, it cannot be used with AD FS |
| Security | Entra holds a salted hash of the AD password hash (derived with PBKDF2, a deliberately slow hashing function, so it is a "hash of a hash"), which cannot be replayed in a pass-the-hash attack on-premises (an attack that signs in with a stolen hash instead of the password) | Password validation never leaves the premises; agents need unconstrained access to DCs, so they cannot sit in a perimeter network | Largest on-premises attack surface (web application proxy servers in the perimeter, certificates) |
| Operational overhead | Lowest: part of the sync process, runs every 2 minutes and the interval is not configurable | Agents on existing servers; three recommended | Highest: two or more AD FS servers and two or more proxies, TLS certificates, health monitoring |
| Resilience | Cloud service; survives an on-premises outage | Needs agents and DCs reachable; failover to PHS is manual through Entra Connect | Farm and DCs must be up; PHS can be a fallback |
Behaviour detail
- Sign-in with PHS. A user changing a password on-premises signs in with it after the next sync, normally minutes. Existing cloud sessions are unaffected. Bulk account disables should be followed by an immediate sync cycle.
- Seamless SSO creates an
AZUREADSSOcomputer account in each forest. Microsoft recommends rolling its Kerberos decryption key (the secret that lets Entra decrypt the Kerberos tickets browsers present) at least every 30 days withUpdate-AzureADSSOForest, once per forest only, because Microsoft documents that running it more than once per forest stops the feature until users' Kerberos tickets expire and are reissued by AD. It works only with PHS or PTA. - Password expiry. If synced users must follow cloud password-expiry rules, enable the
CloudPasswordPolicyForPasswordSyncedUsersEnabledfeature before enabling PHS; accounts that must never expire (service accounts) then needDisablePasswordExpirationset explicitly.
High availability of the sync engine
Microsoft Entra Connect Sync can only be active-passive. A second server in staging mode (a standby copy kept current but not allowed to write to the cloud) imports and synchronizes but does not export, and does not run password sync or password writeback until you disable staging. If it has been in staging for long, password sync needs a catch-up period after it goes active, and newly changed passwords do not work in the cloud until the backlog clears. Staged rollout, mentioned in the pitfalls, is a different feature: it moves chosen user groups to cloud authentication gradually. Check state on either server:
Import-Module ADSync
Get-ADSyncScheduler | Select-Object StagingModeEnabled
Microsoft Entra Connect cloud sync is the lighter alternative, because it does not depend on a single sync server and can provide higher availability. For PTA, install three agents (the first with Connect plus two more) so one can be down for maintenance while another fails.
Writeback
Password writeback (which keeps the on-premises password current after a cloud change) sends a password changed or reset in the cloud (self-service password reset or an administrator in the Entra admin center) back to AD DS. It works with PHS, PTA and federation, enforces your on-premises history, complexity, age and filters, uses only outbound port 443, and cannot reset passwords for protected-group members. Only one active Connect server can use writeback at a time, and it is not reliable with staged rollout enabled for a security group.
Worked example
At 09:00 HR disables a leaver's AD account. Under PTA the leaver's next sign-in is refused, because the agent checks the live account state. Under PHS the disabled state can lag up to 30 minutes, so the leaver may still sign in until about 09:30 unless you force a sync cycle after the change. If that window is unacceptable for a high-risk population, that is the case for PTA, with PHS still enabled as the fallback.
Trade-offs and pitfalls
- Switching methods needs planning; staged rollout lets you move users gradually.
- PTA keeps password validation on-premises but adds an on-premises dependency; PHS is the opposite.
- Federation is the most capable and the most fragile.
Explain the differences between TCP and UDP in terms of connection model, reliability, ordering, and flow/congestion control. For each protocol, name two real-world services that should use it and explain why. Then describe a scenario where you would build a custom reliable protocol on top of UDP rather than simply using TCP.
Sample Answer
Direct answer
TCP is connection-oriented and guarantees reliable, in-order delivery with built-in flow and congestion control, at the cost of handshake setup latency and head-of-line blocking; UDP is connectionless, with no delivery guarantees, ordering, or congestion control, trading reliability for minimal overhead and lower latency. The right choice depends on whether the application can tolerate loss and reordering itself, or needs the transport layer to handle it.
Structured elaboration
| Property | TCP | UDP |
|---|---|---|
| Connection model | Connection-oriented (handshake required) | Connectionless (no setup) |
| Reliability | Guaranteed delivery via retransmission | Best-effort, no retransmission |
| Ordering | In-order delivery guaranteed | No ordering guarantee |
| Flow control | Yes (receive window) | None |
| Congestion control | Yes (built into the protocol) | None (must be built by the application, if needed at all) |
| Overhead | Higher (handshake, ACKs, header size) | Lower (no handshake, smaller header) |
Two examples per protocol: TCP is the right choice for a database connection or a file transfer, where losing or reordering even one byte silently would corrupt the result, and the application has no interest in reimplementing reliability itself. UDP is the right choice for live video/voice calls or DNS queries, where a single lost or late packet is better DISCARDED and moved past (an old, out-of-order audio frame is useless once its playback moment has passed) than retransmitted at the cost of added latency that would make the whole stream feel laggy.
Worked example
A scenario where building custom reliability ON TOP of UDP beats plain TCP: a real-time multiplayer game sending frequent position updates. TCP's strict in-order delivery means a single lost packet blocks EVERY later packet from being delivered to the application until the lost one is retransmitted and received (head-of-line blocking), even if those later packets contain fresher, more relevant position data. Building a thin reliability layer over UDP lets the application decide per-message whether it's worth retransmitting (a player's current position, three updates old, usually isn't worth retransmitting, a newer update has probably already superseded it) rather than being forced into strict in-order delivery for data where "newest wins" matters more than "nothing lost."
Trade-offs & pitfalls
A common mistake is treating "TCP is reliable, UDP isn't" as the END of the analysis; the real question is whether your application's OWN definition of correctness matches TCP's specific guarantees (strict ordering, full reliability) or would be better served by a custom scheme that's more permissive in exactly the ways TCP is rigid. Building your own reliability on UDP is real engineering work (implementing retransmission, sequencing, and congestion awareness yourself), not a shortcut, it's justified specifically when TCP's guarantees don't match what the application actually needs.
What is the difference between a Windows Server role and a feature? Give examples of each and a situation where you would install only a feature.
Sample Answer
Direct answer
A role is the main job a server performs for other computers (DNS Server, DHCP Server, Web Server (IIS), Active Directory Domain Services). A role service is an optional piece inside a role (for example the Common HTTP Features inside Web Server, or Remote Desktop Licensing inside Remote Desktop Services). A feature is a supporting capability that is not a server's main job: it extends the operating system or the tooling (Failover Clustering, BitLocker Drive Encryption, Windows Server Backup, SNMP Service, Telnet Client, and the Remote Server Administration Tools, or RSAT). You install only a feature when the machine should gain a capability or admin tool without becoming a server for that function, such as adding the Active Directory PowerShell module to a management server so you can administer AD without installing the AD DS role.
How the three relate
| Kind | Internal name (example) | What it means | Example situation |
|---|---|---|---|
| Role | Web-Server, DNS, DHCP, AD-Domain-Services | The server provides this service to the network | Build an IIS web server |
| Role service | Web-Common-Http, RDS-Licensing | A selectable component of a role | Add only the static-content components you need |
| Feature | Failover-Clustering, Windows-Server-Backup, RSAT-AD-PowerShell, BitLocker | A capability or tool that is not a role | Cluster two file servers; back up a member server |
The Server Manager wizard shows Server Roles and Features as separate pages, and Install-WindowsFeature installs all three kinds (roles, role services and features) by internal name. Roles often pull in required features automatically as dependencies.
Commands
# What is installed, and what could be installed
Get-WindowsFeature | Where-Object Installed
Get-WindowsFeature -Name Web-Server, RSAT-AD-PowerShell
# Preview first, then install a role with all its role services and the management tools
Install-WindowsFeature -Name Web-Server -IncludeAllSubFeature -IncludeManagementTools -WhatIf
Install-WindowsFeature -Name Web-Server -IncludeManagementTools
# Install only a feature: the AD PowerShell module on a management server (no AD DS role)
Install-WindowsFeature -Name RSAT-AD-PowerShell
Install-WindowsFeature does not install the management tools unless you add -IncludeManagementTools, so a role can end up installed with no console to manage it. -WhatIf shows what would be installed without changing anything. Verify afterwards with Get-WindowsFeature -Name <name> and look at the Installed and InstallState values. For example, Get-WindowsFeature -Name Web-Server, RSAT-AD-PowerShell | Select-Object Name, Installed, InstallState prints a table like this (illustrative values):
Name Installed InstallState
---- --------- ------------
Web-Server True Installed
RSAT-AD-PowerShell False Available
Installed is a simple True or False. InstallState is the fuller status: Installed, Available (not installed, and the files are on the machine), or Removed (the files were deleted from the machine, so an install needs an outside source). Here the web server role is present and the AD PowerShell feature is not yet installed.
When I install only a feature
- Admin workstation or jump server (a hardened administration machine you connect through to manage other servers):
RSAT-AD-PowerShelland the other RSAT tools, so I can manage AD, DNS or DHCP remotely without making that machine a domain controller or DNS server. - High availability:
Failover-Clusteringon two or more file or SQL servers to cluster them; no new role is involved. - Recoverability:
Windows-Server-Backupon a member server that has no backup agent. - Test connectivity from a locked-down server: the Telnet Client feature gives a quick TCP port check, though
Test-NetConnection -Portdoes the same without installing anything.
Trade-offs and pitfalls
- Install the minimum. Every role adds services, open ports and patches, which enlarges the attack surface (the total set of running services, open ports and code an attacker could try to abuse). A feature-only install is usually a smaller commitment than a role.
- Server Core (the install option with no desktop experience) supports a subset of the roles and features, installed by command or remote tools; manage it remotely with RSAT rather than adding a GUI.
- Plan for retirement. The Windows Server Update Services role (
UpdateServices) is still installable, but Microsoft states WSUS is no longer actively developed (existing capabilities and content continue to work), so choose it for new builds only with a migration plan. - Do not confuse a feature with a Windows capability or an application. Some tools (for example WMIC in Windows Server 2025) are Features on Demand (optional Windows components that are not stored on the machine until you add them) added with DISM (Deployment Image Servicing and Management, a command-line tool that adds and removes Windows components), a different mechanism.
Want to create your own tailored preparation guide using our deep research?
Get Started for FreeInterview-Ready Courses
Visual-first, interactive, structured learning paths
Browse Systems Administrator jobs
AI-enriched listings across hundreds of company career pages
Explore Jobs