Microsoft Systems Administrator (Entry Level) - Comprehensive Interview Preparation Guide
Microsoft's Systems Administrator interview process for entry-level candidates typically involves an initial recruiter screening, followed by technical phone assessments, and multiple onsite technical and behavioral rounds. The process focuses on evaluating foundational infrastructure knowledge, troubleshooting ability, understanding of Windows and Linux systems, networking basics, and cultural fit. Entry-level candidates are expected to demonstrate eagerness to learn, strong fundamentals, and basic hands-on experience with system administration tasks.
Interview Rounds
Recruiter Screening
What to Expect
Initial screening call with a technical recruiter to verify your qualifications, assess career motivation, and determine culture fit. This is a combined round that includes both initial recruiter contact and post-technical screening recruiter follow-up. The recruiter will review your background, discuss the Systems Administrator role responsibilities, clarify your availability and location preferences, and provide an overview of Microsoft's interview process. They will also assess your understanding of the role and ensure your expectations align with the position.
Tips & Advice
Be concise and enthusiastic. Have your resume and the job description handy. Prepare a clear 2-3 minute summary of your relevant experience. Ask informed questions about the team and role. Mention specific aspects of the Systems Administrator role that excite you (e.g., infrastructure reliability, automation). Be honest about your entry-level status but emphasize your eagerness to grow in the infrastructure space.
Focus Topics
Technical communication skills
Ability to explain technical concepts clearly without jargon, and to listen and respond thoughtfully to questions.
Practice Interview
Study Questions
Relevant background and hands-on experience
Discussion of any IT support, helpdesk, system administration, or relevant technical experience. Be specific about systems you've worked with.
Practice Interview
Study Questions
Career motivation and role understanding
Clear articulation of why you want the Systems Administrator role, what attracts you to infrastructure/IT operations, and understanding of the day-to-day responsibilities.
Practice Interview
Study Questions
Technical Phone Screen
What to Expect
Technical assessment conducted via phone with a senior Systems Administrator or IT operations engineer. This round evaluates your foundational knowledge of Windows Server administration, Active Directory basics, networking fundamentals, and troubleshooting methodology. You'll be asked scenario-based questions about common infrastructure problems. Expect questions that test your understanding of server setup, user management, patch management, and basic networking. The interviewer will assess your ability to think through problems systematically and ask clarifying questions.
Tips & Advice
Think out loud and walk through your troubleshooting process step-by-step. Ask clarifying questions before attempting to solve problems. For scenario-based questions, break down the problem into components (hardware, network, software, user permissions). Don't pretend to know something you don't—instead explain how you would research or troubleshoot it. Focus on methodology rather than memorization. Be prepared to discuss your hands-on experience with specific tools and systems.
Focus Topics
Server hardware and system components
Basic knowledge of server hardware components (CPU, RAM, storage), BIOS/UEFI, hardware monitoring, and how to diagnose hardware-related issues.
Practice Interview
Study Questions
User account and access permission management
Understanding of user account lifecycle, password policies, security groups, NTFS permissions, share permissions, and access control basics.
Practice Interview
Study Questions
System troubleshooting methodology
Structured approach to troubleshooting: gathering information, identifying symptoms, checking logs, isolating the problem, testing solutions, and documenting results.
Practice Interview
Study Questions
Networking fundamentals and connectivity troubleshooting
Basic knowledge of DHCP, DNS, TCP/IP, network troubleshooting tools (ping, ipconfig, nslookup), network connectivity issues diagnosis, and how servers connect to networks.
Practice Interview
Study Questions
Windows Server fundamentals and Active Directory basics
Basic understanding of Windows Server installation, domain concepts, user and computer account creation in Active Directory, Group Policy basics, and common AD management tasks.
Practice Interview
Study Questions
Onsite Technical Interview - Windows Server and Active Directory
What to Expect
First onsite technical round conducted by a senior Windows infrastructure specialist. This round goes deeper into Windows Server administration and Active Directory management. You may work through scenario-based problems involving AD user management, group policy application, server configuration, and common Windows infrastructure issues. The interviewer will evaluate your depth of Windows Server knowledge, hands-on experience, and ability to work through real-world infrastructure challenges. This may include live demonstrations, whiteboarding discussions, or scenario walk-throughs.
Tips & Advice
Have concrete examples ready from your experience with Active Directory, user management, or Windows Server setup. If the interview includes a practical component, explain your steps clearly. For scenario questions, don't rush—take time to understand the problem fully. Discuss trade-offs and best practices (e.g., security vs. convenience). At entry level, it's acceptable to say 'I haven't encountered this specific scenario, but here's how I would approach it.' Show genuine interest in learning.
Focus Topics
Group Policy configuration and troubleshooting
Understanding Group Policy Objects (GPOs), user and computer policy scopes, policy application, troubleshooting policy application failures, and common policy use cases.
Practice Interview
Study Questions
Troubleshooting Windows Server and infrastructure issues
Approaching Windows Server issues systematically: Event Viewer analysis, system logs, common issues (service failures, permission problems, connectivity), and resolution strategies.
Practice Interview
Study Questions
Windows Server roles and features installation
Installing and configuring Windows Server operating system, installing server roles (DNS, DHCP, File Services), configuring features, and basic server hardening.
Practice Interview
Study Questions
Active Directory user and group management
Creating and managing user accounts, computer accounts, security groups, distribution groups, group nesting, and managing group memberships in AD.
Practice Interview
Study Questions
Onsite Technical Interview - Linux Systems and Networking
What to Expect
Second onsite technical round conducted by a systems engineer with Linux and networking expertise. This round evaluates your foundational Linux knowledge and networking troubleshooting skills, as the job description mentions both Windows and Linux systems. You'll be asked about basic Linux commands, file permissions, user management in Linux, network configuration, and troubleshooting network connectivity issues. Questions may include command-line exercises or scenario-based problems involving multi-OS infrastructure environments.
Tips & Advice
For Linux, focus on practical command-line skills: navigating the file system, checking logs, managing users and permissions, and basic troubleshooting. It's okay to not be as proficient in Linux if you're primarily a Windows person—entry-level admins are still building this skill. For networking, explain troubleshooting steps clearly (ping to verify connectivity, check routing, verify DNS resolution). Use the OSI model as a mental framework for troubleshooting network issues. Draw network diagrams to explain concepts if helpful.
Focus Topics
Multi-OS infrastructure environments
Understanding how Windows and Linux systems coexist in enterprise environments, interoperability considerations, and managing heterogeneous infrastructure.
Practice Interview
Study Questions
TCP/IP and networking concepts
Understanding IP addressing, subnetting, DHCP operation, DNS resolution, routing basics, and how to diagnose network connectivity issues in layered infrastructure.
Practice Interview
Study Questions
Linux fundamentals and command-line operations
Basic Linux commands (ls, cd, grep, find), file system navigation, user and group management in Linux, file permissions (chmod, chown), and viewing system logs.
Practice Interview
Study Questions
Network troubleshooting tools and techniques
Using command-line tools for network diagnostics: ping, ipconfig/ifconfig, nslookup/dig, tracert/traceroute, netstat, arp, port connectivity testing, and interpreting results.
Practice Interview
Study Questions
Onsite Technical Interview - Infrastructure Operations and Reliability
What to Expect
Third onsite technical round conducted by an operations or infrastructure reliability specialist. This round focuses on the operational aspects of systems administration mentioned in the job description: backup and disaster recovery procedures, system performance monitoring, security implementation, and preventive maintenance. You'll discuss how systems are monitored, how backups ensure data protection, how patches are deployed reliably, and how to respond to infrastructure incidents. This round emphasizes reliability, automation potential, and best practices for enterprise infrastructure.
Tips & Advice
Discuss backup and recovery from a practical perspective—RPO/RTO concepts, testing recovery procedures, and lessons learned. For monitoring, explain what metrics matter and how to detect problems early. Discuss patch management as a balance between security and stability. Show awareness of security best practices (least privilege, defense in depth, monitoring). At entry level, focus on understanding these concepts and their importance rather than claiming expertise in implementing them.
Focus Topics
Incident response and troubleshooting methodology
Responding to infrastructure incidents, root cause analysis, escalation procedures, documentation practices, and learning from incidents.
Practice Interview
Study Questions
Infrastructure security implementation
Security best practices for servers and systems, firewall configuration basics, antivirus/malware protection, access controls, encryption basics, and security monitoring.
Practice Interview
Study Questions
System performance monitoring and optimization
Monitoring CPU, memory, disk, and network utilization; identifying performance bottlenecks; analyzing system logs; using monitoring tools; and optimizing resource usage.
Practice Interview
Study Questions
Security updates, patching, and compliance
Software update and patch management processes, deploying patches safely, testing procedures, security vulnerability assessment, and compliance considerations.
Practice Interview
Study Questions
Backup and disaster recovery procedures
Understanding backup strategies, RPO/RTO concepts, recovery procedures, testing recovery, backup verification processes, and ensuring business continuity.
Practice Interview
Study Questions
Onsite Behavioral Interview - Culture and Collaboration
What to Expect
Final onsite round conducted by an HR representative or senior team member to assess culture fit, teamwork, communication skills, and Microsoft's core values. This round evaluates how you collaborate with colleagues, support end-users, handle stress, and align with Microsoft's culture of learning and growth. You'll discuss your work style, experience working in teams, handling difficult situations, learning from failures, and your approach to helping others. Questions will focus on your ability to grow, adapt, and contribute positively to the team and organization.
Tips & Advice
Use the STAR method for behavioral questions (Situation, Task, Action, Result). Focus on specific examples from your experience. At entry level, it's appropriate to discuss learning experiences and how you've grown from challenges. Emphasize teamwork, willingness to help colleagues, and curiosity about technology. Discuss how you handle support tickets or user requests with patience. Ask thoughtful questions about team dynamics and growth opportunities. Show genuine enthusiasm about Microsoft's mission and values.
Focus Topics
Handling challenges and learning from failures
Examples of technical problems you've troubleshot, mistakes you've made and learned from, how you approach unfamiliar situations, and your commitment to continuous learning.
Practice Interview
Study Questions
Work style and stress management
How you prioritize multiple tasks during incidents, manage stress in critical situations, maintain professionalism under pressure, and take care of your wellbeing.
Practice Interview
Study Questions
User support and communication
Providing technical support to non-technical users, explaining technical issues in accessible language, patience with frustrated users, and prioritizing support requests.
Practice Interview
Study Questions
Teamwork and collaboration
Working effectively with colleagues in IT operations teams, supporting each other, sharing knowledge, and collaborating to solve infrastructure problems.
Practice Interview
Study Questions
Frequently Asked Systems Administrator Interview Questions
Describe the steps and commands to add a new physical disk to an existing LVM volume group, create or extend a logical volume, and grow the filesystem online. Mention any differences when using ext4 vs xfs and safety considerations (backups, snapshots).
Sample Answer
Direct answer
The sequence is the same regardless of filesystem: initialize the new disk as an LVM (Logical Volume Manager) physical volume, add it to the existing volume group, extend the logical volume into the newly available space, then grow the filesystem to fill the larger volume, all of which can happen while the filesystem stays mounted. The one real difference between ext4 and XFS shows up at the very last step: resize2fs (ext4) takes the block device directly and can grow or, offline only, shrink; xfs_growfs (XFS) takes the mount point, not the device, and can only ever grow, never shrink, for the life of that filesystem.
The steps, concretely
Given an existing volume group vgdata with a logical volume lvapp mounted at /data, and a new disk /dev/sdb just attached to the host:
pvcreate /dev/sdb
vgextend vgdata /dev/sdb
lvextend -L +100G /dev/vgdata/lvapp
pvcreate writes LVM's metadata header onto the new disk, marking it as available for use. vgextend folds its capacity into the existing volume group's free space pool. lvextend -L +100G grows the logical volume by 100 gigabytes drawn from that pool (you can also use -l +100%FREE to consume all of a VG's remaining free space in one step, rather than specifying an exact size).
Growing the filesystem, ext4:
resize2fs /dev/vgdata/lvapp
Run with no size argument, resize2fs grows the filesystem to fill the entire underlying block device, and it can do this online, while /data stays mounted and in active use.
Growing the filesystem, XFS:
xfs_growfs /data
Note the argument is the mount point, /data, not the device path; XFS's growfs tool operates through the mounted filesystem itself rather than the raw device. Like resize2fs, this runs online with the filesystem mounted and serving traffic throughout.
Worked example with real numbers
Starting from a 150 GiB lvapp that is nearly full, adding a second 150 GiB disk and extending fully:
$ vgextend vgdata /dev/sdb
Volume group "vgdata" successfully extended
$ lvextend -L +150G /dev/vgdata/lvapp
Size of logical volume vgdata/lvapp changed from 150.00 GiB (38400 extents) to 300.00 GiB (76800 extents).
$ resize2fs /dev/vgdata/lvapp
# filesystem now reports 300 GiB total, still mounted throughout
The logical volume doubled from 150 GiB to 300 GiB by consuming the entirety of the newly added disk's capacity, and the extent count doubling exactly in step with the size (38400 to 76800, at LVM's default 4 MiB extent size: 150 GiB / 4 MiB = 38400 extents, confirming the arithmetic) is a useful sanity check that the operation did exactly what you expected before you move on to growing the filesystem itself.
Safety considerations before touching any of this
- Take a backup or an LVM snapshot before extending anything on a volume holding data you cannot afford to lose, even though these operations are designed to be safe: a snapshot (
lvcreate --snapshot -L 20G -n lvapp-snap /dev/vgdata/lvapp) gives you a rollback point in case something unrelated goes wrong mid-operation, at essentially no cost if nothing does. - Verify you are extending the correct logical volume and using the correct new disk, with
lsblkandvgs/pvsbefore running anything: extending the wrong LV, or accidentally initializing a disk that already holds data you meant to keep, as a fresh PV, is a genuinely destructive, hard-to-reverse mistake, and it happens most often when device names like/dev/sdbshift after a reboot on a host with multiple disks. - Check free space in the VG first (
vgsshowsVFreedirectly) solvextenddoes not fail partway through, or worse, silently succeed with a smaller size than intended because you asked for more than was actually available.
Trade-offs and pitfalls
- The ext4-versus-XFS difference here is not cosmetic: if you ever expect to need to shrink a volume (moving data to smaller, cheaper storage; reclaiming clearly overprovisioned capacity), XFS forecloses that option permanently the moment you format with it, while ext4 at least allows it offline. This is a decision to make at filesystem-creation time, not something to discover you need later.
resize2fstargets the block device whilexfs_growfstargets the mount point; running the wrong tool against the wrong kind of argument is a common enough mix-up that it is worth double-checking which filesystem you are actually working with before running either command, especially in a mixed-filesystem environment.- None of these operations validate that the application using the volume actually benefits from more space immediately; a database or log aggregator that was already failing due to being full may need an explicit nudge (a restart, or clearing a stuck disk-full error state) even after the underlying filesystem correctly reports more room available.
You are asked to perform a threat model for an API gateway that routes traffic to multiple backend services. What are the top five threat vectors you would analyze, and what countermeasures would you recommend for each?
Sample Answer
Direct answer
An API gateway routing to multiple backend services sits at the one point in the architecture where every client request converges before fanning out, which makes it both the highest-leverage place to enforce security consistently and the single point whose own compromise or misconfiguration has the broadest blast radius; the five highest-priority threat vectors are authentication/authorization bypass, injection and input tampering, denial of service and abuse, misrouting and server-side request forgery (SSRF) toward internal backends, and credential/secrets exposure at the gateway layer itself.
Structured elaboration
1. Authentication and authorization bypass. A client reaches a backend service without a valid identity, or with a valid identity but insufficient authorization for the specific route requested, either because the gateway's own authentication check is misconfigured (a route accidentally excluded from the authentication requirement) or because a backend service incorrectly trusts that the gateway has already fully validated authorization for that specific request when it has only validated authentication. Countermeasures: enforce authentication centrally at the gateway for every route by default (an explicit opt-out for a genuinely public route, not an opt-in requirement that a new route could silently miss), validate JSON Web Tokens (JWTs) fully (signature, issuer, audience, expiration) rather than only checking for presence, and have each backend service independently re-validate authorization for its own specific resources rather than fully trusting the gateway's authentication as a substitute for its own authorization logic.
2. Injection and input tampering. A malicious or malformed request body, header, or query parameter passes through the gateway unvalidated and reaches a backend service that trusts it, since the gateway's routing function does not inherently include content validation unless explicitly configured to. Countermeasures: a web application firewall (WAF) integrated with the gateway inspecting request content against known injection patterns, and, more fundamentally, schema validation at the gateway for each route's expected request shape, rejecting anything that does not conform before it ever reaches a backend.
3. Denial of service and abuse. A single client, or a distributed set of clients, overwhelms either the gateway itself or a specific backend service through excessive request volume; because the gateway is the single point every request passes through, an under-provisioned or unprotected gateway is a single point of failure for every backend service behind it simultaneously, a materially worse outcome than one backend service alone being overwhelmed. Countermeasures: rate limiting and per-client quotas enforced at the gateway, and gateway capacity provisioned and tested against realistic peak load, not just typical traffic.
4. Misrouting and server-side request forgery toward internal backends. A gateway misconfiguration (a route mapping error, or a routing rule that fails to validate the destination correctly) sends a request to an unintended internal backend, potentially one never meant to be reachable through the gateway's public-facing routes at all; separately, if the gateway itself performs any server-side fetch based on request content (less common, but present in some gateway designs with dynamic backend resolution), that fetch logic is itself a potential SSRF vector against internal infrastructure. Countermeasures: an explicit, reviewed allow-list of valid backend destinations per route, never a dynamically-resolved or wildcard destination, and network-level segmentation ensuring the gateway itself can only reach the specific backend services its routing configuration legitimately targets, not the organization's full internal network.
5. Credential and secrets exposure at the gateway layer. The gateway itself typically holds credentials needed to authenticate to backend services on the client's behalf (an internal service-to-service credential, an API key for a downstream integration); a compromise of the gateway itself, or a logging misconfiguration that captures these credentials in request/response logs, exposes every backend service the gateway integrates with at once, not just one. Countermeasures: the gateway's own credentials for backend authentication should be short-lived and narrowly scoped per backend, not one broad, long-lived credential reused across every downstream integration, and logging configuration should explicitly exclude credential-bearing headers and fields from captured log output.
Worked example
An API gateway routes to three backend services: a public product-catalog service, an authenticated order-processing service, and an internal-only inventory-management service never meant to be reachable from outside the organization. A misconfiguration (vector 4) accidentally exposes a route to the inventory-management service through the gateway's public-facing configuration; combined with an authentication gap (vector 1, the newly-exposed route was not added to the gateway's default-authenticated route set), an unauthenticated external caller can reach the internal inventory service directly. The gateway's rate limiting (vector 3) does not prevent this, since the request volume from a single, patient attacker probing for exactly this kind of exposed route stays well under any reasonable rate threshold. The gap is closed by two independent fixes: correcting the routing configuration to remove the internal service from the public-facing route set (closing vector 4 directly), and separately confirming every route, including ones assumed to already be adequately protected, is included in the gateway's default-authenticated set rather than relying on each route being individually, correctly configured (closing vector 1 as a systemic fix, not just for this one route).
Trade-offs and pitfalls
- The worked example's compromise required two separate gaps (a routing misconfiguration and an authentication gap) to actually manifest, and fixing only one of the two would have left the other quietly present, waiting for the next routing mistake to expose it again; treating these five vectors as independent items on a checklist, rather than recognizing how they compound in a real incident, understates the actual risk of any one gap on its own.
- Rate limiting is necessary but, as the worked example shows, does not catch every abuse pattern, specifically a patient, low-volume reconnaissance attempt looking for a misconfigured route rather than attempting to overwhelm capacity; a design that treats rate limiting as covering "abuse" broadly, rather than specifically volumetric abuse, has a gap for exactly this slower, more deliberate attack pattern.
- Requiring each backend service to independently re-validate authorization, rather than fully trusting the gateway's own authentication check, adds real development overhead across every backend team, and it is precisely what limits the worked example's blast radius if the gateway-level authentication gap had gone undetected longer; a backend service that blindly trusted "the gateway already checked this" would have had no independent check to catch what the gateway itself missed.
- The gateway's own credential-management practice (vector 5) is easy to under-prioritize relative to the more visible, request-facing vectors, since a credential-exposure incident is less immediately visible than a misrouted request; but a compromise here has the broadest blast radius of any of the five vectors, since it affects every backend integration simultaneously, not one route at a time, which is why it belongs on this list at the same priority tier as the more obviously request-facing threats.
The business wants about 100 people to reach internal applications securely from home on Windows Server remote sessions. Describe the deployment you would build: which components, how it stays up if one server fails, how licensing works, and how you protect access from the internet.
Sample Answer
Direct answer
For about 100 remote users I would build a Remote Desktop Services (RDS) deployment with five parts: session hosts (the servers where users' desktops and apps actually run), a Remote Desktop Connection Broker (tracks who is on which host and reconnects them), Remote Desktop Gateway (tunnels RDP over HTTPS port 443, so RDP itself is never exposed), Remote Desktop Web Access (the web page that lists apps and desktops) and a Remote Desktop Licensing server (issues the client access licenses, or CALs). Every infrastructure role is duplicated on a second machine, the broker's state lives in a shared SQL database (SQL Server or Azure SQL Database, a database service the two brokers both read), and the two gateway/web servers sit behind a load balancer (a device or service that spreads incoming connections across several servers and stops sending them to one that stops responding). Licensing is per-user RDS CALs, one per person, installed on a license server. Internet exposure is limited to the gateway and Web Access servers behind one load balancer on TCP 443 (plus UDP 3391 for the gateway's UDP transport), with a trusted certificate and multifactor authentication on the gateway sign-in.
Components and how each stays up
| Component | Count | How it stays up if one fails |
|---|---|---|
| Session hosts | 4 | The broker sends new connections to the remaining hosts; users on a failed host lose that session and reconnect to another |
| Connection Broker (with Licensing co-located) | 2 | Two brokers sharing one SQL database holding the connection information, reached through one DNS name (a load balancer, or DNS round robin, where one name resolves to several server addresses handed out in turn but nothing checks that each server is alive) |
| RD Gateway and RD Web Access (co-located) | 2 | Behind a load balancer with a TCP 443 health probe (the balancer regularly tries a connection to port 443 and removes a server that fails), plus a UDP 3391 load-balancing rule (3391 is the port for the UDP transport the gateway offers alongside HTTPS) |
| SQL database for the broker | 1 managed or existing clustered SQL | Use SQL Server or Azure SQL Database with its own availability |
| License server | 2 (on the broker pair) | Every RD server is configured to use both license servers, so licenses can still be issued if one is down. The 100 CALs are a total across them, not 100 per server |
Microsoft's high-availability guidance is to duplicate each infrastructure role service on a second machine. That is eight servers plus the database (4 session hosts + 2 broker and licensing + 2 gateway and web access), which is the footprint to defend: a smaller one has a single point of failure at some layer. The session hosts, the broker, the gateway and the license server are what a working remote session needs; Web Access is the page where users pick their apps and desktops.
Sizing the session hosts (worked arithmetic)
Capacity is a measured number, not a formula. I run a pilot with the real applications and find how many concurrent users one session host carries at an acceptable response time; assume the pilot shows 35. Worst case all 100 users are connected at once:
- Hosts to carry the load:
ceil(100 / 35)= 3. - One extra host so one can fail or be patched: 3 + 1 = 4 hosts.
- With one host down, 3 hosts carry 100 users: 100 / 3 = 33.3 users each, which is under 35, so the design holds during a failure.
If the pilot measures 50 instead, the answer changes to ceil(100 / 50) = 2, plus 1 = 3 hosts. The measured per-host number is what flips the host count.
Licensing
- Each user and each device that connects to a session host running Windows Server needs an RDS CAL, issued by a Remote Desktop Licensing server.
- How many to buy: 100 users means 100 per-user CALs, not 200 because there are two license servers. Microsoft's roles guidance has both license servers associated with every RD server so issuing continues if one is down; how the purchased license codes are installed across the pair is a licensing-agreement question to settle with your Microsoft reseller before the build.
- Per-user vs per-device: remote workers use their own devices from home, so I would buy per-user CALs (100). Per-user CALs are tied to a user in Active Directory, cannot be tracked in a workgroup, and can be over-allocated, which would breach the licensing terms, so I track usage in Remote Desktop Licensing Manager. Per-device CALs fit shared devices used by several shifts.
- Version rule: the RDS CAL must be the same version as, or newer than, the session host's Windows Server version (a Windows Server 2022 CAL cannot be used on a 2025 session host), and the license server must run the same or a newer Windows Server version than the CALs it holds.
- Grace period: there is a 120 day grace period during which no license server is needed. After it ends, clients must hold a valid CAL to sign in, so the license server is built and activated first, not last.
Protecting access from the internet
- Expose only the gateway tier, never RDP (3389) or the broker. RD Gateway and RD Web Access run on the same two servers behind the same load balancer, so TCP 443 (and UDP 3391 for the gateway) reaches both, while the brokers, session hosts and SQL database stay unreachable from the internet. Place the gateway servers in a perimeter network (a separate network zone between the internet and the internal network, sometimes called a DMZ), with TCP 443 inbound to the load balancer and rules allowing the gateway to reach the internal brokers and hosts.
- A trusted TLS certificate on the gateway and web access servers, and a public DNS name that matches it.
- Authorization policies: RD Gateway requires a Connection Authorization Policy (RD CAP, who may connect) and a Resource Authorization Policy (RD RAP, which internal machines they may reach). By default both are rules stored on the gateway; the CAPs can instead be held on a central Network Policy Server, while RAPs are always evaluated on the gateway itself. Scope users to a named security group and the RAP to the session hosts only.
- Multifactor authentication: with the RD CAPs held on a central Network Policy Server (NPS, the Windows Server role that answers RADIUS authentication requests), the Microsoft Entra multifactor authentication NPS extension, installed on that NPS server and not on the gateway, adds a second factor to the gateway sign-in. The gateway sign-in supports phone call and Approve/Deny push in the Authenticator app, not SMS or typed codes, so plan enrollment accordingly. This is a layer added on top of a working gateway.
- Operations: patch the gateway first, alert on repeated sign-in failures, and give both Web Access servers identical validation and decryption machine keys, as Microsoft's farm procedure requires. Those keys are the secrets IIS uses to sign and decrypt the tokens Web Access issues to a browser; with identical keys, a sign-in started on one Web Access server is accepted by the other when the load balancer switches the user over.
Trade-offs and pitfalls
- Recommendation: this on-premises RDS design if the applications must stay on your servers. Choose Azure Virtual Desktop instead if there is no appetite for running broker, gateway and database infrastructure, or if the team is not already staffed to patch it.
- Windows Internal Database (WID, a lightweight database engine built into Windows) is used by roles that include the Connection Broker, but a highly available broker needs a shared database server (SQL Server or Azure SQL Database). Microsoft also lists WID as being removed from Windows in a future release and advises SQL Server for these roles, so use SQL from the start.
- Co-locating roles saves servers but ties failure domains together. I accept brokers plus licensing on one pair at 100 users; I would separate them at larger scale.
- User data on the session hosts disappears with a failed host. Keep profiles and files on a highly available file share, not on the host's local disk.
Your organisation is adopting Microsoft 365 and needs sign-in with on-premises AD while keeping password policies. Compare the ways to connect the two directories in terms of sign-in behaviour, security, operational overhead and resilience.
Sample Answer
Direct answer
Default to password hash synchronization (PHS) plus Seamless single sign-on (SSO), which signs in domain-joined PCs without a password prompt. It has the least infrastructure, keeps working when on-premises servers are down, and still applies your on-premises password complexity and history rules at change time. Choose pass-through authentication (PTA, where an on-premises agent checks the password against a domain controller) only if you must enforce on-premises lockout, disabled, expired and sign-in-hours state at every single sign-in. Choose federation with Active Directory Federation Services (AD FS) only for needs Microsoft Entra ID cannot meet natively, such as third-party multifactor authentication (MFA) or sign-in with a sAMAccountName. Enable PHS even if you choose PTA or federation, as the fallback.
What each method does at sign-in
A few terms first. A tenant is your organization's own dedicated instance of Microsoft Entra ID. Microsoft Entra Connect is the on-premises software that copies users from AD to the tenant. Federation means the tenant hands the password check to a separate on-premises sign-in service (AD FS), which vouches for the user afterwards.
- Password hash sync (PHS): Connect copies a processed form of each password hash to the cloud, and the cloud itself checks the password. Nothing on-premises is involved at sign-in.
- Pass-through authentication (PTA): the cloud passes the typed password to an on-premises agent, which asks a domain controller whether it is right and returns yes or no.
- Federation (AD FS): the cloud redirects the user to the on-premises AD FS farm, which checks the password and issues a signed token back to the cloud.
For choosing a method, what matters is where the password is checked, what happens when on-premises is down, and how quickly a disabled account is refused; the intervals and agent counts in the table are operating details that follow from that choice.
The options compared
| Password hash sync (PHS) | Pass-through authentication (PTA) | Federation (AD FS) | |
|---|---|---|---|
| Where the password is checked | In the cloud, against a hash of a hash synchronized from AD | On-premises, by an agent that validates against a DC | On-premises, by the federation farm |
| On-premises policy behaviour | Complexity and history apply when the password changes on-premises; expiry in the cloud is "never" by default for synced users; disabled accounts can lag up to 30 minutes; locked-out and expired states are not synced | Disabled, locked out, account expired, password expired and sign-in hours are enforced at each sign-in | Same state checks as PTA, enforced on-premises |
| Seamless SSO | Yes | Yes | No, it cannot be used with AD FS |
| Security | Entra holds a salted hash of the AD password hash (derived with PBKDF2, a deliberately slow hashing function, so it is a "hash of a hash"), which cannot be replayed in a pass-the-hash attack on-premises (an attack that signs in with a stolen hash instead of the password) | Password validation never leaves the premises; agents need unconstrained access to DCs, so they cannot sit in a perimeter network | Largest on-premises attack surface (web application proxy servers in the perimeter, certificates) |
| Operational overhead | Lowest: part of the sync process, runs every 2 minutes and the interval is not configurable | Agents on existing servers; three recommended | Highest: two or more AD FS servers and two or more proxies, TLS certificates, health monitoring |
| Resilience | Cloud service; survives an on-premises outage | Needs agents and DCs reachable; failover to PHS is manual through Entra Connect | Farm and DCs must be up; PHS can be a fallback |
Behaviour detail
- Sign-in with PHS. A user changing a password on-premises signs in with it after the next sync, normally minutes. Existing cloud sessions are unaffected. Bulk account disables should be followed by an immediate sync cycle.
- Seamless SSO creates an
AZUREADSSOcomputer account in each forest. Microsoft recommends rolling its Kerberos decryption key (the secret that lets Entra decrypt the Kerberos tickets browsers present) at least every 30 days withUpdate-AzureADSSOForest, once per forest only, because Microsoft documents that running it more than once per forest stops the feature until users' Kerberos tickets expire and are reissued by AD. It works only with PHS or PTA. - Password expiry. If synced users must follow cloud password-expiry rules, enable the
CloudPasswordPolicyForPasswordSyncedUsersEnabledfeature before enabling PHS; accounts that must never expire (service accounts) then needDisablePasswordExpirationset explicitly.
High availability of the sync engine
Microsoft Entra Connect Sync can only be active-passive. A second server in staging mode (a standby copy kept current but not allowed to write to the cloud) imports and synchronizes but does not export, and does not run password sync or password writeback until you disable staging. If it has been in staging for long, password sync needs a catch-up period after it goes active, and newly changed passwords do not work in the cloud until the backlog clears. Staged rollout, mentioned in the pitfalls, is a different feature: it moves chosen user groups to cloud authentication gradually. Check state on either server:
Import-Module ADSync
Get-ADSyncScheduler | Select-Object StagingModeEnabled
Microsoft Entra Connect cloud sync is the lighter alternative, because it does not depend on a single sync server and can provide higher availability. For PTA, install three agents (the first with Connect plus two more) so one can be down for maintenance while another fails.
Writeback
Password writeback (which keeps the on-premises password current after a cloud change) sends a password changed or reset in the cloud (self-service password reset or an administrator in the Entra admin center) back to AD DS. It works with PHS, PTA and federation, enforces your on-premises history, complexity, age and filters, uses only outbound port 443, and cannot reset passwords for protected-group members. Only one active Connect server can use writeback at a time, and it is not reliable with staged rollout enabled for a security group.
Worked example
At 09:00 HR disables a leaver's AD account. Under PTA the leaver's next sign-in is refused, because the agent checks the live account state. Under PHS the disabled state can lag up to 30 minutes, so the leaver may still sign in until about 09:30 unless you force a sync cycle after the change. If that window is unacceptable for a high-risk population, that is the case for PTA, with PHS still enabled as the fallback.
Trade-offs and pitfalls
- Switching methods needs planning; staged rollout lets you move users gradually.
- PTA keeps password validation on-premises but adds an on-premises dependency; PHS is the opposite.
- Federation is the most capable and the most fragile.
You're partway through a sequenced recovery when a third-party dependency you were counting on stays down longer than expected. Which services do you bring online anyway, how do you handle the transactions that would normally rely on that dependency, and how do you communicate the degraded state to customers in the meantime?
Sample Answer
Direct answer. Bring online everything that doesn't strictly need the down dependency, and put a firm, visible hold on anything that does, rather than letting it silently degrade or silently retry. Decide what "handling a transaction" means without the dependency in a way that never risks a customer being charged, billed, or committed twice: hold, don't guess. And tell customers proactively, in plain language, before they have to ask.
1. Decide what comes online, using criticality, not convenience
This decision should trace back to the business impact analysis, the process that ranks business functions by how much an hour or a day of downtime actually costs, so it isn't made ad hoc mid-incident. Functions that don't touch the down dependency come online first. Functions that touch it only on a non-essential path (browsing, viewing account history) come online in a read-only or informational mode. Functions where the dependency is essential to the transaction itself (authorizing a new charge) stay explicitly gated, not silently attempted.
The authority to declare "we're operating in a degraded state" and to approve which functions run that way should be defined in the continuity plan ahead of time, not improvised in the room. That's usually someone senior enough to own the customer and regulatory risk of the decision, not just whoever is closest to the outage.
2. Handle the affected transactions without guessing
Accept the transaction, record the customer's intent, and hold it in a clearly marked pending state rather than attempting it against a dependency that isn't there, or silently retrying it in the background where the customer can't see what happened. Be honest about the state: "received and pending" is very different from letting a customer believe it went through, and that distinction is also what stops a customer from trying again themselves out of uncertainty.
When the dependency comes back, process the backlog in order and reconcile before declaring the incident closed: confirm, transaction by transaction, that what your system believes happened matches what the dependency's own records show, and follow up individually on anything that doesn't match rather than assuming the queue drained cleanly.
3. Communicate the degraded state to customers
Say something before they have to ask. A visible status message in the product itself, not only on a status page nobody checks mid-transaction, that names what's affected, sets expectations (their action is saved but pending, not lost), and gives a realistic timeframe, even a wide one, beats silence. Update it as the situation changes, and if any pending items need a customer to take action, reach out directly instead of leaving them to notice on their own. Keep the message consistent across channels so a customer who checks two of them doesn't get two different stories on top of the outage itself.
4. Protect service levels for what's still running
The functions still online need their own expectations reset for the duration. It's reasonable to temporarily relax targets for anything adjacent to the affected dependency (background jobs that would also normally touch it) so they don't cause a second incident by retrying aggressively against something that's down. That should be a deliberate, communicated decision, not something that just happens because nobody planned for it.
Worked example
A checkout flow depends on a third-party payment processor down for two hours, well past what anyone expected. Browsing, cart building, and order history don't touch the processor, so they stay fully online. Checkout is gated: a customer can complete every step through "place order," and the order is recorded as pending payment with the cart and the chosen payment method captured, but no charge is attempted. The product tells them directly: order saved, payment processing is delayed, confirmation will follow once it's back, expected within the hour. No background retry loop fires against the processor. When it recovers, pending orders are processed in the order placed, and each is reconciled against the processor's own transaction record before being marked complete, so a customer is never charged twice even if an earlier attempt is later discovered to have partially gone through.
Trade-offs & pitfalls. The tempting shortcut is to keep retrying the transaction in the background hoping the dependency returns soon; that's exactly how a customer ends up double-charged if the retry succeeds silently after they've already tried again themselves out of frustration. Hold and communicate, don't guess and retry. Silence is worse than an honest "we don't know exactly when," because customers who get no information assume the worst or attempt workarounds that make reconciliation harder afterward. And bringing too much online too fast, without a clear degraded-mode decision from someone with the authority to own that risk, is how a team ends up shipping a feature that looks live but isn't actually safe to depend on.
Executives have cut maintenance windows to almost nothing. How do you keep hosts patched and hardened without regular downtime, and what do you tell the business about the risk that remains?
Sample Answer
Direct answer
Separate the two things reboots are doing in the plan: applying fixes, and restarting the machine. Reduce how often a machine must restart (live kernel patching, Windows hotpatching where supported, replacing instead of patching, rolling updates over redundant nodes), and be honest with the business that the rest cannot be eliminated, so give them a number for the remaining exposure and ask for the smallest guaranteed window that bounds it. Never promise "patched with no downtime and no risk".
Step 1: Sort the estate by how it can be updated without downtime
Each row below is a method for a different kind of host, so they are not alternatives for one machine; an estate usually uses several. (Hardening baselines, meaning agreed secure settings, come in Step 2. The word "baseline" in the hotpatch row means something else: a full cumulative update that does require a restart.)
| Host type | No-downtime method | What it does not cover |
|---|---|---|
| Stateless web or application tier behind a load balancer | Replace instances with a freshly built patched image, one batch at a time, draining each batch first | Nothing special; this is the cleanest answer |
| Clustered roles on Windows | Cluster-Aware Updating (CAU), Windows Server's built-in tool for updating a failover cluster (a group of servers that can take over each other's roles): puts each node into maintenance mode (a state in which the node is drained of work), moves the clustered roles off it, installs updates, restarts if needed, brings roles back, then does the next node | Many roles trigger a planned failover, which can cause a brief interruption for connected clients; only continuously available workloads (those designed to keep serving clients while a node moves, such as Hyper-V with live migration or a file server using SMB Transparent Failover) avoid it |
| Linux servers needing kernel fixes | Kernel live patching, which loads a replacement for a faulty kernel function into the running kernel and redirects calls to it, so the fix applies without a restart | Only functions that the kernel's tracing hook (ftrace, the mechanism live patching uses to intercept a call) can intercept can be patched, so some fixes still need a reboot, and user-space libraries need their services restarted |
| Windows Server on Azure, Azure Local (Microsoft's product that extends Azure to your own hardware, managed through Azure Arc) or Azure Arc-connected (servers outside Azure registered for Azure management) | Hotpatch (patches in-memory code without restart), supported only on specific editions: Datacenter: Azure Edition images of Windows Server 2022 and 2025 on Azure and Azure Local, and Windows Server 2025 Standard or Datacenter on Azure Arc-connected machines once Arc hotpatching is enabled (Windows Server 2022 is not supported on Arc) | A new hotpatch baseline (a full cumulative update that resets the cycle) arrives every three months and still requires a restart; non-security updates, .NET and driver or firmware updates are not hotpatched; hotpatch has no automatic rollback |
| Single-instance legacy servers | None | Needs a window, or accepted risk with compensating controls |
The hotpatch facts: the hotpatch baseline refreshes every three months, with hotpatch releases for the following two months, so a planned year is four baselines (restart needed) and eight hotpatch months. An unplanned baseline can replace a hotpatch month if a fix cannot be hotpatched.
Checking one host. After a live patch or hotpatch, confirm it landed on the host itself, because "the console says compliant" is a second-hand view.
- Linux live patch:
uname -rstill prints the old kernel version (a live patch does not change it), so check the live-patch interface instead. The kernel keeps one directory per loaded patch under/sys/kernel/livepatch/, andcat /sys/kernel/livepatch/<patch-name>/enabledreads1while that patch is active (patch names depend on the vendor's package; I could not run this here because a container has no live-patched kernel). - Windows hotpatch: list installed updates with
Get-HotFix, which lists the updates Component Based Servicing installed (columns Source, Description, HotFixID, InstalledBy, InstalledOn), and look for the month's KB number:(Get-HotFix | Sort-Object -Property InstalledOn)[-1]shows the most recent one, andGet-HotFix -Id KB5099999(an illustrative number) returns nothing if that update is absent. For Azure VMs, the VM's Updates page in the Azure portal also shows the hotpatch status.
Step 2: Hardening without downtime
Hardening changes are mostly configuration, not binaries. Roll them out as rolling, per-node changes behind the same redundancy, with the change applied to one node, verified, then repeated. Settings that need a service restart (not a host restart) are handled by draining that node. Build the settings into the golden image so new instances arrive hardened, rather than modifying running hosts.
Step 3: What to tell the business about the remaining risk
State it as a measurable exposure window, an owner, and a decision.
Worked example (illustrative numbers, stated as assumptions): suppose after the techniques above 40 of 500 hosts can only be restarted in a single scheduled window per quarter. The longest a fix can wait for such a host is the gap between windows, about 91 days (365 divided by 4, rounded to the nearest day). Live patching or hotpatch narrows that for kernel and Windows security fixes, but not to zero.
| Question the business asks | Answer you give |
|---|---|
| What is still exposed? | 40 hosts, listed, with the services they run |
| For how long? | Up to about 13 weeks for fixes that need a restart; days for those covered by live patching or hotpatch |
| What reduces it? | Compensating controls on those hosts (tight allow-lists, removal of unused services), plus one extra short window per quarter or funding a second node so the host becomes rolling-updatable |
| Who accepts what remains? | A named executive signs a risk acceptance with an expiry date, reviewed each quarter |
Recommendation: ask for the smallest fixed window that bounds exposure, for example two hours monthly for the 40 hosts, rather than none; each shorter gap proportionally shortens worst-case exposure (monthly gives about 30 days, a third of the quarterly gap).
Trade-offs and pitfalls
- Live patching is a mitigation, not a replacement for reboots: a kernel can be live-patched for months and still drift from the tested baseline.
- Replace-instead-of-patch moves risk into the image pipeline; the pipeline must be patched and tested too.
- Do not hide the residual risk in a spreadsheet. If an incident occurs on an accepted host, the signed acceptance is what shows the business chose the exposure knowingly.
- What would change the recommendation: if the 40 hosts include internet-facing ones, the compensating controls alone are not enough and the second node becomes the priority.
List and interpret common ICMP types and codes you may encounter during ping/traceroute: Destination Unreachable (with host/network/port codes), TTL Expired, Fragmentation Needed, Redirect, and more. For each, explain what the ICMP message tells you about the network problem and what diagnostic step you would take next.
Sample Answer
Direct answer
The ICMP messages you'll see most often during ping/traceroute each map to a specific, distinct network condition: Destination Unreachable (with several sub-codes for host, network, or port), TTL/Time Exceeded, Fragmentation Needed, and Redirect, and reading which specific one you got (not just "it failed") tells you what to check next.
Structured elaboration
- Destination Unreachable, host unreachable: a router along the path has no route to the specific destination HOST; check routing tables at the router that generated this message.
- Destination Unreachable, network unreachable: a router has no route to the destination NETWORK at all, a broader failure than "host unreachable," usually indicating a missing or withdrawn route rather than just an ARP/reachability issue for one specific host.
- Destination Unreachable, port unreachable: the destination host IS reachable, but nothing is listening on the specific port being probed; this is actually the NORMAL, expected response that classic UDP-based traceroute relies on to detect it has reached the final destination (since it deliberately probes an unlikely-to-be-open port).
- TTL Exceeded (Time Exceeded): as covered by traceroute's core mechanism, a router decremented TTL to zero and discarded the packet; seeing this from an intermediate hop is normal and expected during traceroute, but seeing it unexpectedly during regular traffic (not traceroute) can indicate a routing loop or an unusually long path exceeding a low starting TTL.
- Fragmentation Needed (with the "don't fragment" bit set): a router needs to fragment the packet to forward it but is prohibited from doing so, and reports the MTU of the constrained link in its reply, which is exactly the message Path MTU Discovery depends on to let the sender adjust; its absence (being filtered somewhere) is the classic PMTUD black-hole failure mode.
- Redirect: a router is telling the sender that a better next hop exists for this destination than the one currently being used, typically because the sender's default gateway assumption is suboptimal for that specific destination; on modern networks, redirects are frequently disabled or ignored for security reasons (they can be spoofed to redirect traffic maliciously), so seeing them, or specifically NOT seeing them where you might expect one, both carry diagnostic meaning depending on context.
Worked example
A traceroute to a destination shows normal TTL-Exceeded replies through several hops, then a "Destination Unreachable, network unreachable" from an intermediate router instead of ever reaching the destination. This specific sub-code tells you the FAILURE is at the ROUTING layer at that specific router (it has no path to the destination network at all), distinct from what you'd see if the destination itself were simply not responding (which would produce silence or timeouts, not an explicit unreachable message) or if only a specific port were closed (which would be "port unreachable," implying the network path and host are both fine).
Trade-offs & pitfalls
Treating all "Destination Unreachable" messages as equivalent is a common mistake; the sub-code (host, network, port, and several others) carries real diagnostic information about WHERE and WHAT kind of failure occurred, and conflating them (for example, treating a normal "port unreachable" completion of a UDP traceroute as an error) leads to misreading perfectly healthy tool output as a failure.
Explain the differences between local user accounts and domain user accounts on Windows. In your answer include: authentication scope, where profiles and user data are stored, how password and lockout policies apply, single-sign-on behavior, and common scenarios where you would choose a local account over a domain account.
Sample Answer
Direct answer
A local account lives in the Security Account Manager (SAM), a small account database stored on one specific Windows machine, and is only ever authenticated by that machine; a domain account lives once in Active Directory and is authenticated by a domain controller (a server that holds the directory and validates logons), so the same identity works across every machine joined to that domain. Every other difference, where the profile lives, how password and lockout policy is applied, whether single sign-on (SSO, logging in once and being trusted by other resources without re-entering credentials) works, and when you'd deliberately choose one over the other, follows from that one structural fact.
Structured elaboration
| Aspect | Local account | Domain account |
|---|---|---|
| Authentication scope | Validated against the SAM database on that one machine only; unrecognized anywhere else unless deliberately duplicated | Validated by a domain controller, typically via Kerberos (with NTLM as a fallback); recognized on every machine joined to the domain |
| Profile and user data location | C:\Users\<username> on that single machine, tied to a security identifier (SID) that only means something locally | C:\Users\<username> on whichever machine is used, but tied to a domain-wide SID; a brand-new, empty local profile is still created on each new machine the first time that account logs in, unless roaming profiles or Folder Redirection are configured to follow the user across machines |
| Password and lockout policy | Set by the Local Security Policy on that one machine; can silently differ machine to machine | Set once by Group Policy (GPO, a domain-level policy object) at the domain, with optional Fine-Grained Password Policies to layer stricter exceptions on specific groups or organizational units; a lockout follows the account everywhere, not just one machine |
| Single sign-on behavior | None beyond the one machine; each machine's local account is an independent identity, so reaching a network resource normally means re-entering credentials | Real SSO: log in once at a domain-joined machine and the domain controller's Kerberos ticket lets you reach other domain resources (file shares, another domain machine, printers) without re-authenticating |
Why domain accounts get real single sign-on and local accounts don't. When you log into a domain-joined machine with a domain account, the domain controller issues a ticket (Kerberos's ticket-granting mechanism) that the machine can present on your behalf to other domain resources. A local account has no domain controller vouching for it; it is only ever known to the one machine that created it, so there is nothing for a second resource to trust without being handed the same raw credential again.
Common scenarios where a local account is the right choice over a domain account:
- A standalone machine that will never join the domain at all: an isolated workshop kiosk, or a host deliberately segmented off the domain's trust boundary for security reasons (a DMZ system).
- Break-glass emergency access: keeping a local administrator account specifically so a technician can still log in and diagnose the machine when the network path to a domain controller is unavailable, since local authentication has no dependency on reaching the domain at all.
- A machine-scoped service or application account: something intentionally confined to exactly one machine, since it structurally cannot be used to authenticate anywhere else, which shrinks the blast radius if that one credential is compromised.
- Initial setup and imaging, before the machine has even joined the domain yet.
Worked example
Two concrete situations make the distinction tangible. First: a branch office loses its wide-area network link to the data center overnight. A technician arriving on-site the next morning cannot authenticate a domain account on the affected workstation, because there is no reachable domain controller to validate it, but the local Administrator account (a break-glass account kept for exactly this case) still authenticates instantly, since the SAM database it checks against lives on that same machine. Second: a new hire, jsmith, already has a domain account used daily on their assigned desktop. On their first day working from a shared lab machine, logging in with the exact same domain account still creates a brand-new, empty C:\Users\jsmith profile on that lab machine (no desktop shortcuts, no saved settings) even though it is unambiguously the same identity, because without roaming profiles or Folder Redirection configured, the profile is per-machine while only the underlying account is domain-wide.
Trade-offs and pitfalls
The most common mistake is assuming the domain's password and lockout policy automatically covers the built-in local Administrator account on a domain-joined machine; it does not, since that account is still governed by local policy on that machine, which is exactly why Microsoft's LAPS (Local Administrator Password Solution) exists, to randomize and centrally manage that one account's password per machine rather than leaving it to drift. A second, genuinely dangerous pitfall is administrators duplicating the same local account name and password across many machines as an informal substitute for single sign-on convenience: this creates a real lateral-movement risk, because compromising the password hash on one machine effectively compromises the identical local account on every other machine sharing it, which is the exact mechanism pass-the-hash attacks rely on. A third pitfall is treating a fresh, empty local profile on a new machine as a sign that "the account didn't work," when it is simply expected behavior for a domain account without roaming configured. Finally, keeping any accounts local at all trades away the domain's centralized policy enforcement and access-review visibility in exchange for machine-level autonomy and resilience when the network is unavailable, so local accounts should be deliberately minimized to the specific cases above and tracked, since they fall outside routine domain-wide access reviews by construction.
Explain TCP Selective Acknowledgment (SACK): how SACK blocks are represented in the TCP options, and how SACK lets a sender avoid retransmitting segments the receiver already has after a single loss event. What does a sender do differently once SACK is enabled versus a sender using only cumulative ACKs?
Sample Answer
Direct answer
Selective Acknowledgment (SACK) lets a receiver tell the sender exactly which non-contiguous blocks of data it has ALREADY received, so after a loss the sender only has to retransmit the specific missing segment(s), not everything that came after it.
Structured elaboration
Without SACK, TCP uses cumulative acknowledgment: an ACK only confirms "I have received everything up through this byte, contiguously." If segment 3 of a 10-segment flight is lost but segments 4 through 10 all arrive fine, the receiver can only ACK up through the end of segment 2, it has no way to tell the sender "I actually already have 4 through 10, I'm just missing 3." A sender using only cumulative ACKs, upon detecting the loss, may end up retransmitting segments 3 through 10 (everything the receiver hasn't cumulatively acknowledged), even though 4 through 10 were never actually lost.
With SACK enabled (negotiated via a permitted option in the handshake, then carried on subsequent ACKs), the receiver's ACK can include SACK blocks, explicit ranges of sequence numbers it holds that are NOT contiguous with the main acknowledged run, in the example above, a SACK block spanning segments 4 through 10. Now the sender knows precisely that only segment 3 needs retransmitting.
Worked example
Say a sender has segments with sequence ranges [1000-1500), [1500-2000), [2000-2500), ... up to [4500-5000), and segment [2000-2500) is lost in transit while everything else arrives. Without SACK: the receiver's ACKs stay pinned at ack=2000 (the last contiguous byte received) even as segments up through 5000 keep arriving; the sender, upon detecting the loss (via duplicate ACKs all saying ack=2000), knows only that SOMETHING after 2000 needs resending and, in older/naive implementations, could resend everything from 2000 onward. With SACK: the same duplicate ACKs at ack=2000 now also carry a SACK block like sack=2500-5000, telling the sender explicitly that only the single segment [2000-2500) is actually missing, so it retransmits exactly that one segment and nothing else.
Trade-offs & pitfalls
SACK is most valuable on connections with a large amount of data in flight (a large window relative to segment size) and where losses are isolated rather than in a solid burst, since that's exactly the scenario where "retransmit everything after the gap" wastes the most bandwidth compared to "retransmit only the gap." On a connection with a tiny window, or where an entire flight is lost at once (nothing left to selectively acknowledge), SACK provides little advantage.
Explain the Linux Filesystem Hierarchy Standard (FHS). Describe the purpose and typical contents of these directories: /, /etc, /var, /usr, /opt, /home, /tmp, /srv, and /proc. For each, give one example of a file or service an SRE would place there and why that location is appropriate.
Sample Answer
What the FHS is and why it exists
The Filesystem Hierarchy Standard (FHS) is a specification maintained by the Linux Foundation that defines where categories of files belong on a Linux system: configuration here, variable runtime data there, user binaries somewhere else. Distributions (Debian, Red Hat Enterprise Linux, and most others) largely follow it. The payoff is portability of operational knowledge: an engineer who knows the FHS can find a service's config, logs, and data on an unfamiliar box without exploring first, and packaging tools and automation (Ansible, systemd unit files, backup scripts) can rely on paths existing in predictable places; systemd here is the init system and service manager most current Linux distributions use to start and supervise services.
The directories, what lives there, and a placement example
| Path | Purpose | Example an SRE would place there | Why that location |
|---|---|---|---|
/ | The root of the whole tree; everything else is mounted under it | N/A (it's the mount point) | Every other path is defined relative to it; it must exist before anything else can |
/etc | Host-specific, static configuration files (plain text, not binaries) | /etc/nginx/nginx.conf | Config is read at startup, edited by humans/config-management, and should never contain executables, so backup and audit tooling can treat /etc as "the config surface" |
/var | Variable data that changes while the system runs: logs, spool queues, caches | /var/log/nginx/access.log | Data here grows and rotates constantly, so /var is commonly its own filesystem/partition sized and monitored separately from the OS |
/usr | Read-only(-ish), distribution-managed program files and libraries shared across the system, installed by the package manager | /usr/bin/nginx, /usr/lib/x86_64-linux-gnu/libssl.so | Package managers own this tree; it can be mounted read-only or shared over NFS (Network File System, a protocol for exposing a filesystem tree to other machines over the network) by multiple hosts because it doesn't change per-host |
/opt | Self-contained third-party or vendor software that doesn't fit the distro's package layout | /opt/datadog-agent/bin/agent | Vendor installers ship a whole directory tree they own; keeping it out of /usr avoids collisions with package-manager-owned files |
/home | Personal directories for interactive users | /home/alice/.ssh/authorized_keys | Per-user data that should be backed up and quota-managed independently of the OS |
/tmp | Temporary files any user or process may need, normally cleared on reboot or by a periodic cleaner | A build tool's scratch directory, e.g. /tmp/pip-install-xyz | World-writable (with the sticky bit set, which means only a file's owner can delete or rename it) so any process can drop scratch data without needing a dedicated writable path |
/srv | Data served by this host to others (web content, FTP trees) | /srv/www/example.com/public_html | Separates "content this server serves to the world" from /var (the server's own operational state) and from /home (personal data) |
/proc | A virtual, in-memory filesystem exposing live kernel and process state; nothing here is on disk | /proc/<pid>/status to read a process's current memory and state | Because it's generated by the kernel on read, it always reflects the current instant, which is exactly what live diagnostics need and a real file could never guarantee |
Worked example: where a new internal tool's files land
Say the team ships an internal metrics agent as a .deb package. Its systemd unit file (defining how it starts) goes in /etc/systemd/system/, its own tunables go in /etc/metrics-agent/config.yml, its binary is installed by the package manager to /usr/bin/metrics-agent, and its running output (rotated logs, a small on-disk queue of unsent metrics) lives under /var/log/metrics-agent/ and /var/lib/metrics-agent/. If instead the vendor shipped it as a self-extracting tarball outside the distro's package manager, the whole tree would go under /opt/metrics-agent/ so it never collides with anything dpkg/rpm thinks it owns.
Trade-offs and pitfalls
/usrhistorically held a stricter separation from/(so/could mount before other filesystems were available at early boot), but most modern distributions merge/bin,/sbin, and/libinto/usrvia symlinks (the "usr-merge"). The FHS still describes the logical roles even where physical layout has consolidated.- Putting large or fast-growing data (a database, container images) directly under
/varwithout giving it its own filesystem is a common outage cause:/varfilling up can starve logging and package operations for the whole host, so production hosts typically give/var, and sometimes/var/logspecifically, a dedicated partition or volume. - Third-party software that ignores
/optand scatters files into/usr/localor/usrcan conflict with the package manager during upgrades;/usr/localis the FHS-sanctioned spot for locally-built software that isn't managed by the distro's packages, which is a subtly different bucket from/opt's vendor-installed trees.
Want to create your own tailored preparation guide using our deep research?
Get Started for FreeInterview-Ready Courses
Visual-first, interactive, structured learning paths
Browse Systems Administrator jobs
AI-enriched listings across hundreds of company career pages
Explore Jobs