Windows Server Administration Questions
Operating and maintaining Windows Server hosts and the roles they run: installing and managing server roles and features (Server Manager, PowerShell), building a production-ready server from first boot, Windows service configuration, startup types and dependencies, IIS web hosting, application pools and TLS bindings, Remote Desktop Services and administrative remote access, print services, NTFS and SMB share permissions, DFS namespaces, folder redirection and roaming profiles, Failover Clustering and clustered file services, Windows performance counters, Event Viewer diagnosis and event collection, PowerShell administration and remoting at scale, Just Enough Administration, backup and recovery of application and file servers, servicing channels and lifecycles across Windows Server releases including 2025, running WSUS (deprecated but still supported) or ConfigMgr as the Windows update infrastructure, and boot or offline recovery. Windows counterpart to Linux administration, often in mixed estates. Ground covered elsewhere: Active Directory domains and Group Policy, DNS and DHCP service design, identity and privileged-access models, security incident response, host security hardening baselines, and patch cadence, rollout rings and patch compliance.
Your company wants users' profiles and data to follow them across laptops, desktops and remote session hosts. Compare roaming profiles, folder redirection and container-based profile solutions, then recommend an approach for branch users on slow links and another for pooled virtual desktops, covering the trade-offs and storage you would plan for.
Sample Answer
Direct answer
Use Folder Redirection with Offline Files (Windows keeps a cached local copy of network files, so users can keep working when the link is down and changes sync back later) for branch users on slow links, keeping any roaming profile small and limited to their own computers, because a roaming profile is copied across the link at every sign-in and sign-out. Use FSLogix profile containers (a virtual hard disk file per user (VHD or VHDX) on a file share, attached at sign-in) for pooled virtual desktops, because there the profile must follow the user to whichever machine in the pool they land on and must not be copied at all. Size container storage from measured usage, not from the per-user ceiling.
The three approaches compared
| Roaming user profiles | Folder Redirection (with Offline Files) | Container-based profiles (FSLogix, user profile disks) | |
|---|---|---|---|
| How it works | The profile loads from a file share to the local computer at sign-in, merges with any local copy, and the local copy merges back to the server at sign-out | A known folder such as Documents points at a path on a file share; applications use it as if local | A virtual disk (VHD or VHDX) holding the profile is mounted at sign-in, and a filter driver (software that sits between applications and the file system and redirects file access) makes applications see a local profile |
| Network cost | Whole profile copied at sign-in and sign-out | Files fetched as used; Offline Files caches them for offline use | No copy; reads and writes go to the disk file |
| Sign-in time on a slow link | Grows with profile size | Small, since only settings roam | Small, as nothing is copied |
| Weak points | Large profiles, profile versions differ between Windows versions | Needs a reachable server, or Offline Files to hide an outage | Depends on file share availability and storage design |
| Best for | Small settings-only profiles on a few trusted computers | Document-centric users, branch offices, mobile users | Pooled or multi-session virtual desktops |
User profile disks (the Remote Desktop Services feature that keeps each user's profile in a VHDX file on a share) are also a container approach, and Microsoft's roaming-profile article points to them for Start menu roaming on session hosts. FSLogix is the current option for pooled desktops and is not limited to virtual desktops.
Branch users on slow links
Recommendation: Folder Redirection plus Offline Files for the data, and a roaming profile only if settings must follow users between computers. Both Folder Redirection and roaming profiles can be limited to a user's primary computers, using the Group Policy settings "Download roaming profiles on primary computers only" and "Redirect folders on primary computers only" (they depend on the msDS-PrimaryComputer attribute in Active Directory, an attribute on the user's account that lists the computers designated as that user's own). That keeps profile data off shared or conference-room machines.
The cost of getting this wrong is arithmetic. Assume a dedicated 10 Mbit/s branch link, a full-rate transfer and no protocol overhead:
10 Mbit/s500 MB×8=400 s10 Mbit/s50 MB×8=40 sA 500 MB roaming profile costs about 6.7 minutes at sign-in; redirecting the large folders so the profile is 50 MB brings that to about 40 seconds. The 500 MB and 50 MB figures are illustrative; measure your own profiles. Real transfers are slower, and sign-out repeats the copy in the other direction.
Storage and share design (from Microsoft's deployment guidance):
- Deploy Folder Redirection first on existing local profiles, then roaming profiles, so profiles start small.
- Keep the roaming profile share separate from shares used for redirected folders, to prevent inadvertent offline caching of the profile folder. The share's caching mode (
New-SmbShare -CachingMode) controls offline caching. - On a clustered file share, disable continuous availability (an SMB share setting that lets open files survive a failover to another cluster node) for roaming profiles to avoid performance problems. If the share sits behind Distributed File System (DFS) Namespaces, give the folder a single target, and if DFS Replication copies it, let users reach only the source server, or users will make conflicting edits.
- Do not place Folder Redirection, home directories or roaming profiles on a Scale-Out File Server; Microsoft marks them "not recommended" there because they generate many writes that must be committed immediately. Use a File Server for general use.
- If you support two Windows versions, keep separate profile versions per OS; each gets its own folder, so profiles double in count and storage, and changes made on one version do not roam to the other. Redirect common folders so files are visible on both.
- Use the permissions Microsoft lists for the profile share: System Full control; Administrators Full control on the folder only; Creator Owner (a built-in identity that stands for whoever created the item) Full control on subfolders and files only; the user group List folder and Create folders on the folder only; everything else removed. Access-based enumeration and Encrypt data access can be enabled on the share.
Pooled virtual desktops
Recommendation: FSLogix Profile Container for each user on Server Message Block (SMB) file shares, with everything in the single profile container. FSLogix also has a separate Office container (ODFC, which holds only Office data such as Outlook and OneDrive cache); it is optional and is mainly paired with another roaming-profile product.
Settings that matter (from the FSLogix configuration reference and examples):
| Setting | Value to use | Why |
|---|---|---|
Enabled | 1 | Required |
VHDLocations | One UNC (Universal Naming Convention, \\server\share) path to the share | Where containers live |
VolumeType | VHDX | Preferred over VHD for supported size and fewer corruption cases |
SizeInMBs | Default 30000 | The maximum container size; containers are extended to it at sign-in; it can be raised later but not lowered |
IsDynamic | Default 1 | Disk usage grows only as needed, up to SizeInMBs |
ProfileType | 0 (default) | A user's container is mounted through only one connection at a time, which keeps things simple and fast |
DeleteLocalProfileWhenVHDShouldApply | 1 | Prevents a stray local profile masking the container; Microsoft warns FSLogix permanently deletes the local profile in that case, so confirm nothing needed lives only in local profiles |
ClearCacheOnLogoff (Cloud Cache only) | 1 | Saves local disk on pooled machines |
Two of these settings are marked required in Microsoft's reference, so containers do not work without them: Enabled and VHDLocations (with Cloud Cache, CCDLocations takes the place of VHDLocations). VolumeType and SizeInMBs have defaults (vhd and 30000) that this design overrides or confirms, and the other rows tune behaviour or apply only when Cloud Cache is used.
Availability: listing several paths in VHDLocations does not give resilience. Microsoft describes the list as locations to search for the user's container, and if none holds one a new container is created in the first listed location, so a user whose existing container is not found in any searched location gets a new, empty profile. Microsoft's high-availability article also states that standard containers provide no resiliency of their own and rely entirely on the storage provider, and each container still lives in exactly one location, so a lost share still stops the users whose containers are on it. For resilience use Cloud Cache (an FSLogix option that keeps a local cache and writes each container to two or more storage providers, meaning separate storage locations such as file shares or Azure storage accounts; CCDLocations lists them, with at least two), and set HealthyProvidersRequiredForRegister to 1 as in Microsoft's example so a user cannot start a local cache when a provider is down. Do not change FlipFlopProfileDirectoryName, SIDDirNameMatch, SIDDirNamePattern, VHDNameMatch or VHDNamePattern on an existing deployment, because those five settings control how FSLogix names and finds each user's folder and disk, so after a change it looks for the old containers under new names, will not find them and will create new empty profiles.
Currency note: Microsoft's FSLogix pages warn that a Windows change in the April 2026 Windows Server update, still worded as upcoming on those pages although that month has now passed, moves the default Kerberos encryption type (Kerberos is the Windows sign-in ticket protocol) from RC4, an older cipher, to AES-SHA1, a stronger encryption and integrity combination, and that file shares hosting containers that have not been upgraded to AES-SHA1 might have access problems afterwards. The link to profiles: the session host reaches the container share over SMB and proves who the user is with Kerberos tickets, so a change in the default ticket encryption can break access to the share unless the storage also supports AES-SHA1. Confirm the storage accounts and file servers hosting containers support AES-SHA1 before that update is installed, and if the session hosts already have it, check now.
Storage planning for 500 users at the default ceiling:
500×30000 MB=15000000 MBConverting that needs one stated convention. Treating a megabyte as 1,048,576 bytes (so the unit is really a mebibyte, MiB), one container is 30000 / 1024 = 29.3 GiB and the total is 15,000,000 / 1,048,576 = 14.3 TiB. Treating a megabyte as 1,000,000 bytes, the total is 15,000,000 x 1,000,000 bytes = 15 TB. These are two readings of the same 15,000,000 "MB" figure, not one quantity in two units: 15,000,000 x 1,048,576 bytes is 15.7 TB (14.3 TiB), while 15,000,000 x 1,000,000 bytes is 15 TB (13.6 TiB), a gap of about 5 percent, so state which convention you use when you order storage. Either way it is the worst case when every container reaches its maximum, not the expected usage. Because dynamic containers grow only as used, run a pilot, measure the actual size of each container, and plan on a high percentile plus headroom. For illustration, if the measured average were 8 GiB, the demand would be 500 x 8 = 4,000 GiB, which is 4,000 / 1,024 = 3.9 TiB, well below the 14.3 TiB ceiling. Use the ceiling for alerts and quotas, and use measured data for purchases. Also plan antivirus exclusions (listed in the FSLogix prerequisites) and test the share and NTFS (New Technology File System) permissions before the pilot.
Pitfalls
- Using roaming profiles for pooled desktops: sign-in and sign-out time grow with the profile, because the profile is copied each time.
- Putting containers on a share that is not highly available and calling it done; a share outage then stops every sign-in.
- Letting the working data live in the profile. Redirect or keep documents in a separate store so profile size stays bounded.
- Assuming a container's size limit can be shrunk later; it cannot, so start with a deliberate ceiling.
A business-critical Windows service spikes to 100% CPU, stops responding for a few minutes, then recovers on its own. It only happens in production. How do you find out what it is doing during the spike without making things worse for users?
Sample Answer
Direct answer
Bound the problem with cheap counters (Windows performance counters are built-in meters for CPU, memory, disk and network), arm a trigger-based capture so the evidence is taken at the moment of the spike, and keep the capture light. The two tools are ProcDump, a free Microsoft Sysinternals command-line tool that writes a dump when a trigger fires, and Windows Performance Recorder (WPR), which records a CPU trace for the whole machine. A dump is a file holding a snapshot of a process's memory plus the call stack of every thread. A call stack ("stack" for short) is the list of functions a thread is in the middle of running, innermost first. A clone dump takes the snapshot from a copy of the process, so the original is paused only briefly.
Then read the stacks. Several dumps taken while the CPU stays high show whether the process is spinning (one thread repeating the same code without making progress) in one place or doing different work. Attaching a debugger or taking repeated full dumps of a busy service is what makes things worse.
Step 1: confirm and bound the spike
Get-Counter samples performance counters. Compare the machine total with the process counter:
Get-Counter -Counter '\Processor(_Total)\% Processor Time','\Process(*)\% Processor Time' -SampleInterval 2 -MaxSamples 5
\Process(name)\% Processor Time sums across cores: Learn's own example output shows _total at 395 on a multi-core machine. So 100 means one full core. On an 8-core server, one fully busy thread is 100 / 8 = 12.5 percent of the machine, which is why a service can freeze while the machine total looks calm. Add \Processor(*)\% User Time and compare it with % Processor Time to see how much of the busy time is outside user mode. User mode is the program's own code; kernel mode is Windows itself and its drivers working on the program's behalf, for example handling disk or network requests.
Step 2: arm the capture
Two tools, started before the next spike and left running:
procdump -accepteula -r -a -mp -n 3 -s 5 -c 90 -u -at 120 AppSvc D:\Dumps
-c 90 -utriggers at 90 percent of one core.-umakes the threshold per core, so it matches theProcesscounter's scale.-s 5requires the CPU to stay above the threshold for 5 seconds, so a blip does not trigger.-n 3writes up to three dumps before ProcDump exits.-rdumps a clone of the process, which keeps the pause short. On Windows 8.1 and later it uses a process snapshot.-askips a trigger that would suspend the process for a prolonged time because the concurrent dump limit is exceeded, and-at 120abandons a collection that takes more than 120 seconds.- Which switches matter:
-c,-uand-sdecide when a dump is taken.-r,-aand-atdecide how much the capture can hurt the service.-mpand-ndecide what is saved and how many files.-accepteulaonly accepts the Sysinternals licence prompt so an unattended run does not stop on it. -mpwrites a MiniPlus dump with the private memory needed to read stacks and heap objects. The default-mmmini dump is smaller but shows less. A full dump (-ma) is the largest, and ProcDump writes managed-code processes (common language runtime, CLR) as full dumps with-mpregardless.
The service name works in place of a process ID. ProcDump names each file PROCESSNAME_YYMMDD_HHMMSS.dmp by default, so the three files sort in time order. Write dumps to a data disk with room: if each full dump is about the size of a service's 6 GiB of private memory, three of them need about 3 x 6 = 18 GiB, and more once image and mapped memory are counted, because a full dump includes all memory. Learn puts a MiniPlus dump at 10 to 75 percent of a full one, except that managed (CLR) processes are written as full dumps, so budget for the full size there.
The watcher below records a CPU trace at the moment the service's own CPU stays high, plus a top-five process snapshot. Against stand-in Get-Counter, Get-CimInstance and wpr functions (the real tools run only on Windows) and a made-up sample sequence of 10, 95, 97, 20, 92, 98, 99, 96, 94, it behaves like this: the first run of two high samples is discarded because 20 resets the count, it triggers after five consecutive high samples (92, 98, 99, 96 and 94), writes the top-process CSV, and then calls wpr -start CPU and wpr -stop.
param(
[Parameter(Mandatory)][string] $ServiceName,
[double] $ThresholdCores = 0.9,
[int] $ConsecutiveSamples = 5,
[int] $SampleIntervalSeconds = 2,
[int] $RecordSeconds = 60,
[string] $OutDir = 'D:\Diagnostics'
)
$svcPid = (Get-CimInstance Win32_Service -Filter "Name='$ServiceName'").ProcessId
$instance = (Get-Counter -Counter '\Process(*)\ID Process').CounterSamples |
Where-Object { $_.CookedValue -eq $svcPid } |
Select-Object -First 1 -ExpandProperty InstanceName
$counter = "\Process($instance)\% Processor Time"
$hits = 0
while ($hits -lt $ConsecutiveSamples) {
# The process counter sums every core, so 100 means one full core.
$value = (Get-Counter -Counter $counter).CounterSamples[0].CookedValue
if ($value -ge ($ThresholdCores * 100)) { $hits++ } else { $hits = 0 }
Start-Sleep -Seconds $SampleIntervalSeconds
}
New-Item -ItemType Directory -Force -Path $OutDir | Out-Null
$top = (Get-Counter -Counter '\Process(*)\% Processor Time').CounterSamples |
Where-Object { $_.InstanceName -notin '_total', 'idle' } |
Sort-Object CookedValue -Descending | Select-Object -First 5 InstanceName, CookedValue
$top | Export-Csv -Path (Join-Path $OutDir 'top-processes-at-spike.csv') -NoTypeInformation
wpr -start CPU
Start-Sleep -Seconds $RecordSeconds
wpr -stop (Join-Path $OutDir 'cpu-spike.etl') "$ServiceName CPU spike"
What the stand-in run printed (real output from running the script above against the stand-ins, so the values are the made-up sequence, not a real server):
sample 10
sample 95
sample 97
sample 20
sample 92
sample 98
sample 99
sample 96
sample 94
wpr -start CPU
wpr -stop <OutDir>\cpu-spike.etl AppSvc CPU spike
and the file top-processes-at-spike.csv it wrote:
"InstanceName","CookedValue"
"appsvc","94"
"sqlservr","12"
"msmpeng","7"
Read the CSV like this: each value is a percentage of one core, so appsvc at 94 is almost one whole core, and the _total and idle rows are filtered out so the list shows real processes. Here the service itself is the busiest process by a wide margin, which points at its own code rather than at a neighbour such as the antivirus (msmpeng) or the database (sqlservr).
Windows Performance Recorder (WPR) wpr -start CPU records in memory mode by default, which is bounded. File mode (-filemode) writes to an unbounded file that can fill the disk, so avoid it unattended. The trace is an event trace log (.etl) file that Windows Performance Analyzer opens to show which functions were on the CPU.
Step 3: open the dumps and read the stacks
Open each .dmp file in WinDbg, Microsoft's debugger (part of Debugging Tools for Windows, which Learn lists as available standalone or with the Windows SDK and WDK). For a .NET service, Learn notes that the Visual Studio debugger is often the easiest way to start. Load symbols (the files that map addresses back to function names) first, otherwise the stack shows raw addresses. Two commands do most of the work, and both are documented on Learn:
!runaway 3lists how much user-mode and kernel-mode CPU time each thread has used. Learn describes it as a quick way to find which threads are spinning out of control. Learn says it works during live debugging or on dump files created by.dump /mtor.dump /ma, so if it prints nothing useful for a ProcDump file, compare the~*kcstacks instead of the thread times.~*kcprints a clean call stack for every thread in the process.~*means all threads, andkcshows only the module and function name on each line.
Illustrative output (made-up service and function names, trimmed to the user-mode time and to the busiest thread, which ~12kc prints on its own):
0:000> !runaway 3
User Mode Time
Thread Time
12:2b8c 0:00:47.3120
7:1f40 0:00:02.1090
0:000> ~12kc
# Call Site
00 AppSvc!PriceRules::MatchPattern
01 AppSvc!PriceRules::Evaluate
02 AppSvc!RequestWorker::Run
03 kernel32!BaseThreadInitThunk
04 ntdll!RtlUserThreadStart
How to read it: thread 12 has burned 47 seconds of CPU while the others used 2, so it is the one to look at. Its stack reads from the bottom up as the story of the thread: Windows started it (RtlUserThreadStart), it ran a request worker, which evaluated price rules, which is now inside MatchPattern, the function actually on the CPU. If thread 12 shows the same function at the top in all three dumps, taken seconds apart, it is stuck in that code, not finishing and moving on.
Step 4: read the evidence
| What you see | What it suggests |
|---|---|
The same thread and the same stack in all three dumps (in the example above, thread 12 with MatchPattern on top each time) | A loop or spin: one code path stuck (a retry loop, runaway regular expression, lock spin) |
| Many threads with different stacks, all in one library | Heavy legitimate work, such as a batch or a cache rebuild |
| Time mostly in kernel mode, small user time | A driver or a flood of I/O: look at the storage and network counters |
| Spike times line up with a scheduled task or antivirus scan | An external trigger: check the Task Scheduler history and the Application and System logs for that minute |
Step 5: fix and verify
Address what the stack names (a bad query, an unbounded loop, an unthrottled batch), then leave the watcher running through the next occurrence and compare.
Guardrails: not making it worse
- A dump pauses the process while memory is copied. The clone, the trigger duration and
-nall limit that pause. Never collect repeatedly without a cap. - Dumps contain the process's memory, including secrets. Restrict the folder's access list and delete the dumps after analysis.
- Do not attach a live debugger to a business-critical service in production.
- Sample a few named counters every 2 seconds. Avoid wildcard-everything runs on a struggling server.
Pitfalls
- Reading a process value of 100 on a 16-core server as a full machine: it is one core, 100 / 16 = 6.25 percent.
- Capturing after the spike is over: the evidence is gone, so the trigger must be armed first.
- Trusting one dump: you need several to tell a spin from busy work.
You have just built a new Windows Server and the team needs to manage it over Remote Desktop, but only from the management subnet 10.10.0.0/16 and only by clients that authenticate before a session is created. How do you set that up, and what else has to be true for remote access to work?
Sample Answer
Direct answer
Four settings, then a test from both sides. (1) Turn Remote Desktop on (registry value fDenyTSConnections = 0, where the registry is Windows' database of settings and a value is one named setting in it; or use Server Manager). (2) Require Network Level Authentication (NLA), which makes the client authenticate before the server builds a session, so unauthenticated clients never get a desktop. (3) Scope the built-in Remote Desktop Windows Firewall rule group so it only accepts the management subnet 10.10.0.0/16. (4) Make sure the right people are allowed in: members of the local Administrators group can already connect, and anyone else needs to be in Remote Desktop Users. Remote access also depends on the Remote Desktop Services service (TermService, the service name Windows uses for it) running and listening on port 3389, a network path that carries 3389 from the management subnet, and no Group Policy overriding your settings.
Setup
# 1. Allow Remote Desktop connections (0 = allow, 1 = deny)
Set-ItemProperty -Path 'HKLM:\SYSTEM\CurrentControlSet\Control\Terminal Server' -Name fDenyTSConnections -Value 0
# 2. Require Network Level Authentication on the RDP listener
Get-CimInstance -Namespace root/cimv2/TerminalServices -ClassName Win32_TSGeneralSetting |
Where-Object TerminalName -eq 'RDP-Tcp' |
Invoke-CimMethod -MethodName SetUserAuthenticationRequired -Arguments @{ UserAuthenticationRequired = 1 }
# 3. Enable the built-in Remote Desktop rules, but only for the management subnet
Get-NetFirewallRule -DisplayGroup 'Remote Desktop' |
Set-NetFirewallRule -Enabled True -RemoteAddress 10.10.0.0/16
# 4. Let the server-admin group in (Administrators can already connect)
Add-LocalGroupMember -Group 'Remote Desktop Users' -Member 'CONTOSO\ServerAdmins'
The registry value lives under HKLM\SYSTEM\CurrentControlSet\Control\Terminal Server, and -RemoteAddress accepts a subnet in prefix notation such as 10.10.0.0/16. The number after the slash says how many of the address's 32 bits are fixed. In 10.10.0.0/16 the first 16 bits (10.10) are fixed and the last 16 can be anything, so the range is 10.10.0.0 to 10.10.255.255, which is 2^16 = 65,536 addresses. A client at 10.10.4.20 is inside the range; 10.11.4.20 and 192.168.1.5 are outside it.
Step 2 is the least obvious line. The Remote Desktop settings are exposed through WMI (Windows Management Instrumentation, Windows' built-in management interface), which the CIM cmdlets query. Get-CimInstance ... Win32_TSGeneralSetting returns one settings object per listener, and Where-Object TerminalName -eq 'RDP-Tcp' keeps the one for the standard RDP listener. Invoke-CimMethod -MethodName SetUserAuthenticationRequired calls the method that sets that object's UserAuthenticationRequired property, and UserAuthenticationRequired = 1 means authenticate at connection time, which is NLA. The RDP listener uses port 3389 over both TCP and UDP, and the built-in group covers both, which is why the scoping is applied to every rule in the group rather than to one.
What else has to be true
- The service is up and listening:
TermServicerunning and bound to port 3389 (netstat -ano | find "3389", then match the PID, the process ID number that Windows gives each running program, toTermServicewithtasklist /svc). - No policy overrides you. If
fDenyTSConnectionsor the NLA setting keeps reverting, a Group Policy object is overriding the local value (Microsoft documents a policy-level copy offDenyTSConnections, and WMI exposes which source set the NLA value). To see the source, runGet-CimInstance -Namespace root/cimv2/TerminalServices -ClassName Win32_TSGeneralSetting | Where-Object TerminalName -eq 'RDP-Tcp' | Select-Object UserAuthenticationRequired, PolicySourceUserAuthenticationRequired. The second property reads 0 when the server's own setting is in force, 1 when Group Policy set it, and 2 when it is the default. A 1 means the value will keep reverting until you change the GPO. Fix it in the GPO, not on the server. - The user is permitted: Administrators, or members of Remote Desktop Users, and the account is not locked out.
- The path is open: any network firewall between the management subnet and the server must also allow 3389 from
10.10.0.0/16. Host firewall scoping is only one layer. - The client supports NLA. Older clients that cannot do NLA will be refused, which is the intent; do not turn NLA off to accommodate them.
- No broader rule exists. Another enabled inbound allow rule that matches port 3389 would let other networks in despite your scoped group. List them:
Get-NetFirewallRule -Direction Inbound -Enabled True -Action Allow | Get-NetFirewallPortFilter | Where-Object LocalPort -eq 3389.
Verify from both sides
# From a host inside 10.10.0.0/16: should succeed
Test-NetConnection -ComputerName newserver -Port 3389 -InformationLevel Detailed
# From a host outside that subnet: should fail (TcpTestSucceeded : False)
Test-NetConnection -ComputerName newserver -Port 3389
A successful detailed result looks like this (abbreviated, with illustrative addresses; Microsoft's Test-NetConnection examples show the full field list):
ComputerName : newserver
RemoteAddress : 10.10.1.25
RemotePort : 3389
InterfaceAlias : Ethernet
SourceAddress : 10.10.4.20
TcpTestSucceeded : True
TcpTestSucceeded is the line that matters. True means the host completed a TCP connection to port 3389, so the firewall let it through. From a host outside 10.10.0.0/16 the same command ends with TcpTestSucceeded : False, which is the result you want there. Note that this tests the network path only; a real RDP sign-in additionally needs NLA and the user's permission.
On the server, Get-NetFirewallRule -DisplayGroup 'Remote Desktop' | Get-NetFirewallAddressFilter shows the remote address of each rule as 10.10.0.0/16.
Trade-offs and pitfalls
- Scope by address and require NLA together. NLA stops unauthenticated clients from consuming a session; the address scope stops the rest of the network from even reaching the login. Either alone is weaker.
- Do not expose 3389 to the internet. Put administrative access behind a VPN or a gateway; the subnet restriction here assumes the management network is already trusted and segmented.
- Disabling the firewall to "see if that fixes it" removes the scope and is how this restriction gets lost in practice; test with the rule, not without it.
- Use a group, not individual accounts, for Remote Desktop Users, so access follows role changes.
Explain how NTFS permissions and share permissions work together to determine what a user can actually do on a Windows file share. How do explicit allow and deny entries and inheritance change the outcome?
Sample Answer
Direct answer
A user's access to a file over the network is decided in two independent layers, and the result is the more restrictive of the two. The share permissions (set on the shared folder, apply only to network access) are checked first; the NTFS permissions (the access control list, or ACL, stored on the folder or file in the file system) apply whether the user comes over the network or sits at the console. Within each layer, a user's allow entries from all their groups add up, and an explicit deny beats an allow. So: compute the user's rights in the share layer, compute them in the NTFS layer, and take the lower of the two.
How the two layers combine
- Share layer. The share has its own ACL with three levels: Read, Change and Full Control. Add up every allow that applies to the user or their groups; any deny removes what it covers. Result: the highest level the user holds.
- NTFS layer. The folder or file has an ACL made of access control entries (ACEs), each granting or denying rights such as Read, Modify or Full Control to a user or group. Again, allows from all of the user's groups add up and a deny wins over an allow at the same level.
- Effective access over the network = the lower of the two. A share that grants Change and NTFS that grants only Read gives Read. NTFS Full Control behind a Read-only share is still Read.
- How each layer is evaluated. Windows reads the entries of an ACL in order and stops as soon as every requested right has been granted by allow entries, or one requested right is hit by a deny. If it reaches the end with a requested right still ungranted, access is denied by default. So the order of entries matters, and the next section gives the order.
- Local access ignores the share layer. Share permissions do not restrict a person sitting at the server or in a remote desktop session on it; only NTFS applies. That is why the share layer is a coarse outer gate and NTFS is where the real detail belongs.
Explicit allow, explicit deny and inheritance
- No entry means no access. If nothing grants a right, it is implicitly denied.
- A deny beats an allow for the same user or group at the same level, even when the allow comes from a different group. That is why denies are best kept for excluding a small subset from a larger allowed group.
- Explicit permissions take precedence over inherited ones. An inherited deny does not block a user who has an explicit allow, and explicit entries are evaluated before inherited ones; Learn's order for an ACL is: explicit entries before inherited ones; within the explicit group, denies before allows; inherited entries next, those from the parent folder first, then the grandparent and so on, with denies before allows at each level. The Hana row in the worked example below shows an explicit allow beating an inherited deny, and the Cleo row shows an explicit deny beating an explicit allow.
- Inheritance means a folder's ACEs flow down to the files and subfolders created in it. Breaking inheritance on a subfolder (and removing the inherited entries) is how you give a subfolder a different, narrower list.
- Because the two layers are independent, changing one never changes the other, which is also how people end up debugging "I gave them Modify but they cannot save": the share layer is still Read.
Worked example
Share \\fs1\Finance (folder D:\Shares\Finance). Share permissions: Finance-Staff = Change, Finance-Leads = Change, Auditors = Read, Interns = Read. NTFS on the folder: Finance-Staff = Modify, Auditors = Read and Execute, Finance-Leads = Full Control, Interns = explicit deny of Modify. The subfolder Payroll has inheritance broken and only Payroll-Team = Modify. The subfolder Reports keeps inheritance on (so it receives all of the folder's entries as inherited entries, including the Interns deny) and adds one explicit entry of its own: Finance-Staff = Modify.
| User (member of) | Where | Share layer | NTFS layer | Effective over the network |
|---|---|---|---|---|
| Ana (Finance-Staff) | Finance root | Change | Modify | Modify-level (change files) |
| Ben (Auditors) | Finance root | Read | Read and Execute | Read |
| Cleo (Finance-Staff, Interns) | Finance root | Change | Modify allowed through Finance-Staff, but the explicit deny of Modify for Interns is read first | None: the deny wins over the Finance-Staff allow |
| Dev (Finance-Staff) | Payroll | Change | No entry for him | None |
| Eli (Payroll-Team, Finance-Staff) | Payroll | Change | Modify | Modify-level |
| Fay (Interns) | Finance root | Read | Explicit deny of Modify | None |
| Hana (Finance-Staff, Interns) | Reports | Change (Finance-Staff Change and Interns Read add up to Change) | Modify: the explicit allow is read first and already grants every right asked for, so the inherited deny of Modify is never reached | Modify-level |
| Gus (Finance-Leads) | Finance root | Change | Full Control | Change level: the share caps him, so over the network he cannot do what Full Control would allow beyond Change |
The table is a model that treats Change and Modify as the same level and takes the lower of the two layers. Each row applies that rule. Every user in the table has a share entry through one of their groups (Gus through Finance-Leads, which is why that group is in the share list above); a user with no share entry at all would get no network access whatever the NTFS list says.
The Hana row, step by step. On Reports the entries are read in this order: (1) explicit allow, Finance-Staff, Modify; then the inherited ones: (2) deny, Interns, Modify; (3) allow, Finance-Staff, Modify; (4) allow, Auditors, Read and Execute; (5) allow, Finance-Leads, Full Control. Hana asks to write a file. Entry 1 matches her through Finance-Staff and grants write, so every requested right is granted and the check stops before entry 2. Cleo has the same two groups but on the Finance root, where the Interns deny is explicit and is read before any allow, so she gets nothing. Two people with the same two group memberships end up with opposite results because of where the deny sits, which is why explicit denies on a parent folder deserve a second look when a subfolder is meant to behave differently. Gus is the reverse case: the share layer caps him over the network, but he can still use the full NTFS rights by connecting locally (for example, through a remote desktop session on fs1), since the share layer does not apply there.
Trade-offs and pitfalls
- Recommendation: keep the share permissive and simple (for example Authenticated Users = Change, or per-group Change and Read) and do all fine-grained control with NTFS groups. One ACL to audit beats two that must agree.
- Avoid deny entries where possible. Cleo's problem is fixed by removing her from a group, not by adding another deny that later surprises someone.
- Grant to groups, never to individuals, so that access follows role changes.
- Test as the user. Compute the answer on paper, then confirm the NTFS half with the Effective Access tab in the folder's advanced security settings, which estimates a chosen user's rights from the NTFS ACL, and check the share layer separately (for example,
Get-SmbShareAccess -Name Finance). - Tools:
New-SmbSharecreates the share and sets the share layer with-FullAccess,-ChangeAccessand-ReadAccess;icaclsand the Security tab manage the NTFS layer.
You are building a 4-node Windows failover cluster stretched across two datacenters. How do you validate it before production, and which quorum and witness configuration would you choose to avoid split-brain when the link between sites fails? Work through what happens if one site goes down.
Sample Answer
Direct answer
Use two nodes in each datacenter plus a witness (a small tie-breaking vote that holds no workload) in a third location. That gives 5 votes, and the cluster needs a majority, which is 3. Validate with Test-Cluster before any production role is placed on it. If one datacenter fails, the other keeps 2 node votes plus the witness vote, so it holds 3 of 5 and carries on without human action. If only the link between the sites fails, both sides are alive, but only the side that can still reach the witness reaches 3 votes, so exactly one side keeps running (the witness can give its vote to only one side) and split-brain (two halves each believing they own the same data) cannot happen.
Terms used here
-
Heartbeat: a small message the nodes send each other every second or so to show they are alive. A node whose heartbeats stop arriving for long enough is treated as unreachable.
-
Fault domain: a named group of nodes that tend to fail together, such as all the nodes in one site; the cluster uses it to place roles and replicas sensibly.
-
Storage Replica: the Windows Server feature that copies a volume's blocks from a server in one site to a server in another, so each site holds its own copy.
-
Asymmetric storage: each site sees only its own disks, not the other site's, which is why validation reports storage errors on a stretch design.
-
Forced quorum: an administrator command that makes one node start the cluster without a majority; it is a last resort, covered under "What happens when one site goes down".
-
Quorum: the cluster may run only while it can see more than half of the total votes. It exists to prevent split-brain.
-
Vote: each node has one vote, and so does an active witness.
-
Dynamic quorum and dynamic witness: the cluster adjusts votes automatically so the total stays odd. The witness vote is switched on when the number of voting nodes is even and off when it is odd.
Validate before production
- Run the validation tests.
Test-Cluster -Node ...runs cluster, inventory, network, storage and system test categories (among others) and writes a report (-ReportNamenames it).Test-Cluster -Listshows every test and category, and-Includeor-Ignorenarrows a run. - Run storage tests before roles go live. Storage tests skip disks that are online and owned by a running clustered role, because testing them means taking them offline. To include them you must stop the role first, or pass
-Diskwith-Force, which takes the disk offline for the duration. Doing this on a live cluster is an outage, which is why validation belongs before production. - Expect noise from asymmetric storage. In a stretch design each site has its own storage replicated to the other (for example with Storage Replica), so Microsoft's own walkthrough says to expect storage errors in validation. Read each failure and confirm it is explained by the topology instead of waving the whole report through.
- Check the replication link. With Storage Replica,
Test-SRTopologychecks the Storage Replica requirements, including the link. Synchronous replication (a write is acknowledged only after both sites hold it, so no data is lost) is recommended at about 5 ms average round-trip latency. A slower link means asynchronous replication, which can lose recent writes that had not yet reached the other site. - Multi-subnet detail. If each site has its own subnet, the cluster IP address needs an additional address resource in the other site, joined with an OR dependency.
- Run a failure drill. Before go-live, power off both nodes in one site (the Storage Replica guide describes cutting power to both nodes) and confirm roles restart in the other site. Repeat for the other site, then for a cut inter-site link.
$nodes = 'node-a1', 'node-a2', 'node-b1', 'node-b2'
# 1. Validate before the cluster carries production roles.
Test-Cluster -Node $nodes -ReportName 'StretchCluster-PreProd'
# 2. Describe the sites so the cluster knows which nodes fail together.
New-ClusterFaultDomain -Name SiteA -Type Site -Description 'Primary' -Location 'Datacenter A'
New-ClusterFaultDomain -Name SiteB -Type Site -Description 'Secondary' -Location 'Datacenter B'
Set-ClusterFaultDomain -Name node-a1 -Parent SiteA
Set-ClusterFaultDomain -Name node-a2 -Parent SiteA
Set-ClusterFaultDomain -Name node-b1 -Parent SiteB
Set-ClusterFaultDomain -Name node-b2 -Parent SiteB
(Get-Cluster).PreferredSite = 'SiteA'
# 3. The fifth vote: a cloud witness reachable from both sites over HTTPS (port 443).
Set-ClusterQuorum -CloudWitness -AccountName '<StorageAccountName>' -AccessKey '<StorageAccountAccessKey>'
Get-ClusterQuorum
The site definitions let the cluster treat the two nodes in a site as one failure domain, and PreferredSite names where roles and the replication source should sit when everything is healthy. An OR dependency, in the multi-subnet case, means the cluster name comes online if either of its two IP address resources (one per subnet) is online. The cluster name must be 15 characters or fewer.
Choosing the witness
| Option | Needs | Choose it when |
|---|---|---|
| Cloud witness | Azure Standard general purpose v2 storage account, outbound HTTPS (port 443) from every node | Both sites have internet egress and no third datacenter exists |
| File share witness | An SMB (Server Message Block) 2 or later share, on a Windows server in the same Active Directory (AD) forest if domain joined, dedicated to this cluster, physically separate from both sites | A third site, such as a colocation facility, already exists |
| Disk witness | A shared disk visible to all nodes | Not suitable here, because the two sites do not share storage |
My pick is the cloud witness: it sits in Azure Storage, outside both datacenters, and needs no server to patch. What would flip it: no internet egress from one site (then that site can never count the witness vote), or a rule against cloud dependencies (then use a file share witness on a server in a third location).
Vote arithmetic
Total votes: 4 nodes + 1 witness = 5. Majority: floor(5/2) + 1 = 3.
| Event | Votes the side that keeps running holds | Needed | Outcome |
|---|---|---|---|
| Inter-site link cut, both sites reach the witness | The side that takes the witness first has 2 + 1 = 3, the other has 2 | 3 | One side continues, the other stops its cluster service |
| Site A powered off, witness reachable | Site B has 2 + 1 = 3 | 3 | Site B continues, roles restart there |
| Site A off and witness unreachable at the same moment | Site B has 2 | 3 | Cluster stops, forced quorum needed |
| Witness placed inside Site A, Site A off | Site B has 2 | 3 | Cluster stops: the witness failed together with the nodes it should protect |
| Link cut and witness unreachable from both sides | 2 and 2 | 3 | Neither side has a majority, both stop (safe, but an outage) |
| 4 nodes, no witness | Dynamic quorum zeroes one node's vote, total 3, majority 2 | 2 | In a 2-and-2 split the side holding two voting nodes has 2 of 3 and continues; the side with only one voting node has 1 of 3 and stops. Which side that is depends on which node had its vote zeroed (worked example below the table) |
Why one side gets the witness: the witness is a single shared resource (in the cloud witness case, a blob in the storage account that the cluster uses for voting arbitration), and it can give its vote to only one side at a time. When the link is cut and both sides can still reach it, the two sides compete, and the side that claims it first holds the vote while the other is refused. The votes alone do not favour either site, because both sides hold two nodes. The PreferredSite setting in the script says where roles and the replication source should sit when everything is healthy, but this design does not depend on it to settle a split: the outcome is safe whichever side wins. To influence which side survives, set PreferredSite, then repeat the cut-link drill from step 6 and record which side actually keeps running instead of assuming.
Worked example for the last row, with nodes a1, a2 in Site A and b1, b2 in Site B and no witness. The cluster has four voting nodes, which is even, so dynamic quorum zeroes one node's vote automatically. Say it zeroes b2: votes are a1 = 1, a2 = 1, b1 = 1, b2 = 0, total 3, majority 2. The link is cut and each side counts: Site A has 1 + 1 = 2 of 3 and keeps running; Site B has 1 + 0 = 1 of 3 and stops. Had the cluster zeroed a2 instead, Site A would hold 1 of 3, Site B 2 of 3, and Site B would win. The designer does not choose, which is the reason to add the witness: with it the total is 5 and the outcome no longer depends on an arbitrary node.
What happens when one site goes down
- The nodes in the surviving site stop receiving heartbeats from the failed site's nodes.
- They count votes: 2 nodes + witness = 3 of 5. They hold quorum, so the cluster stays online.
- Roles that were on the failed nodes are restarted on the surviving nodes.
- Storage: with synchronous Storage Replica the replica in the surviving site becomes the writable copy with no lost data. With asynchronous replication the unreplicated tail of writes is lost.
- When the failed site returns, its nodes rejoin automatically and replication resynchronizes.
If the witness is also gone, the remaining nodes have 2 of 5 and stop. As a last resort run Start-ClusterNode -ForceQuorum on one surviving node. That node's copy of the cluster configuration becomes the authoritative copy and is pushed to the others, so recent configuration changes can be lost. When the other site comes back, start its nodes with -PreventQuorum so they do not form a second competing cluster, then let them rejoin.
Pitfalls
- A witness in either datacenter, or on the same power or network path as one.
- A file share witness on DFS (Distributed File System) or replicated storage: not supported, because it can let two halves each run independently.
- Rotating the storage account key behind a cloud witness without updating every cluster that uses it. Update with the secondary key first, then regenerate the primary.
- Treating forced quorum as routine. It trades the safety guarantee for availability.
Unlock Full Question Bank
Get access to all 34 Windows Server Administration interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.