Active Directory Architecture and Management Questions
Designing, operating and recovering Active Directory Domain Services and the directory estate around it. Covers logical and physical structure (forests, trees, domains, OUs, trusts, schema extension, FSMO roles, Global Catalog, RODCs), sites and replication topology, domain controller placement, promotion, upgrade and functional levels, DC locator and AD-integrated DNS, domain join, Kerberos and NTLM authentication including SPNs and delegation, token size and SID history, LDAP binds and query tuning, Group Policy design, processing order, filtering, deployment and troubleshooting, user, group, computer and service account management including PowerShell account scripting and account lockout investigation, delegation of control and tiered administration, fine-grained password policy, backup, authoritative restore, forest recovery and USN rollback, AD hardening against Kerberoasting, DCSync and Golden Ticket attacks, forest migration, and hybrid identity with Microsoft Entra ID (Microsoft Entra Connect, Cloud Sync, password hash sync, pass-through authentication, federation, password writeback). Questions are asked from the directory administrator's and architect's seat. Platform-neutral identity protocols and lifecycle design, Windows file-server administration (shares, NTFS permissions, profiles), Linux directory integration, and generic DNS and DHCP service operations are covered elsewhere.
You must consolidate two legacy AD forests into a new forest with a different namespace. Lay out your migration plan and what you do about the problems it will throw up.
Sample Answer
Direct answer
Run it as a coexistence migration in waves: build the new forest, connect it to the old with a trust, migrate groups with their members, then users, then workstations and servers, carrying each account's old SID (security identifier, the unique ID that permissions are actually written against) in sIDHistory (an account attribute that stores those old SIDs) so access keeps working while resources are re-permissioned, then clean up. Keep the old accounts (disabled, not deleted) until each wave is validated so rollback is real. The problems to plan for are SID filtering on the trust, SID-history token bloat, UPN and account-name conflicts in the new namespace, passwords, service accounts and applications, and cloud identity matching.
Plan
| Phase | Work | Exit gate |
|---|---|---|
| 0. Inventory | Users, groups (with nesting and scope), computers, service accounts, SPNs (service principal names, the names Kerberos uses to find services), trusts, Group Policy objects, DNS, applications that hard-code domain names, Entra Connect sync | Signed-off migration scope and a naming map from old to new |
| 1. Foundation | New forest, DNS name resolution between the two namespaces, a trust, tool servers and permissions in both forests | Test migration of a pilot user passes logon and file access |
| 2. Pilot wave | A small friendly group, including a service account and one application | Rollback exercised once |
| 3. Production waves | Groups first with their members, then users, then computers; resource ACLs re-pointed at the new SIDs | Per-wave validation list passed |
| 4. Cutover | Last sync of changes, source accounts disabled | Helpdesk quiet for an agreed period |
| 5. Cleanup | Remove sIDHistory, remove the trust, retire the old forest | Token sizes and ACLs checked |
One user, start to finish (illustrative)
Ann Lee has account ann in the old forest, with SID ending in RID 1105 in the old domain. The file server fs01 has a folder whose permission list names that old SID. Before migration, Ann's token holds her old SID plus her old group SIDs, and she opens the folder. After migration, the new account ann gets a new SID ending in RID 2210 in the new domain, and the old SID is copied into its sIDHistory. Her new token now holds the new SID, the old SID and her migrated groups, and fs01 finds the old SID in the token and grants access, with no change to the folder. Cleanup later re-points the folder's permissions to the new SID, after which the sIDHistory value can be deleted and the token shrinks.
Tooling: ADMT or alternatives
The Active Directory Migration Tool (ADMT) v3.2 is Microsoft's free tool for migrating users, groups and computers between domains, inter-forest or intra-forest, plus security translation for local user profiles. Its download page describes v3.2 as the final release (re-released with bug fixes and no new features) and lists Windows Server 2003 through 2012 R2 as the system requirements. So check before committing whether your source and target versions fit.
| Option | Choose it when |
|---|---|
| ADMT | Both forests sit on versions it lists, and you accept that v3.2 is the final release |
| A commercial migration suite | Target forest runs newer Windows Server than ADMT supports, or you need vendor support and rollback tooling |
| Create fresh accounts and re-permission resources | Few resources and little shared data, so SID history is not worth its risks |
My recommendation is to decide on SID history first: if the estate has large shared file, print and application ACL footprints, use a supported tool that preserves SID history and plan the cleanup; if it does not, skip SID history entirely and re-permission.
SID history and why it is delicate
When a user is migrated with SID history, the old SID goes into sIDHistory on the new account. At logon the access token gets both the new SID and the old SID, and resources compare ACL entries against both, so a file server still listing the old SID grants access without any ACL change. The same applies to groups: global groups can only contain members from their own domain, so when a user moves, the global groups they belong to must move too, and their old SIDs can be kept in the new group's sIDHistory.
Problems it creates:
- SID filtering. SID filtering is a trust setting that discards any SID in a user's token that does not belong to the trusted domain. Because
sIDHistorySIDs belong to the old domain, a filtering trust throws them away and the old access stops working, which is why a migration needs filtering relaxed. A trust normally filters out SIDs that do not belong to the trusted domain, precisely to stop someone who controls a domain controller (DC) in a trusted domain from writingsIDHistoryto grant themselves rights. A migration trust needs that filtering relaxed for the migration period, and it must be turned back on afterwards. For an external trust that is the quarantine setting (Active Directory Domains and Trusts, ornetdom trustwith/Quarantine:No); for a forest trust,netdom trusthas/EnableSIDHistory:Yes, which Microsoft documents for outbound forest trusts and says to enable only if you trust the administrators of the trusted forest. - Token bloat. Every
sIDHistorySID counts toward the user's token. Combined with nested groups and the roughly 1,010 group SID limit in the access token, long-lived SID history causes slow or failing logons. Plan to re-point access control lists (ACLs, the permission lists on files and folders) at the new SIDs during the waves, and removesIDHistoryin the cleanup phase once the ACLs are re-pointed, because the rollback path below relies on it until then. - Dependencies. Anything that authorizes by the old SID keeps working only because of SID history, for example permissions on roaming profile shares, certification authority permissions or software distribution shares, so removal needs testing against those too.
UPN and account-name conflicts
- A UPN (user principal name, the name@suffix sign-in name) must be unique across the whole forest. In the new namespace you will add a new UPN suffix and decide the mapping rule up front, for example keep the same prefix and change the suffix.
- A
sAMAccountName(the pre-Windows-2000 logon name) must be unique within a domain. If two source forests havejsmith, one of them collides when both land in the same new domain. Detect collisions in the inventory phase, not during a wave: export both forests, group by name, and resolve each clash with a naming rule and an owner. - For Microsoft 365 or Entra ID sync, the new UPN suffix must be added and verified as a custom domain in the tenant. A suffix that is not verified makes Entra ID compute a
<mailNickname>@<tenant>.onmicrosoft.comsign-in name instead. A UPN change on a synced user also triggers recalculation of the cloud UPN, so users must sign in with the new name. - Duplicate
proxyAddressesor UPN values across the two forests cause Entra sync errors such asAttributeValueMustBeUnique.
Cloud identity matching
The cloud object is tied to the on-premises one by the sourceAnchor attribute (also called immutableId): one value that must stay the same for the life of the object, so Entra Connect can tell that the cloud user and the on-premises user are the same person. A soft match, by contrast, is Entra's fallback of matching on attributes such as the primary email address when the anchor does not line up, and InvalidSoftMatch is the error raised when that attempt finds a conflict. If the old forest used objectGUID, the new account in the new forest has a different objectGUID, so Entra Connect cannot match it and raises InvalidSoftMatch or creates a duplicate. The fix is to use ms-DS-ConsistencyGuid as the sourceAnchor and copy the old value into the new account's ms-DS-ConsistencyGuid before the new account is first imported by Entra Connect, since the value cannot be changed after the object is imported. How to copy the value, with illustrative numbers: suppose Ann's old account has objectGUID 6f1c2a3e-4b5d-4e6f-8a9b-0c1d2e3f4a5b. Its sourceAnchor in the cloud is the Base64 form of that GUID's 16 bytes, here Piocb11Lb06KmwwdLj9KWw==. Before the new account is first imported, write the same 16 bytes into the new account (the account running these commands needs permission to write this attribute on the new account):
$oldUser = Get-ADUser -Identity ann -Server old-dc.old.example
$bytes = $oldUser.ObjectGUID.ToByteArray()
Set-ADUser -Identity ann -Server new-dc.new.example -Replace @{ 'mS-DS-ConsistencyGuid' = $bytes }
The first line reads the old account, the second takes its GUID as 16 raw bytes (the Base64 text above is just how those bytes are shown in the cloud), and the third replaces the attribute on the new account, using its LDAP name as the -Replace parameter requires. If the old forest already stored a value in its own mS-DS-ConsistencyGuid, copy that value instead. Check the result by reading the attribute back before the first sync. One current caution: since 1 July 2026 Microsoft Entra ID blocks a hard match (a match on the anchor value) when the target cloud user already has onPremisesObjectIdentifier set or is assigned or eligible for a privileged role, and the export fails (Microsoft lists InvalidHardMatch for the privileged-role case and an AttributeUpdateNotAllowed error or a hard-match hardening message when onPremisesObjectIdentifier is already set). A user already synced from the old forest may fall into that group, so prove this method on a pilot account in the real tenant, and follow Microsoft's documented hard match recovery paths for an existing tenant, before the production waves. Never let both forests project the same user at once; scope each wave so one forest is authoritative, and inspect the pending exports on a staging server first.
Passwords
Choose per wave. Either carry existing passwords using a tool mechanism that supports it (confirm the tool's requirements in its own documentation), or issue temporary passwords with a forced change at next logon plus self-service reset. If users are synced to Entra ID, the "must change at next logon" behaviour for synced users relies on the Entra feature that supports temporary passwords, which Microsoft says should only be used when self-service password reset and password writeback are enabled.
Service accounts, computers and applications
Service accounts need the owner of each application, since their SPNs, delegation settings and stored credentials all change. Computers are re-joined and local profiles are translated to the new SIDs. Applications that reference the old domain by name need a test plan and a cutover date of their own.
Rollback
- Source accounts stay intact and are only disabled after the wave passes validation.
- The trust and SID history stay until cleanup, so a user can sign back in to the old account and still reach migrated resources.
- To roll a wave back: re-enable the old accounts, reverse the Entra sync scoping for that wave, and check the pending exports against the accidental-delete threshold (500 by default) before exporting.
- Do not decommission the old forest until the final wave has been stable for an agreed period.
Pitfalls
- Leaving the migration trust unfiltered after cutover.
- Migrating users before the global groups they depend on.
- Discovering UPN or
sAMAccountNamecollisions during a production wave. - Forgetting that
sIDHistoryis part of every token.
You must allow limited resource access among three separate forests while keeping most resources isolated. Propose a trust topology, justify the direction and type of each trust, and explain the security risks each trust introduces and how you contain them.
Sample Answer
Direct answer
Use forest trusts, one-way, between a resource forest and each account forest that needs it, with selective authentication turned on, and create no trust between the two forests that should stay isolated. Forest trusts are transitive inside the two forests they join (transitive here means the trust covers every domain in both forests, not only the two root domains where it was created) but they do not chain onward to a third forest, so isolation of the third is the default, not something you bolt on. Then contain each trust with selective authentication (outside users may sign in only to computers you name), SID history disabled (old identifiers from migrated accounts are ignored), delegation blocked (a service cannot act as the user on other servers) and tight group design. Each of these is explained where it is used below.
Scenario and topology
Contoso (contoso.com, forest A) runs the shared file servers and an application. Fabrikam (fabrikam.com, forest B) is an acquired company whose users need the Contoso application. Wingtip Toys (wingtiptoys.com, forest C) hosts one lab application that Contoso users must reach. B and C must not see each other's resources.
| Trust | Direction | Type | Why |
|---|---|---|---|
| T1 | Contoso trusts Fabrikam (Contoso is the trusting, resource forest; Fabrikam is the trusted, account forest) | Forest trust, one-way | Fabrikam users reach Contoso resources; Contoso users gain nothing in Fabrikam |
| T2 | Wingtip trusts Contoso (Wingtip is trusting; Contoso is trusted) | Forest trust, one-way | Contoso users reach the Wingtip lab app; Wingtip users gain nothing in Contoso |
| none | Fabrikam and Wingtip | n/a | Forest trusts are not extended to a third forest. B and C have no path to each other |
Check: Fabrikam can reach only Contoso resources, Contoso users can reach Wingtip resources, and Fabrikam cannot reach Wingtip because T1 and T2 are separate forest trusts. If B and C later need to share, a separate trust between them is required.
Direction in plain terms
In a one-way trust the users come from the trusted forest and the resources live in the trusting forest. Draw the arrow of access from account forest to resource forest, then check you have not given the reverse. For this design the arrows are:
Fabrikam (accounts) ---> Contoso (resources) T1: Contoso is trusting, Fabrikam is trusted
Contoso (accounts) ---> Wingtip (resources) T2: Wingtip is trusting, Contoso is trusted
Fabrikam x Wingtip no trust, so no arrow and no path
Read each arrow as "people from here can use things over there". The forest at the arrow head is the one that trusts; it is the one that decides who may touch its servers.
Terms used below. Forest trust: a trust between the root domains of two whole forests, covering every domain in each (a plain domain trust, called an external trust, covers only the two domains it joins). ACE: access control entry, one line in a permission list that names a user or group and what it may do. Domain local group: a group that can be used in permissions only inside its own domain, but can contain accounts and global groups from trusted forests. Global group: a group that holds accounts from its own domain and can be used in permissions in any trusted domain. Tier 0: the accounts and servers that control identity itself, such as domain controllers and the accounts that administer them. Allowed to authenticate: a permission on a computer object in the trusting forest that lets one named outside user or group sign in to that computer when selective authentication is on.
Prerequisites and creation
- DNS. A forest trust needs one of: a shared root DNS server for both namespaces, conditional forwarders in each forest for the other namespace, or secondary zones.
- Permissions. Domain Admins of the forest root domain or Enterprise Admins create the trust. A password is set between the sides. Enterprise Admins in both forests can create both sides at once with a generated password. The
netdom trustcommand cannot create a forest trust, so use the Active Directory Domains and Trusts snap-in. - Trust password rotation. The trusting domain's PDC emulator changes the trust password every 30 days.
Risk specific to each trust
- T1: Fabrikam's administrators and any compromise of Fabrikam accounts become a route into Contoso. Contain it with selective authentication limited to
app01, SID history off and no administrative use of Fabrikam accounts on Contoso hosts. - T2: a compromised Contoso account can reach the Wingtip lab application. Contain it with selective authentication on Wingtip's outbound trust, access granted only to one Contoso group, and no Contoso admin accounts used there.
Risks every trust introduces and how to contain them
| Risk | Containment |
|---|---|
| Accounts from the trusted forest are treated as authenticated in the trusting forest, so a broadly granted "Authenticated Users" ACE includes them | Turn on selective authentication on the outgoing side, then grant "Allowed to authenticate" on only the specific computer objects, then grant resource permissions to groups |
| SID history (old identifiers carried over when an account was migrated between domains) can grant access the trusting forest never intended | Keep SID history disabled with netdom trust ... /enablesidhistory:No unless you trust the other forest's administrators |
| A privileged trusted-forest account used on trusting-forest hosts | Do not use admin accounts across trusts; keep Tier 0 separate |
| Kerberos delegation across the trust (a service acting as the user on other servers) | Disable full delegation on outbound trusts with /enabletgtdelegation:No |
| Stale or unmonitored trusts | Verify with netdom trust ... /verify; review annually; remove unused trusts |
| Compromise of the trusted forest | Treat the trusted forest as part of the trusting forest's attack surface; the controls above limit blast radius, they do not remove it |
Selective authentication applies on outbound forest and external trusts; the forest-wide alternative gives outside users the same level of access as local users, which is exactly what you do not want here.
netdom trust contoso.com /domain:fabrikam.com /SelectiveAUTH:Yes /userD:FABRIKAM\admin /passwordD:*
netdom trust contoso.com /domain:fabrikam.com /enablesidhistory:No /userD:FABRIKAM\admin /passwordD:*
netdom trust contoso.com /domain:fabrikam.com /enabletgtdelegation:No /userD:FABRIKAM\admin /passwordD:*
netdom trust contoso.com /domain:fabrikam.com /verify /userD:FABRIKAM\admin /passwordD:*
Repeat on the wingtiptoys.com side for T2 with the arguments reversed.
What each line does. The first name after netdom trust is the trusting forest, /domain: is the trusted one, and /userD with /passwordD:* is an administrator of the trusted domain whose password is prompted for instead of typed on the line.
/SelectiveAUTH:Yesturns on selective authentication: outside users can sign in only to computers where they hold Allowed to authenticate./enablesidhistory:Nomakes the trusting forest ignore SID history in tickets from the other forest./enabletgtdelegation:Noblocks Kerberos full delegation across the trust, so a service in the other forest cannot receive a forwardable copy of a user's ticket-granting ticket./verifychecks that the trust password (the shared secret) matches on both sides.
The Microsoft reference says /SelectiveAUTH and /enablesidhistory are valid only on outbound forest trusts (and external trusts for selective authentication), which is why they run on the trusting side.
Illustrative shape of a successful /verify (not captured from a live run; the exact wording varies by Windows version):
The trust between contoso.com and fabrikam.com has been successfully verified.
The command completed successfully.
Read it as: the secret on both sides agrees and each side can reach a domain controller of the other. A failure prints an error instead, most often a name resolution, firewall or RPC problem, which is what to check after a network change.
Worked example
A Fabrikam finance user needs the Contoso application on server app01. Fabrikam adds the user to FAB-Finance-Global. In Contoso, app01 has "Allowed to authenticate" granted to FAB-Finance-Global, a Contoso domain local group APP-Users contains FAB-Finance-Global, and the application ACL grants APP-Users. The same user cannot reach fileserver01 in Contoso: the account is trusted, but not allowed to authenticate to that computer. Here FAB-Finance-Global is a global group (it holds Fabrikam accounts), APP-Users is a domain local group (it is used in the application's ACE), and putting the first inside the second is the usual pattern across a trust.
Trade-offs
- Selective authentication adds administrative work for every new server.
- Two separate trusts are easier to reason about than a hub that implies reach.
Create a disaster recovery plan for AD in a hybrid cloud estate. How do you handle an on-premises DC failure and what happens to sync and sign-in during and after it?
Sample Answer
Direct answer
Plan the hybrid directory as three separate things that fail separately: the on-premises domain controllers (DCs), the Microsoft Entra Connect sync server (the server that copies identities from Active Directory to Microsoft Entra ID, Microsoft's cloud identity directory, formerly called Azure AD), and the sign-in method users depend on. A DC failure is handled by having other DCs and seizing roles (forcing a role onto a surviving DC when the old holder is gone for good) if needed. A sync-server failure is handled by a warm staging server. A forest-wide loss is handled by forest recovery, with sync switched off until the restored directory has been inspected. Targets: recovery time objective (RTO, how long recovery may take) and recovery point objective (RPO, how much data may be lost) differ by tier, and I propose them below for the owner to confirm.
Proposed targets by failure tier
These are targets to agree with the business, not measured figures.
| Failure | What users notice | Mechanism | Target |
|---|---|---|---|
| One DC lost, others in the site | Nothing, if clients locate another DC | At least two writable DCs per site; seize any FSMO roles the lost DC held; metadata cleanup | RTO: no outage. RPO: only changes not yet replicated off that DC |
| Sync server lost | Cloud sign-in continues; new users, attribute changes and password changes stop reaching the cloud | Staging-mode standby, or rebuild | RTO: 1 hour for exports to resume with a warm standby (password catch-up after the switch can take longer, see the failover section); Microsoft says a rebuild usually completes within a few hours |
| All on-premises DCs lost | Depends on sign-in method (see the sign-in table) | Forest recovery, as described in the forest-wide loss section | RTO per drill results |
FSMO means Flexible Single Master Operations, the five single-holder DC roles.
On-premises DC failure: ordered steps
- Confirm clients are authenticating against other DCs:
Get-ADDomainController -Discover -Service PrimaryDCand-Service GlobalCatalogshow which DC the locator hands out. - Find out whether the dead DC held roles:
Get-ADDomain | Format-List InfrastructureMaster, RIDMaster, PDCEmulatorandGet-ADForest | Format-List DomainNamingMaster, SchemaMaster. - If it did and it will not return, seize the roles to a healthy DC (the documented seize procedure), then run metadata cleanup, which removes the dead DC's leftover records from the directory so partners stop trying to replicate with it.
- Check replication:
Get-ADReplicationFailure -Target <domain> -Scope Domain. - On the sync server, confirm the AD connector still completes imports (the run history in Synchronization Service Manager, the sync server's own console) and run
Start-ADSyncSyncCycle -PolicyType Delta. - Add the replacement DC and verify it.
What happens to sync and sign-in
| Sign-in method | Depends on on-premises during an outage? | Notes |
|---|---|---|
| Password hash synchronization (PHS) | No for sign-in itself. User authentication happens in Microsoft Entra ID against the synchronized hash | New password changes made on-premises do not reach the cloud until sync runs, so users may need the old password in the cloud |
| Pass-through authentication (PTA) | Yes. Agents (small services installed on-premises that check each password against a DC) validate every sign-in | Microsoft recommends at least 3 agents per production tenant, with a system limit of 40, installed close to the DCs. All agents lose function if no DC is reachable |
| Federation (AD FS, Active Directory Federation Services) | Yes, on the federation servers, which in turn need DCs | PHS can be enabled as a fallback if the federation service has an outage |
The three methods differ in one respect worth reading first: whether sign-in still works when the datacenter is dark. With PHS it does, because Microsoft Entra ID holds a hash of the password. With PTA and AD FS it does not, because each sign-in is checked on-premises. Two concrete users show the difference (names illustrative):
- Dana is on PHS. The DCs are down, and she signs in to a cloud app with her usual password: it works. Had she changed her password on-premises less than two minutes before the sync server died (password hash sync runs every 2 minutes, so anything older had already reached the cloud), the cloud still holds the old hash, so she must keep using the old password in the cloud until the sync server is back and catches up.
- Lee is on PTA. The DCs are down, and the agents have nothing to ask: Lee cannot sign in until a DC is reachable again.
PHS runs every 2 minutes; the object and attribute sync cycle runs every 30 minutes by default. Keep one or two cloud-only emergency administrator accounts so you can still manage the tenant if the on-premises side is down.
Sync server loss and failover
Entra Connect supports active-passive only: exactly one server may be exporting. A server in staging mode imports and synchronizes but does not export, and does not run password sync or password writeback (sending cloud password changes back to Active Directory), even if those features were selected.
Failover order:
- Make sure the old active server cannot export. If it is reachable, switch it to staging mode in the wizard; if not, shut it down or isolate it so it cannot return unexpectedly.
- On the standby, confirm it has synchronized recently and run a sync cycle:
Get-ADSyncSchedulershould showStagingModeEnabledas True before the switch. - Check the pending exports for surprises. In the
binfolder under the Connect install,csexport "<connector name>" %temp%\export.xml /f:xwrites the changes the connector is about to export into an XML file, andCSExportAnalyzer %temp%\export.xml > %temp%\export.csvturns it into a spreadsheet, one row per pending change. - Untick staging mode in the wizard and let it start the sync process.
- Confirm export run profiles are running in the Synchronization Service.
How to read what these steps show. On the standby, an excerpt of Get-ADSyncScheduler looks like this before the switch (property names are the cmdlet's own; layout illustrative):
SyncCycleEnabled : True
MaintenanceEnabled : True
StagingModeEnabled : True
NextSyncCyclePolicyType : Delta
StagingModeEnabled : True means this server imports and synchronizes but suppresses export, which is exactly the state you want before you review the pending changes; after step 4 it reads False. In export.csv, the OMODT column is the object-level change (Add, Update or Delete) and AMODT is the attribute-level change. A handful of Update rows for users you expect is normal. A column of Delete rows for an organizational unit you did not mean to remove is the stop sign: do not untick staging mode. The same list is available in Synchronization Service, under Connectors: select the Microsoft Entra ID connector, Search Connector Space, scope Pending Export, tick Delete.
Consequences to plan for: after staging is turned off, password sync resumes from its last watermark (its bookmark of the last password change it processed), which can mean hours of catch-up in a large estate, and newly changed passwords do not work in the cloud until the backlog clears. Do not restart the sync services during catch-up. Periodically promote the standby to active temporarily so the backlog stays small. Only one active server can use password writeback at a time.
The sync server holds no unique data: it can be rebuilt from Active Directory and Entra ID because the sourceAnchor attribute (a stable ID stored on each synced object, used to match the on-premises object to its cloud twin) re-matches existing objects. What must be saved is the configuration: custom sync rules, filtering, and a current export of the server configuration. Also, a delta sync must happen at least once every 7 days (this applies to staging servers too), or a full synchronization will be required. If the sync database is on SQL Server rather than the bundled SQL Express, SQL Always On availability groups and clustering are supported, and mirroring is not.
Forest-wide loss: what changes in Entra after restore
A restored directory goes back to its backup time. Without precautions, Entra Connect would push that old state to the cloud:
- Objects created after the backup would be treated as gone from the directory and queued as deletes in Entra ID. The default accidental-delete threshold of 500 stops an export that contains more than 500 deletes (it halts before deleting anything and logs warning event 116), which is a safety net, not a plan.
- A restored user whose password changed after the backup has the old hash on-premises, and a password sync of that user would overwrite the cloud password with it (Microsoft's password hash sync article says a synchronized password overwrites the existing cloud password; I did not find documentation of exactly when a restored directory triggers that sync, so treat this as a risk to check rather than a certainty). For example, backup at 02:00, Dana changes her password at 10:00 (the cloud gets the new hash), the forest is restored to the 02:00 state at 14:00: the on-premises copy now has the old password, and if an unprotected sync pushed it over the cloud copy, her working password would silently stop working.
So after a restore, keep the sync server in staging mode (or disable the scheduler with Set-ADSyncScheduler -SyncCycleEnabled $false) until the directory is validated, run the imports, inspect the pending exports for deletes and unexpected changes, then re-enable exports. The restore itself follows the forest recovery guide: isolate, restore one writable DC per domain, seize roles, clean up metadata, reset the krbtgt password twice (at least 10 hours apart by default), then redeploy the remaining DCs.
Testing
- Quarterly: fail the sync role over to the standby and back, timing the export resuming and the password sync catch-up.
- Annually at least: a forest recovery drill in an isolated network, including the step where the sync server is held back.
- Monthly: confirm PTA agent count and health if PTA is in use, and that PHS is enabled as the fallback if federation is in use.
Pitfalls
- Believing "we have PHS" means the on-premises side no longer matters for password changes and new users.
- Two Connect servers exporting at once.
- Letting sync run against a freshly restored AD before checking pending deletes.
Design Active Directory for a company with 50,000 users across six regions that needs 24x7 authentication and low logon latency. Where do you draw forest and domain boundaries, where do domain controllers go and of what kind, and what would you give up?
Sample Answer
Direct answer
Build one forest with one domain. The six regions become six sets of AD sites, not six domains. Put two to three writable domain controllers (DCs) in each regional hub site, make every DC a global catalog, add read-only DCs (RODCs) in branches that lack physical security or bandwidth, and place the PDC emulator in the best-connected hub. Use fine-grained password policies (password and lockout rules that apply to chosen groups instead of the domain's single default rule) and OU (organizational unit) delegation (giving a regional team rights over just their OU) for regional differences. What you give up is isolation: everything shares one security boundary, one schema and one set of replicas, and wrong-password and lockout handling for regional users goes through a PDC emulator that may be far away.
Where the boundaries go
- Domain. The administrative and replication boundary: all DCs in a domain hold a complete copy of that domain's data and share its security policies.
- Forest. The security boundary. An administrator in one domain cannot be stopped from reaching data in another domain of the same forest, so extra domains add no security.
- Sites. Map subnets to locations so clients find a nearby DC and replication is scheduled over the real links. This, not domains, delivers low logon latency.
Domain count: six domains would multiply operations-master roles (a single forest has one schema master and one domain naming master; each domain has its own RID (relative identifier) master, PDC emulator and infrastructure master). One domain: 5 roles. Six domains: 2 + 3 x 6 = 20 roles. It would also bring multi-domain global catalog planning and universal-group logon dependencies, without any security gain. Regional differences are handled inside one domain: fine-grained password policies give different password and lockout rules to groups of users, and OU delegation gives regional admins rights over their OU only.
Split off a separate domain or forest only for a stated reason: a legal entity or acquisition that must stay isolated, a regulation requiring that one region's administrators cannot affect another's (that needs a separate forest, because a domain boundary does not stop a malicious domain administrator), or a pristine administrative forest for high-value assets. Data residency rules are a special case: every DC in a domain, including an RODC, holds all of that domain's objects (an RODC withholds passwords and any attributes in its filtered attribute set, but still holds every object), and a global catalog holds a partial copy of every object in the forest. If a law forbids a region's user records on foreign DCs, only a separate forest fully answers it. Check the regulation's actual wording before paying that cost.
Domain controller layout
| Component | Decision |
|---|---|
| Regional hub site (6) | 2 to 3 writable DCs, all global catalogs |
| Branch site | RODC if physical security or bandwidth is poor; otherwise users authenticate over the WAN to the hub |
| Global catalog | Every DC. In a single-domain forest this costs no extra storage, CPU or replication traffic. (If domains were ever added, Microsoft's guidance is a global catalog at locations with more than 100 users, and universal group membership caching for smaller ones) |
| PDC emulator | The best-connected hub. The forest root domain's PDC emulator is the authority for time and should take it from an external source; with one domain, that is this DC |
| Infrastructure master | Irrelevant when every DC is a global catalog |
| Functional level | A functional level is a switch that turns on newer AD features and sets the oldest Windows Server version a DC may run. Set it to the highest the DCs support: the Windows Server 2016 level for 2016, 2019 and 2022 DCs (2019 and 2022 introduced no new level); the Windows Server 2025 level only when all DCs run Windows Server 2025 |
| Features | AD Recycle Bin enabled, system state backups, DSRM password archived |
Why two or three per hub: 24x7 means a DC can be patched or fail while logons continue. Microsoft's capacity guide recommends a peak-period CPU target between 40 and 60 percent of capacity and an N + 1 design, in which service continues at acceptable quality when one system fails; its worked site example plans for 40 percent at peak with every DC up (N + 1) and 60 percent otherwise (N, which covers the case after a failure), and it gives 80 to 100 percent at peak as the ceiling for the N case. This design uses the example's two targets: 40 percent at peak in normal operation and 60 percent after one DC fails. Call a hub's peak load L, measured in "whole DCs' worth" of work. With N DCs sharing evenly, each runs at L / N normally and at L / (N - 1) after one fails.
- Three DCs at 40 percent each: L = 3 x 0.40 = 1.2. After one fails, the two remaining DCs carry 1.2 / 2 = 60 percent each, which meets the failover target.
- Two DCs at 40 percent each: L = 2 x 0.40 = 0.8. After one fails, the single remaining DC carries 0.8 / 1 = 80 percent, above the 60 percent target. Against the guide's looser 80 to 100 percent ceiling that two-DC hub would only just fit, so the 60 percent target is the safety margin the guide's own example plans for.
The rule that follows: the remaining DCs must carry L at no more than 60 percent, so L <= 0.6 x (N - 1). Two DCs are enough when L <= 0.6 (one DC could carry the whole region at 60 percent), three when L <= 1.2, four when L <= 1.8. A hub whose measured peak is 1.5 DCs' worth needs four, because 1.5 / 2 = 75 percent fails the target with three DCs but 1.5 / 3 = 50 percent passes with four. That is why 12 to 18 writable DCs (6 x 2 to 6 x 3) is a planning range for the six hubs, to be confirmed by measuring each hub's peak, not a fixed number.
Sizing (from Microsoft's capacity planning method)
The failover math above decides how many DCs a hub gets. The three estimates below decide how big each DC is: database size drives disk and RAM, and the CPU estimate gives a starting core count.
- Database: 40 to 60 KB per user, so 50,000 users is 50,000 x 40 KB = 2.0 GB to 50,000 x 60 KB = 3.0 GB for user objects. Computers and groups add to it, so measure.
- RAM: database + base operating system (about 4 GB) + agents (about 0.2 GB LSASS, 0.1 GB monitoring, 0.2 GB antivirus) + 1 GB cushion. With a 3 GB database, 3 + 4 + 0.2 + 0.1 + 0.2 + 1 = 8.5 GB. Add 33 percent growth (Microsoft's own sizing example uses a 33 percent growth estimate over a server's 3 to 5 year hardware life): 8.5 x 1.33 = 11.3 GB, so provision 16 GB, the next standard memory size above 11.3.
- CPU: the guide's benchmark is about 1,000 concurrent users per core, and the guide itself warns that a general users-per-core figure is inapplicable to many environments because client behaviour varies so much. Average 50,000 / 6 = 8,333 users per region. Treating every user as concurrent is the cautious assumption, and it gives roughly 8.3 cores' worth at that benchmark. Treat it as a starting point and replace it with measured LSASS (the Windows security process) processor time at peak.
Trade-offs: what you give up
- One blast radius. (Blast radius means how much is damaged when one thing is compromised or breaks.) A compromised domain admin owns the forest. Mitigate with tiered administration (separate admin accounts for DCs, servers and workstations, so a credential stolen on a workstation cannot reach the DCs) and secure admin hosts; keep high-value assets in a separate pristine forest (a small, tightly locked-down forest holding only the most valuable accounts and servers) if the risk justifies it.
- PDC emulator distance. Password changes replicate preferentially to the PDC emulator, bad-password results are forwarded to it, and lockout is processed there. Users in a distant region pay that WAN latency on wrong-password attempts and lockouts. Choose its location by where most users are, and monitor this path.
- Every DC holds everything. Fine for 3 GB; a problem only for residency rules (above).
- Shared schema and policy. A schema update or a bad domain-level GPO (Group Policy object) reaches all six regions. Use OU-scoped GPOs and staged rollouts.
- Regional autonomy. Regional admins get OU delegation, not domain admin.
Failure modes to design for
- WAN down: hubs authenticate locally, RODC branches serve only cached accounts, so cache the branch users and their computers.
- Virtual DC restored from a snapshot: USN (update sequence number) rollback, where the restored DC reuses numbers its partners already saw and so silently stops replicating its newer changes. Use hypervisors with VM-GenerationID support (a counter the hypervisor changes when a VM is restored or copied, letting a Windows Server 2012 or later DC notice and recover) and AD-aware backups only.
A penetration test shows an attacker with ordinary domain credentials reached Domain Admin. Which AD attack paths are likely, and how would you harden against them and detect them?
Sample Answer
Direct answer
Starting from ordinary domain credentials, the realistic routes to Domain Admin are weak points that let one low-privilege identity borrow a privileged one: an over-privileged service account that can be Kerberoasted, accounts that skip Kerberos pre-authentication, unconstrained delegation, replication rights that allow DCSync (abusing directory replication to pull password data), permissive access control lists (ACLs) on privileged objects and groups, and admin credentials left on lower-trust machines. Harden by shrinking and isolating privilege, moving service accounts to group Managed Service Accounts (gMSAs) with AES, removing stray replication rights, and deploying Windows LAPS; detect by auditing service-ticket encryption types, directory-service access to the domain object, and changes to protected objects.
Attack paths, fixes and detection
Terms used in the table, in plain words:
- Kerberos ticket and TGT. After sign-in, Kerberos gives the user a ticket-granting ticket (TGT), a signed proof of identity, which they present to get a separate ticket for each service. The
krbtgtaccount is the account whose key signs every TGT, so its secret lets an attacker forge any identity. - Pre-authentication. The step where a user proves they know their password before the DC hands over anything encrypted with it.
- RC4 and AES. Two encryption algorithms Kerberos can use. RC4 is the older and weaker one and is faster to crack, so the goal is to see AES in tickets, not RC4.
- Protected Users. A built-in security group whose members get stricter rules: no NTLM, no weak encryption, no delegation, short-lived TGTs.
- "Sensitive and cannot be delegated". An account flag that stops any service from reusing that account's identity on its behalf.
- Pass-the-hash. Using a stolen password hash directly as proof of identity, without cracking it.
- Tiering (Tier 0). Splitting admin work by what is controlled: Tier 0 is the DCs and anything that can control them, and Tier 0 accounts never sign in to lower tiers.
Order of work: the plan below starts with the path the pen test actually used. It then shrinks privilege, because that limits what every other path can reach, and fixes Kerberoastable service accounts, which need only an ordinary domain account and so are reachable by every employee. AS-REP roasting is just as cheap for an attacker and is fixed by clearing one flag on each account the audit finds. Delegation abuse, replication rights and ACL abuse need a foothold on a specific host or a mistaken permission, so they are found by audit and fixed after the paths that every employee can reach.
| Path | How it works | Harden | Detect |
|---|---|---|---|
| Kerberoasting | Any user can request a service ticket for an account with a service principal name (SPN, the name that identifies a service to Kerberos). Part of the ticket is encrypted with that account's password-derived key, so a weak password can be cracked offline | Move to gMSA, remove SPNs from privileged users, long random passwords, configure AES | Event 4769 with ticket encryption type 0x17 (RC4) for user-account services; Microsoft expects 0x11 or 0x12 |
| AS-REP roasting | An account flagged "does not require Kerberos pre-authentication" answers a ticket-granting-ticket request without proof of the password, so the reply can be cracked offline | Clear the flag; audit for it | Authentication events for those accounts from unusual clients |
| Unconstrained delegation | A service trusted for delegation can impersonate clients that connect to it, so a privileged account that touches that host exposes its identity | Remove it; use constrained or resource-based delegation; mark admins "sensitive and cannot be delegated" or put them in Protected Users | Monitor changes to delegation settings on accounts |
| DCSync | An account holding Replicating Directory Changes and Replicating Directory Changes All on the domain object can ask a DC to replicate password data, including the krbtgt secret | Remove those rights from everything that is not a DC or a deliberate sync account | Event 4662 on the domain object whose properties carry the right GUIDs, from a non-DC account |
| ACL abuse and AdminSDHolder | Write or modify-permission rights on a privileged group, a protected account or the AdminSDHolder object (the template whose ACL is re-applied to protected accounts and groups every 60 minutes by default, so a permission an attacker adds to a protected group is reset, but one added to the AdminSDHolder object itself is copied to all of them) | Review ACLs; restrict who can edit them | Event 4662 with Write Property or WRITE_DAC on those objects |
| Credential theft on lower tiers | Admins sign in to ordinary workstations or servers; local administrator passwords are shared | Secure admin hosts, tiering, Windows LAPS, Protected Users | 4624/4769 for admin accounts on machines they should never use |
| Editable policy objects | Rights to edit a Group Policy Object (GPO) linked to the domain root or the Domain Controllers OU let an attacker push settings to every machine under it | Restrict edit rights on those GPOs | Changes to those GPOs |
Ordered plan
- Reproduce the tester's exact chain and fix that link first. Every hardened path matters less than the one already proven.
- Shrink privilege. Remove permanent membership in Domain Admins, Enterprise Admins and Administrators; grant temporary membership when needed; keep admin accounts off ordinary hosts.
- Fix service accounts. Create a Key Distribution Service (KDS) root key (the forest-wide secret from which DCs derive gMSA passwords; without it no gMSA can be created) once per forest with
Add-KdsRootKey -EffectiveImmediately. Even then, domain controllers wait up to 10 hours from the key's creation, so that all of them have replicated it, before a gMSA can be created and used; in a single-DC test lab only,Add-KdsRootKey -EffectiveTime ((Get-Date).AddHours(-10))skips the wait. Then create gMSAs withNew-ADServiceAccount -Name ... -DNSHostName ... -PrincipalsAllowedToRetrieveManagedPassword ... -KerberosEncryptionType AES256. Set the AES (Advanced Encryption Standard) encryption type explicitly:-KerberosEncryptionTypesets the account'smsDS-SupportedEncryptionTypesattribute, and when that attribute is not configured the domain controllers fall back to the domain's default supported encryption types, which are not necessarily AES-only. Use the audit queries in the worked example to find accounts that do not require Kerberos pre-authentication and clear that flag, and to find non-DC objects trusted for unconstrained delegation and move them to constrained or resource-based delegation. On Windows Server 2025, delegated managed service accounts (dMSA) are also available: they are created withNew-ADServiceAccount -CreateDelegatedServiceAccount, and Microsoft documents them for migrating services that run under normal user accounts. - Deploy Windows LAPS (Local Administrator Password Solution), which rotates each machine's local administrator password and stores it in Active Directory or Microsoft Entra ID, so one stolen local password no longer opens other machines.
- Protect the admins. Put human admin accounts in Protected Users (no NTLM, no DES or RC4 in Kerberos pre-authentication, no delegation, ticket-granting tickets limited to 4 hours at the 2012 R2 domain functional level). Never put service or computer accounts in it, and never enrol every privileged account before testing, because there is no workaround if the restrictions lock you out.
- Remove replication rights. The directory-sync connector account for Microsoft Entra Connect legitimately holds both replication rights when password hash synchronization is used (an express installation grants them for that purpose), so list it, treat that server as tier-0, and remove everything else.
- Turn on detection (below), then retest with the tester.
Worked example
svc-sql has the SPN MSSQLSvc/sql01.contoso.com:1433, belongs to Domain Admins and has a six-year-old password. The tester requests its ticket (event 4769, encryption type 0x17), cracks it offline and signs in as a domain admin. The fix is three changes: remove svc-sql from Domain Admins, replace it with a gMSA that has only the rights SQL needs, and delete the old account. Audit for the other paths with:
$domainDn = (Get-ADRootDSE).defaultNamingContext
$bitAnd = '1.2.840.113556.1.4.803'
# 1. User accounts that carry a service principal name (Kerberoastable)
Get-ADUser -LDAPFilter '(&(objectCategory=person)(objectClass=user)(servicePrincipalName=*)(!(sAMAccountName=krbtgt)))' `
-Properties servicePrincipalName, pwdLastSet, adminCount, 'msDS-SupportedEncryptionTypes' |
Select-Object SamAccountName, adminCount,
@{ n = 'PasswordSetOn'; e = { [datetime]::FromFileTime($_.pwdLastSet) } },
@{ n = 'EncTypes'; e = { $_.'msDS-SupportedEncryptionTypes' } },
@{ n = 'SPNs'; e = { $_.servicePrincipalName -join '; ' } } |
Format-Table -AutoSize
# 2. Accounts that do not require Kerberos pre-authentication (0x400000 = 4194304)
Get-ADUser -LDAPFilter "(userAccountControl:${bitAnd}:=4194304)" | Select-Object SamAccountName
# 3. Non-DC objects trusted for unconstrained delegation (0x80000 = 524288; 516 = Domain Controllers group)
Get-ADObject -LDAPFilter "(&(|(objectCategory=person)(objectCategory=computer))(userAccountControl:${bitAnd}:=524288)(!(primaryGroupID=516)))" |
Select-Object Name, ObjectClass
# 4. Who can replicate secrets out of the domain (DCSync rights on the domain head)
$getChanges = [guid]'1131f6aa-9c07-11d1-f79f-00c04fc2dcd2' # Replicating Directory Changes
$getChangesAll = [guid]'1131f6ad-9c07-11d1-f79f-00c04fc2dcd2' # Replicating Directory Changes All
(Get-Acl -Path "AD:\$domainDn").Access |
Where-Object { $_.AccessControlType -eq 'Allow' -and $_.ObjectType -in @($getChanges, $getChangesAll) } |
Select-Object IdentityReference,
@{ n = 'Right'; e = { if ($_.ObjectType -eq $getChangesAll) { 'Changes-All' } else { 'Changes' } } }
How to read the audit queries above:
userAccountControlis one number in which each bit is a switch. The long number1.2.840.113556.1.4.803in the filter is the LDAP "bitwise AND" matching rule:(userAccountControl:1.2.840.113556.1.4.803:=4194304)means "this bit is set". 4194304 is 0x400000, the "do not require Kerberos pre-authentication" switch, and 524288 is 0x80000, the "trusted for delegation" switch.primaryGroupID516 is the Domain Controllers group, so(!(primaryGroupID=516))excludes real DCs, which are legitimately trusted for delegation.- The two GUIDs are the fixed identifiers of the extended rights Replicating Directory Changes and Replicating Directory Changes All. Query 4 lists who holds them on the domain object; an account that holds both can DCSync.
What a hit looks like (illustrative output, using the svc-sql account from this example):
SamAccountName adminCount PasswordSetOn EncTypes SPNs
-------------- ---------- ------------- -------- ----
svc-sql 1 9/14/2020 8:12:03 AM MSSQLSvc/sql01.contoso.com:1433
Query 1 hit: a user account with an SPN, adminCount 1 (meaning it is or was in a protected group such as Domain Admins), a six-year-old password, and a blank EncTypes (the attribute is not set, so the account falls back to the domain's default encryption).
SamAccountName
--------------
legacy-batch
Query 2 hit: one account without pre-authentication.
Name ObjectClass
---- -----------
APP-OLD01 computer
Query 3 hit: a non-DC computer trusted for unconstrained delegation.
IdentityReference Right
----------------- -----
BUILTIN\Administrators Changes
BUILTIN\Administrators Changes-All
CONTOSO\svc-oldsync Changes
CONTOSO\svc-oldsync Changes-All
Query 4 rows: the built-in groups and your sync account are expected; svc-oldsync holding both rights is the DCSync-capable account to investigate.
Hunt in a domain controller's Security log for service tickets that are not AES:
# On a domain controller: service tickets issued with anything other than AES128 (0x11) or AES256 (0x12)
$xpath = "*[System[EventID=4769]] and *[EventData[Data[@Name='Status']='0x0' and Data[@Name='TicketEncryptionType']!='0x11' and Data[@Name='TicketEncryptionType']!='0x12']]"
Get-WinEvent -LogName Security -FilterXPath $xpath -MaxEvents 2000 |
ForEach-Object {
$d = @{}
([xml]$_.ToXml()).Event.EventData.Data | ForEach-Object { $d[$_.Name] = $_.'#text' }
[pscustomobject]@{ Account = $d.TargetUserName; Service = $d.ServiceName; EType = $d.TicketEncryptionType; Client = $d.IpAddress }
} |
Group-Object Account, Service, EType |
Sort-Object Count -Descending |
Select-Object -First 20 Count, Name
What the event looks like (illustrative output of the hunt query):
Count Name
----- ----
37 alice@CONTOSO.COM, svc-sql, 0x17
4 bob@CONTOSO.COM, svc-web, 0x17
Each row is one user (TargetUserName), one service account (ServiceName) and one encryption type, with how many tickets matched. In event 4769 (a service ticket was requested), TicketEncryptionType 0x12 is AES256 and 0x11 is AES128, so a normal log for a modern service is all 0x12. 0x17 is RC4. One user asking for one RC4 ticket for an old service is routine; 37 RC4 requests from one user in a short window for an account that is in Domain Admins is what Kerberoasting looks like. Event 4624 is a successful sign-in; for an admin account the question is whether the machine named in it is one that admin should ever use. Event 4662 means an operation was performed on a directory object, which is written only when the SACL described under Trade-offs and pitfalls is in place.
Trade-offs and pitfalls
- RC4 cannot simply be switched off everywhere on day one, because old services and accounts without AES keys will fail; start by finding who still requests it.
- gMSA is not available to every legacy application, so keep a long-random-password and rotation process for those.
- DCSync detection needs a system access control list (SACL, the part of an object's security descriptor that says which accesses to log) entry and the auditing category in place first, otherwise event 4662 is never written.
- A sensible-looking Protected Users rollout can lock out the admins themselves if their accounts lack AES keys. Change their password first.
Unlock Full Question Bank
Get access to all 18 Active Directory Architecture and Management interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.