Active Directory Architecture and Management Questions
Designing, operating and recovering Active Directory Domain Services and the directory estate around it. Covers logical and physical structure (forests, trees, domains, OUs, trusts, schema extension, FSMO roles, Global Catalog, RODCs), sites and replication topology, domain controller placement, promotion, upgrade and functional levels, DC locator and AD-integrated DNS, domain join, Kerberos and NTLM authentication including SPNs and delegation, token size and SID history, LDAP binds and query tuning, Group Policy design, processing order, filtering, deployment and troubleshooting, user, group, computer and service account management including PowerShell account scripting and account lockout investigation, delegation of control and tiered administration, fine-grained password policy, backup, authoritative restore, forest recovery and USN rollback, AD hardening against Kerberoasting, DCSync and Golden Ticket attacks, forest migration, and hybrid identity with Microsoft Entra ID (Microsoft Entra Connect, Cloud Sync, password hash sync, pass-through authentication, federation, password writeback). Questions are asked from the directory administrator's and architect's seat. Platform-neutral identity protocols and lifecycle design, Windows file-server administration (shares, NTFS permissions, profiles), Linux directory integration, and generic DNS and DHCP service operations are covered elsewhere.
Create a disaster recovery plan for AD in a hybrid cloud estate. How do you handle an on-premises DC failure and what happens to sync and sign-in during and after it?
Sample Answer
Direct answer
Plan the hybrid directory as three separate things that fail separately: the on-premises domain controllers (DCs), the Microsoft Entra Connect sync server (the server that copies identities from Active Directory to Microsoft Entra ID, Microsoft's cloud identity directory, formerly called Azure AD), and the sign-in method users depend on. A DC failure is handled by having other DCs and seizing roles (forcing a role onto a surviving DC when the old holder is gone for good) if needed. A sync-server failure is handled by a warm staging server. A forest-wide loss is handled by forest recovery, with sync switched off until the restored directory has been inspected. Targets: recovery time objective (RTO, how long recovery may take) and recovery point objective (RPO, how much data may be lost) differ by tier, and I propose them below for the owner to confirm.
Proposed targets by failure tier
These are targets to agree with the business, not measured figures.
| Failure | What users notice | Mechanism | Target |
|---|---|---|---|
| One DC lost, others in the site | Nothing, if clients locate another DC | At least two writable DCs per site; seize any FSMO roles the lost DC held; metadata cleanup | RTO: no outage. RPO: only changes not yet replicated off that DC |
| Sync server lost | Cloud sign-in continues; new users, attribute changes and password changes stop reaching the cloud | Staging-mode standby, or rebuild | RTO: 1 hour for exports to resume with a warm standby (password catch-up after the switch can take longer, see the failover section); Microsoft says a rebuild usually completes within a few hours |
| All on-premises DCs lost | Depends on sign-in method (see the sign-in table) | Forest recovery, as described in the forest-wide loss section | RTO per drill results |
FSMO means Flexible Single Master Operations, the five single-holder DC roles.
On-premises DC failure: ordered steps
- Confirm clients are authenticating against other DCs:
Get-ADDomainController -Discover -Service PrimaryDCand-Service GlobalCatalogshow which DC the locator hands out. - Find out whether the dead DC held roles:
Get-ADDomain | Format-List InfrastructureMaster, RIDMaster, PDCEmulatorandGet-ADForest | Format-List DomainNamingMaster, SchemaMaster. - If it did and it will not return, seize the roles to a healthy DC (the documented seize procedure), then run metadata cleanup, which removes the dead DC's leftover records from the directory so partners stop trying to replicate with it.
- Check replication:
Get-ADReplicationFailure -Target <domain> -Scope Domain. - On the sync server, confirm the AD connector still completes imports (the run history in Synchronization Service Manager, the sync server's own console) and run
Start-ADSyncSyncCycle -PolicyType Delta. - Add the replacement DC and verify it.
What happens to sync and sign-in
| Sign-in method | Depends on on-premises during an outage? | Notes |
|---|---|---|
| Password hash synchronization (PHS) | No for sign-in itself. User authentication happens in Microsoft Entra ID against the synchronized hash | New password changes made on-premises do not reach the cloud until sync runs, so users may need the old password in the cloud |
| Pass-through authentication (PTA) | Yes. Agents (small services installed on-premises that check each password against a DC) validate every sign-in | Microsoft recommends at least 3 agents per production tenant, with a system limit of 40, installed close to the DCs. All agents lose function if no DC is reachable |
| Federation (AD FS, Active Directory Federation Services) | Yes, on the federation servers, which in turn need DCs | PHS can be enabled as a fallback if the federation service has an outage |
The three methods differ in one respect worth reading first: whether sign-in still works when the datacenter is dark. With PHS it does, because Microsoft Entra ID holds a hash of the password. With PTA and AD FS it does not, because each sign-in is checked on-premises. Two concrete users show the difference (names illustrative):
- Dana is on PHS. The DCs are down, and she signs in to a cloud app with her usual password: it works. Had she changed her password on-premises less than two minutes before the sync server died (password hash sync runs every 2 minutes, so anything older had already reached the cloud), the cloud still holds the old hash, so she must keep using the old password in the cloud until the sync server is back and catches up.
- Lee is on PTA. The DCs are down, and the agents have nothing to ask: Lee cannot sign in until a DC is reachable again.
PHS runs every 2 minutes; the object and attribute sync cycle runs every 30 minutes by default. Keep one or two cloud-only emergency administrator accounts so you can still manage the tenant if the on-premises side is down.
Sync server loss and failover
Entra Connect supports active-passive only: exactly one server may be exporting. A server in staging mode imports and synchronizes but does not export, and does not run password sync or password writeback (sending cloud password changes back to Active Directory), even if those features were selected.
Failover order:
- Make sure the old active server cannot export. If it is reachable, switch it to staging mode in the wizard; if not, shut it down or isolate it so it cannot return unexpectedly.
- On the standby, confirm it has synchronized recently and run a sync cycle:
Get-ADSyncSchedulershould showStagingModeEnabledas True before the switch. - Check the pending exports for surprises. In the
binfolder under the Connect install,csexport "<connector name>" %temp%\export.xml /f:xwrites the changes the connector is about to export into an XML file, andCSExportAnalyzer %temp%\export.xml > %temp%\export.csvturns it into a spreadsheet, one row per pending change. - Untick staging mode in the wizard and let it start the sync process.
- Confirm export run profiles are running in the Synchronization Service.
How to read what these steps show. On the standby, an excerpt of Get-ADSyncScheduler looks like this before the switch (property names are the cmdlet's own; layout illustrative):
SyncCycleEnabled : True
MaintenanceEnabled : True
StagingModeEnabled : True
NextSyncCyclePolicyType : Delta
StagingModeEnabled : True means this server imports and synchronizes but suppresses export, which is exactly the state you want before you review the pending changes; after step 4 it reads False. In export.csv, the OMODT column is the object-level change (Add, Update or Delete) and AMODT is the attribute-level change. A handful of Update rows for users you expect is normal. A column of Delete rows for an organizational unit you did not mean to remove is the stop sign: do not untick staging mode. The same list is available in Synchronization Service, under Connectors: select the Microsoft Entra ID connector, Search Connector Space, scope Pending Export, tick Delete.
Consequences to plan for: after staging is turned off, password sync resumes from its last watermark (its bookmark of the last password change it processed), which can mean hours of catch-up in a large estate, and newly changed passwords do not work in the cloud until the backlog clears. Do not restart the sync services during catch-up. Periodically promote the standby to active temporarily so the backlog stays small. Only one active server can use password writeback at a time.
The sync server holds no unique data: it can be rebuilt from Active Directory and Entra ID because the sourceAnchor attribute (a stable ID stored on each synced object, used to match the on-premises object to its cloud twin) re-matches existing objects. What must be saved is the configuration: custom sync rules, filtering, and a current export of the server configuration. Also, a delta sync must happen at least once every 7 days (this applies to staging servers too), or a full synchronization will be required. If the sync database is on SQL Server rather than the bundled SQL Express, SQL Always On availability groups and clustering are supported, and mirroring is not.
Forest-wide loss: what changes in Entra after restore
A restored directory goes back to its backup time. Without precautions, Entra Connect would push that old state to the cloud:
- Objects created after the backup would be treated as gone from the directory and queued as deletes in Entra ID. The default accidental-delete threshold of 500 stops an export that contains more than 500 deletes (it halts before deleting anything and logs warning event 116), which is a safety net, not a plan.
- A restored user whose password changed after the backup has the old hash on-premises, and a password sync of that user would overwrite the cloud password with it (Microsoft's password hash sync article says a synchronized password overwrites the existing cloud password; I did not find documentation of exactly when a restored directory triggers that sync, so treat this as a risk to check rather than a certainty). For example, backup at 02:00, Dana changes her password at 10:00 (the cloud gets the new hash), the forest is restored to the 02:00 state at 14:00: the on-premises copy now has the old password, and if an unprotected sync pushed it over the cloud copy, her working password would silently stop working.
So after a restore, keep the sync server in staging mode (or disable the scheduler with Set-ADSyncScheduler -SyncCycleEnabled $false) until the directory is validated, run the imports, inspect the pending exports for deletes and unexpected changes, then re-enable exports. The restore itself follows the forest recovery guide: isolate, restore one writable DC per domain, seize roles, clean up metadata, reset the krbtgt password twice (at least 10 hours apart by default), then redeploy the remaining DCs.
Testing
- Quarterly: fail the sync role over to the standby and back, timing the export resuming and the password sync catch-up.
- Annually at least: a forest recovery drill in an isolated network, including the step where the sync server is held back.
- Monthly: confirm PTA agent count and health if PTA is in use, and that PHS is enabled as the fallback if federation is in use.
Pitfalls
- Believing "we have PHS" means the on-premises side no longer matters for password changes and new users.
- Two Connect servers exporting at once.
- Letting sync run against a freshly restored AD before checking pending deletes.
Design a backup and disaster recovery plan for a five-domain forest with a one-hour RPO and an eight-hour RTO. Say how you would test it.
Sample Answer
Direct answer
Use layers, because one mechanism cannot meet both targets. Replication across several domain controllers (DCs) per domain covers losing a DC or a site, with a recovery point (RPO) of roughly the last replication. Hourly system state backups of one designated writable DC per domain, plus nightly backups of a second one, cover logical corruption and forest-wide failure (a system state backup is the set that includes the AD database, SYSVOL and the registry), with the one-hour RPO applying only if the failure time is known. A written, practiced forest recovery plan covers the eight-hour recovery time (RTO). RPO is how much data you can lose; RTO is how long recovery may take.
Design
Layer 1: replication. Two or more writable DCs per domain in each hub, so an ordinary DC failure costs no data beyond changes not yet replicated.
Layer 2: backups.
| What | Frequency | Count per day (5 domains) | Notes |
|---|---|---|---|
| System state backup of the designated restore DC in each domain | Hourly | 5 x 24 = 120 | A dedicated restore DC per domain (one DC chosen in advance as the one you would restore) makes recovery repeatable and scriptable. |
| System state backup of a second writable DC per domain (include the PDC emulator, the role whose SYSVOL copy is usually best) | Nightly | 5 | At least two writable DCs per domain should be backed up so there are choices. |
| Full server backup of the restore DC | Weekly | 5 per week | Needed to rebuild onto different hardware: a system state backup cannot be restored onto a fresh install. |
ntdsutil snapshot of the database | Hourly, kept 24 hours | 120 | ntdsutil snapshot makes a point-in-time copy of the AD database volume; dsamain mounts such a copy as a read-only directory you can query, to find the last safe state without restarting a DC in DSRM (Directory Services Restore Mode, the offline repair boot mode). |
That is 125 system state backup runs a day, plus the snapshots and the weekly full backup. System state is taken with wbadmin start systemstatebackup -backuptarget:<drive>: or a backup product that uses the AD-aware Volume Shadow Copy Service writer (the AD component that makes the database files consistent while they are copied). Layers 1 to 3 are what the targets require; the snapshot row, DC cloning and immutable offsite copies shorten recovery or protect the backups. Measure the load of hourly runs in the drill before committing to it.
Offsite and local copies. Keep a local copy for speed and an offsite or immutable copy for survival. Retention for any backup must stay inside the tombstone lifetime (the period deleted-object records are kept), or the deleted-object lifetime if the AD Recycle Bin is on, whichever is smaller. A backup older than that cannot safely be restored. Also store, offline and with the backup: Domain Admin and DSRM passwords and BitLocker recovery keys.
Layer 3: recovery plan. A topology table of every DC (name, OS, FSMO roles, global catalog (GC) status, RODC status, backup, DNS, virtual or not; FSMO means Flexible Single Master Operations, the five single-holder roles), a chosen restore DC per domain, and a recovery order: forest root domain first, one DC per domain restored in isolation, cleaned, reconnected; the remaining DCs then redeployed (virtual DC cloning speeds this up). Root first because it holds the forest-wide roles (schema master, domain naming master) and the other domains find it for DNS and trusts. In isolation means the restored DC is kept off the network of the other DCs, so it cannot replicate with a DC that still holds the corruption or the attacker's changes, and other DCs cannot overwrite it. Cleaned means removing the directory records of the DCs you are not restoring (metadata cleanup, because those servers will be rebuilt later) and taking over any FSMO role it lacks (seizing a role means claiming it without the old holder's cooperation, unlike a transfer, because the old holder is gone). Reconnected means bringing it back onto the network once those steps are done. RODC (read-only domain controller) backups cannot be used to restore a writable DC.
What the RPO really means
One hour is achievable when a DC, a site, or a whole domain is lost but some DC still holds recent data, and when a bad change is noticed fast and the last hourly backup is clean. It is not achievable when the failure is a compromise or corruption of unknown start time. Microsoft's guidance is to restore from a backup taken a few days before the failure, trading recent data for safety from reintroducing the bad data. So the answer to the business is two numbers: RPO of about an hour for known-time accidents, and RPO equal to the chosen safe restore point for malicious or unknown-time events. Say that explicitly rather than promising one hour for everything.
RTO budget
These are planning allowances to be replaced by timings measured in the drill, not measurements.
| Step | What happens | Hours |
|---|---|---|
| Decide, declare, isolate all writable DCs | Incident lead decides it is a forest recovery; every writable DC is taken off the network so nothing replicates | 1.0 |
| Restore first DC in the root domain, clean it | Restore the system state onto the root domain's restore DC, delete the metadata of the other DCs, seize the roles, reset the krbtgt password the first time; if the restored DC was a global catalog, clear that flag (Microsoft's guide does this for a restored global catalog in a multi-domain forest) | 2.0 |
| Restore one DC in each of the other 4 domains, in parallel | Same restore and cleaning for each domain, by separate people at the same time; the forest has several domains, so clear the global catalog flag on each restored DC (Microsoft's forest recovery guide warns that a restored global catalog from a more recent backup than another domain's can introduce lingering objects) | 2.0 |
| Validation (replication, DNS, trusts, global catalog, test logons) | Reconnect the restored DCs on an isolated common network, check DNS, trusts, replication and a test logon in each domain, then add the global catalog to a root-domain DC (Microsoft's guide adds it only after the restored DCs are reconnected and replication is checked; user logons need a global catalog) | 1.5 |
| Redeploy the minimum extra DCs for load | Promote or clone more DCs so that logon capacity is enough | 1.0 |
| Total | 7.5, leaving 0.5 hour of margin |
One step cannot fit in any 8-hour window: the krbtgt account (the account whose key signs all Kerberos tickets) must be reset twice, at least 10 hours apart by default, so the second reset lands roughly 10 hours after the first. The reason, from Microsoft's forest recovery guide: 10 hours is the default maximum lifetime of a user ticket and of a service ticket, and krbtgt remembers its two most recent passwords. Two resets clear the old password from that history, but if the second came sooner, clients still holding tickets signed with the old key would be rejected. Microsoft's forest recovery guide lists the krbtgt reset among the initial steps on the first restored DC of each domain, so in this plan the first reset lands at about hour 3.0 in the root domain and hour 5.0 in the other four, which puts the second resets at about hour 13.0 and hour 15.0 (deferring the first reset to the end of validation, hour 6.5, would push the second to hour 16.5). Every variant falls outside the 8-hour window. Define the RTO as "directory usable by users" and treat the second reset as scheduled post-recovery work, which is my reading of the documented 10-hour minimum.
How to test it
- Run a full forest recovery drill at least once a year, and again when membership of Enterprise Admins or Domain Admins changes, so the people who would execute it have done it.
- Restore the real offsite copies onto an isolated virtual network, not the local convenience copy.
- Time each step against the table above. Record the age of the backup used. Pass means users can log on and DNS resolves within 8 hours, and the age of the data restored is within the stated RPO for that scenario.
- Include an authoritative SYSVOL step (making the restored DC's SYSVOL the source copy that the other DCs then copy from), the password and key inventory, and a restore of a single deleted object.
The same design for a mid-size single-domain estate
A mid-size organisation with one domain and a handful of DCs uses the same layers with fewer moving parts: hourly backup of the restore DC (24 per day) plus one nightly backup of a second DC (25 per day), and the recovery path loses the "other domains in parallel" step, so the 2.0 hours for that row drops out and the budget total is 5.5 hours.
Pitfalls
- Backups never restored, so the first test is the real event.
- Restoring a backup older than the tombstone lifetime.
- Restoring from a snapshot as a substitute for a backup.
- Restoring an FSMO holder "for simplicity" instead of restoring a plain DC and seizing the roles.
Sales users need a mapped network drive to the sales file share, a logon-time compliance script, and a locked-down user registry setting, all through Group Policy. How would you build the policy, limit it to the right people and machines, and test it before wide rollout?
Sample Answer
Direct answer
Put three things in one Group Policy Object (GPO, a bundle of settings applied to users or computers): a Group Policy Preferences drive map (a mapped drive that users may still change), a logon script under User Configuration, and a registry-based policy setting (enforced, so users cannot undo it). Link it first to a small pilot organizational unit (OU, a container for users and computers), scope it with security filtering (who receives the whole GPO) plus item-level targeting (who receives one preference item), prove it with gpresult and Group Policy Modeling, then widen the link.
Build
| Need | Where it lives | Notes |
|---|---|---|
| Mapped drive to the sales share | User Configuration, Preferences, Windows Settings, Drive Maps | Action Update (changes only the settings you define and creates the drive if missing; Replace deletes and recreates it). Location is a UNC path (a network path to a share) such as \\fs01\Sales. Tick Reconnect to restore the drive at each logon. Drive Maps run in the user's security context by default. |
| Compliance script at logon | User Configuration, Policies, Windows Settings, Scripts (Logon) | Scripts run only when a user logs on or off, never at the 90-minute background refresh. For something that must re-run on a schedule use the Scheduled Tasks preference extension (creates scheduled or immediate tasks). |
| Locked-down registry setting | User Configuration, Administrative Templates if a setting exists, otherwise a registry-based policy | Policy is enforced; a preference is not and can be changed by the user. Example from the cmdlet documentation: screen saver timeout 900 seconds. |
New-GPO -Name 'Sales - User Environment' |
New-GPLink -Target 'OU=Sales Pilot,OU=Staff,DC=corp,DC=contoso,DC=com' -LinkEnabled Yes
Set-GPRegistryValue -Name 'Sales - User Environment' `
-Key 'HKCU\Software\Policies\Microsoft\Windows\Control Panel\Desktop' `
-ValueName 'ScreenSaveTimeOut' -Value 900 -Type DWord
Set-GPPermission -Name 'Sales - User Environment' -TargetName 'Sales-Users' `
-TargetType Group -PermissionLevel GpoApply
# Downgrade the default entry from Read + Apply to Read only (do not delete it)
Set-GPPermission -Name 'Sales - User Environment' -TargetName 'Authenticated Users' `
-TargetType Group -PermissionLevel GpoRead -Replace
The hive in -Key decides the side: HKCU lands in User Configuration. Drive map and logon script are added in the Group Policy Management Console (GPMC) editor. The last command narrows the default security filtering so that only Sales-Users applies the GPO. It keeps Read for Authenticated Users on purpose: since the MS16-072 security update, user policy is retrieved using the computer's security context, so Microsoft's guidance is to give Authenticated Users (or, when you filter, Domain Computers) Read permission on the GPO. Removing the entry on the GPO's Scope tab instead leaves computer accounts unable to read it, and user settings can then fail to apply even for Sales-Users. Confirm with Get-GPPermission -Name 'Sales - User Environment' -All: Sales-Users should show GpoApply and Authenticated Users GpoRead.
Click path in the Group Policy Management Console (GPMC): under Group Policy Objects, right-click the GPO and choose Edit. For the drive, expand User Configuration, Preferences, Windows Settings, then right-click Drive Maps, New, Mapped Drive. Set Action to Update, Location to \\fs01\Sales, a drive letter such as S:, tick Reconnect, and use the Common tab for item-level targeting. For the script, go to User Configuration, Policies, Windows Settings, Scripts (Logon/Logoff), double-click Logon and use Add to point at the script.
Scoping: people and machines
- User settings follow the user object, so link where the Sales users live. A GPO applies to every user in that OU unless filtered.
- Security filtering works on the whole GPO and cannot be used selectively on different settings inside it.
- Item-level targeting works per preference item and is evaluated on the client by the preference extension. Each targeting item evaluates to true or false, and several can be combined with AND or OR. Put a Security Group targeting item on the drive map (it applies when the user or computer is a member of the named group) and, if the drive must appear only on office PCs, add an Organizational Unit targeting item for the computer's OU, or a Site or IP Address Range item (the name Microsoft's documentation gives the IP range check; that item does not support IPv6 addresses).
- To limit by machine type, link one WMI filter (a Windows Management Instrumentation (WMI) query evaluated on the destination computer; the GPO applies only if it returns true) to the GPO. A GPO can hold one WMI filter.
Test before wide rollout
- Back up the GPO:
Backup-GPO -Name 'Sales - User Environment' -Path \\fs01\gpbackup. - Pilot OU with two test users and one workstation. Run
gpupdate /forceon the client, then log off and on so logon-time items run. - Prove it:
gpresult /h pilot.html(the GPO must appear as applied),net usefor the drive, check the log written by the script, and read the registry value. - Preview other populations without logging in as them using Group Policy Modeling in GPMC. Modeling is a wizard that simulates what would apply to a chosen user and computer location and reports the GPOs and settings that would win. It reads no client, so it answers "what if Raj moved into the Sales OU" before anyone is moved, while
gpresultreports what a client actually applied. - Widen: relink to the Sales OU with the
Sales-Usersfilter; unlink is the rollback for policy settings. A preference needs its own reverse item (for the drive: action Delete).
Worked example
The pilot OU holds two test users: Ann, a member of Sales-Users, and Raj, who is not. After gpupdate /force and a fresh logon, gpresult /h for Ann lists the GPO as applied, drive S: appears and the log file is written. For Raj the GPO does not apply and no drive appears, which proves the filter works before any real Sales user is touched.
The gpresult summary for the user scope has this shape (abridged, illustrative; run gpresult /r /scope user as the pilot user). For Ann:
USER SETTINGS
Applied Group Policy Objects
-----------------------------
Sales - User Environment
For Raj the same GPO moves to the other list, with the reason:
The following GPOs were not applied because they were filtered out
-------------------------------------------------------------------
Sales - User Environment
Filtering: Denied (Security)
Read it as: the GPO name under "Applied Group Policy Objects" is the "must appear as applied" check, and a GPO listed under "filtered out" with a Security reason means the security filtering did its job.
Pitfalls
- A preference that a policy also sets loses to the policy. For example, when a registry preference sets the screen saver timeout to 600 seconds and the policy in this GPO sets 900, the user ends up with 900.
- The preference option Apply once and do not reapply (the item is applied the first time only) silently stops the drive from being repaired: a drive the user later disconnects or deletes is never put back.
- A script that needs the share before the drive exists should map by UNC path, not by letter.
- If a change to Location or Reconnect does not reach clients that already have the drive, switch the action to Replace, which deletes and recreates the item.
Your incident team believes an attacker holds persistent forged-ticket access to the domain after compromising a DC. How does that attack work, how would you detect it, and how do you evict the attacker?
Sample Answer
Direct answer
Every Kerberos ticket-granting ticket (TGT) is protected with the secret of the krbtgt account, which exists only on domain controllers (DCs). An attacker who copies that secret from a compromised DC can mint TGTs offline for any account, with any group memberships, and keep minting them after passwords change: a "golden ticket". Pulling the secret out of a DC by abusing directory replication is called DCSync. Detection is probabilistic, so the response does not depend on proving it. You evict by closing the attacker's access first, then rotating the krbtgt password twice with a wait of at least the maximum ticket lifetime between the resets, together with resetting privileged credentials and rebuilding or restoring DCs you cannot trust.
How the attack works
A normal day in two steps. Step one, the AS exchange: you type your password, the DC checks it and hands back a TGT sealed with the krbtgt secret. Step two, the TGS exchange: to open a file share you show that TGT to the DC, which unseals it, trusts what is inside, and hands back a service ticket for the share. A forged TGT skips step one entirely.
In Kerberos the client first talks to the Authentication Service (AS exchange) and receives a TGT, then presents that TGT to the Ticket-Granting Service (TGS exchange) to obtain service tickets. The Key Distribution Center (KDC) on the DC accepts any TGT it can decrypt with the krbtgt key. If the attacker holds that key (for example through directory replication or by copying the directory database), they can build a TGT without ever doing the AS exchange, naming a real or invented account and the group identifiers they want, with a lifetime they choose. Default policy limits legitimate TGTs to 10 hours, renewable for 7 days, but the forged ticket is not issued by the KDC, so those limits do not constrain the attacker's ability to create new ones. It stays usable until the key it was signed with is no longer accepted.
Detection (probabilistic)
- Service tickets without a matching logon. Event 4769 ("a Kerberos service ticket was requested") is logged on a DC; event 4624 ("an account was successfully logged on") is logged on the computer the user signed in to. Both carry a Logon GUID, an identifier for that sign-in session, so the GUID in 4769 lets you find the matching 4624. Illustrative fields to read in a 4769: Account Name
jsmith@CORP.CONTOSO.COM, Service Namecifs/file01, Ticket Encryption Type0x12, Client Address10.1.4.20, Logon GUID{...}. A TGS request for an account that has no AS exchange and no logon trail anywhere is suspicious. Collect events from all DCs, because the AS exchange may have been handled by another one. - Encryption type. Microsoft's monitoring guidance expects 0x11 or 0x12 (AES) in 4769; RC4 (0x17) requests for accounts that normally use AES stand out.
- Accounts that should not be active: 4769 for disabled accounts, accounts that should never be used, or from unexpected client addresses.
- How the key was taken: event 4662 ("an operation was performed on an object") on the domain object, from an account that is not a DC, showing the directory replication rights (such as Replicating Directory Changes All), indicates DCSync.
- Exposure window: the
pwdLastSetofkrbtgt(the attribute that records when the password was last changed). If it has not changed in years, every past theft of that secret is still valid.
A well-forged AES ticket looks normal, which is why evidence of absence proves nothing.
Eviction plan (ordered)
- Contain first. Isolate the compromised DC(s), disable the access that was used, remove persistence (look at AdminSDHolder and protected-group ACLs, replication rights, GPOs on the domain and DC OU). If the attacker can still read the directory, a new
krbtgtkey will simply be stolen again. - Reset privileged credentials next: every member of Enterprise Admins, Domain Admins, Schema Admins, Server Operators, Account Operators and similar groups. Recreate group managed service accounts (gMSAs) with a new KDS root key (the master key their automatic passwords are derived from) if the database was exposed.
- Reset
krbtgttwice. In Active Directory Users and Computers (Advanced Features on), right-clickkrbtgtand choose Reset Password. The value you type does not matter, because the system generates a strong password regardless. Do it on a writable DC; read-only DCs (RODCs) each have their ownkrbtgt_<number>account, which this does not reset. An RODC's key signs tickets only for the accounts that RODC serves, so one exposed RODC is a separate, smaller job: reset that RODC'skrbtgt_<number>account the same way if it was compromised. - Wait between the resets at least as long as the maximum lifetime for user and service tickets (10 hours by default; longer if you raised it), and confirm replication reached every DC.
- Reset the second time. The account keeps its two most recent passwords, so the second reset pushes the attacker's old key out of history.
- Verify and watch. Confirm every DC shows the new change time, run
repadmin /replsum, and keep hunting for 4769 failures and RC4 requests.
Why twice, and why wait
The account stores two passwords. After reset 1 the KDC accepts tickets signed with the new key and the previous one, so the attacker's forged tickets still work. After reset 2 the original key is gone from history, so tickets signed with it fail. If you did the second reset immediately, every legitimate TGT issued before the first reset would also be rejected and everyone would re-authenticate at once. Waiting at least the ticket lifetime lets legitimate TGTs refresh onto the key that survives. Timeline with clock times (illustrative): reset 1 at 08:00. A user who signed in at 07:30 holds a TGT valid to 17:30 under the old key, which the DC still accepts as the previous password. The latest any pre-reset TGT can run is 08:00 plus 10 hours, so 18:00. Reset 2 at 18:05 or later then discards the old key, and by then legitimate users have refreshed onto the new one. This timing is the documented guidance; shortening it is an availability decision for the incident commander, not a documented procedure.
Run this after the first reset. It confirms every DC has the new value and that the minimum wait has passed before you run the second:
$domainDn = (Get-ADRootDSE).defaultNamingContext
$minWait = [timespan]::FromHours(10) # raise this if Maximum lifetime for user or service ticket is above 10 hours
$rows = Get-ADObject -SearchBase "OU=Domain Controllers,$domainDn" -LDAPFilter '(objectClass=computer)' -Properties dNSHostName |
ForEach-Object {
$k = Get-ADUser -Identity krbtgt -Server $_.dNSHostName -Properties pwdLastSet
[pscustomobject]@{ DC = $_.dNSHostName; KrbtgtSetOn = [datetime]::FromFileTime($k.pwdLastSet) }
}
$rows | Sort-Object KrbtgtSetOn | Format-Table -AutoSize
$distinct = @($rows.KrbtgtSetOn | Select-Object -Unique)
if ($distinct.Count -ne 1) {
'WAIT: DCs disagree on the krbtgt change, replication has not converged.'
}
elseif (((Get-Date) - $distinct[0]) -lt $minWait) {
"WAIT: only $([int]((Get-Date) - $distinct[0]).TotalMinutes) minutes since the last reset."
}
else {
'READY: all DCs agree and the minimum wait has passed. Run the next reset.'
}
Illustrative output 95 minutes after reset 1, with all DCs in agreement:
DC KrbtgtSetOn
-- -----------
dc1.corp.contoso.com 10/6/2026 8:00:12 AM
dc2.corp.contoso.com 10/6/2026 8:00:12 AM
dc3.corp.contoso.com 10/6/2026 8:00:12 AM
WAIT: only 95 minutes since the last reset.
The table shows when each DC last saw the change; one differing row would print the first WAIT message instead (replication not converged). READY appears only when all DCs agree and the 10 hour minimum has passed.
Choosing the recovery option
| Option | Use when | Cost |
|---|---|---|
| Targeted eviction (steps above, rebuild the touched DCs) | The compromise window and the touched DCs are well understood | Moderate; confidence depends on scope |
| Forest recovery from a pre-compromise backup, in isolation | The directory itself is untrusted | Loses all changes since the backup; long outage |
| Migrate critical assets to a pristine forest | Trust in the existing forest cannot be rebuilt | Highest, but strongest assurance |
Trade-offs and pitfalls
- Resetting
krbtgtbefore you have stopped the attacker is wasted effort, because they can read the new key again. - A single reset is not enough: the old key stays in the password history.
- When the directory database was exposed, resetting trust passwords and recreating group Managed Service Accounts (gMSAs) can be part of the same job.
You have lost the only writable domain controller in a site and have known-good backups. How do you recover, and how do you decide which kind of restore to perform?
Sample Answer
Direct answer
Decide first whether the lost domain controller (DC) should be restored at all. If any other healthy writable DC still holds the domain, do not restore it. Remove the dead DC's metadata, seize any flexible single master operations (FSMO) roles it held, and promote a fresh DC in the site, which copies a current directory from its partners. Restore from backup only when the domain has no other writable DC: that is a nonauthoritative restore of the directory (the DC comes back as the backup left it, then accepts newer changes from partners) plus an authoritative restore of SYSVOL (the shared folder that holds Group Policy files and logon scripts; marking it authoritative makes this copy the one other DCs rebuilt later copy from) on that first DC. Seizing a role, mentioned here and used below, means another DC takes over a role without the old holder's cooperation; transferring is the graceful hand-over while the old holder is alive. An authoritative restore of individual objects is a third tool, used to bring back deleted objects rather than to replace a server. A backup older than the tombstone lifetime is unusable for all of them.
How to decide
A nonauthoritative restore brings the DC back as the backup left it and then lets normal replication from partners bring it up to date. An authoritative restore marks chosen objects so that they overwrite what partners hold. System state is the backup set that includes the directory database and SYSVOL.
| Situation | Do this | Restore type |
|---|---|---|
| Lost DC, other healthy writable DCs exist for the domain | Seize roles if needed, clean up metadata, promote a new DC | None. The new DC replicates a current copy |
| Lost DC was the only writable DC in the domain | Isolate, restore system state, verify, then rebuild around it | Nonauthoritative restore of Active Directory Domain Services (AD DS) and authoritative restore of SYSVOL, on this first DC only |
| Objects deleted by mistake, DCs healthy | Restore from the Active Directory Recycle Bin when it is enabled and the objects are inside the deleted-object lifetime; otherwise restore them from backup | Recycle Bin restore needs no backup; the backup route is an authoritative restore of those objects |
| Directory corrupted or compromised across the forest | Follow the forest recovery plan | Restore one DC per domain from a backup that predates the damage |
| Virtual DC reverted with a snapshot | Treat as an incident, not a backup | Microsoft does not recommend a VM snapshot as a substitute for backing up a DC |
Why a backup older than the tombstone lifetime is unusable
Deleted objects are kept as tombstones for the tombstone lifetime and then purged. That lifetime is 180 days in forests created with Windows Server 2003 SP1 or later and 60 days in forests upgraded from Windows 2000 or from Windows Server 2003 without a service pack. If the AD Recycle Bin is on, a backup is good only for the lesser of the deleted-object lifetime and the tombstone lifetime.
The question gives no backup date, so take an example: a backup taken on 5 March 2026 and used on 5 October 2026 is 214 days old. That is 34 days past a 180-day lifetime, which expired on 1 September 2026, and 154 days past a 60-day one, which expired on 4 May 2026. Objects deleted in that gap have been purged on every other DC, but the restored copy still contains them. Those are lingering objects: objects deleted and purged elsewhere that survive on the stale DC and can be replicated back. A DC that has not replicated within the tombstone lifetime has its inbound replication stopped, logs Event ID 2042 (it has been too long since this machine last replicated) and reports error 8614 (the time since last replication exceeds the tombstone lifetime), because the two sides' views of deleted objects may differ.
Ordered procedure, rebuild path (other writable DCs exist)
Scenario: site Pune had a single DC, pune-dc1, which is gone. hq-dc1 and hq-dc2 serve corp.contoso.com from another site.
- Scope it. From
hq-dc1runrepadmin /replsumfor current replication health andnetdom query fsmoto see whetherpune-dc1held any of the five roles.repadmin /replsumsummarizes inbound and outbound replication per domain controller, with alargest deltaand afails/totalcount for each. Illustrative reading for this scenario: the rows forhq-dc1andhq-dc2should show no failures between them, whilepune-dc1will keep showing failures for as long as its metadata still exists, which is expected here and stops after step 3.netdom query fsmoprints the current holder of each of the five roles. A role still pointing atpune-dc1is the one to seize. Clients in Pune will find DCs in other sites meanwhile, with WAN latency. - Seize only the roles it held. In
ntdsutil:roles, thenconnections, thenconnect to server hq-dc1.corp.contoso.com, thenquit, then the matching command:seize schema master,seize naming master,seize pdc,seize rid masterorseize infrastructure master. Seizure is for a holder that will never return, so wipe the old machine rather than reconnecting it. Seize the RID master only when the old holder cannot come back, because two live holders of the role could hand out the same RIDs. Microsoft documents that the schema master, domain naming master, RID master and infrastructure master are active only after the new holder has inbound replicated its partition since its directory service started (the PDC emulator has no such requirement), so do not expect those roles to work the instant the command returns. - Clean up metadata. Delete the
pune-dc1computer object from the Domain Controllers OU in Active Directory Users and Computers, tick "This Domain Controller is permanently offline", and confirm. That removes the replication identity. Then, from a healthy DC, runnltest /dsderegdns:pune-dc1.corp.contoso.com, which deregisters the DNS host records of the named host, and check the DNS zone afterwards for leftovers. - Promote the replacement in the existing Pune site, replicating from a healthy partner:
$params = @{
DomainName = 'corp.contoso.com'
SiteName = 'Pune'
InstallDns = $true
ReplicationSourceDC = 'hq-dc1.corp.contoso.com'
Credential = (Get-Credential)
}
Install-ADDSDomainController @params
If the WAN is slow, point -InstallationMediaPath at installation media instead of pulling everything over the link.
5. Verify. repadmin /replsum, dcdiag /v, netdom query fsmo, and the Global Catalog check box on the new DC's NTDS Settings object in Active Directory Sites and Services (or (Get-ADDomainController -Filter { Name -Eq 'NEW_DC_NAME' }).IsGlobalCatalog, which returns True or False). A global catalog server is a DC that also holds a searchable partial copy of every domain in the forest. Healthy means no failures in the repadmin /replsum rows for the live DCs, every dcdiag test reporting that it passed, and role holders that are live DCs.
Ordered procedure, restore path (no other writable DC)
- Keep the DC off the production network (unplugged, or on an isolated virtual network).
- Know the Directory Services Restore Mode (DSRM) password (the special boot mode in which the directory database is offline and can be restored) and have a system state backup. A full server backup used for server recovery is not enough: the backup must explicitly include system state.
- Restart the DC into Directory Services Restore Mode (the reason you need the DSRM password) and restore with
wbadmin start systemstaterecovery -version:<backup version> -authsysvol. The-authsysvolswitch marks SYSVOL authoritative and belongs only on the first DC restored for the domain, because only one copy should be the authoritative source that the DCs rebuilt later copy from. The command reports progress while it restores and finishes with a restart. - Check that the data is sound. If not, repeat with an older backup that is still inside the tombstone lifetime.
- Seize all roles on this DC, clean up the metadata of every other writable DC you are not restoring, raise the available RID pool by 100,000, and reset the DC computer account and
krbtgtpasswords twice, as the Microsoft forest recovery guide specifies. Why: a relative ID (RID) is the per-account suffix of a security identifier (SID), and each DC hands out RIDs from a pool. Accounts created after the backup were lost, but their SIDs may still sit in access lists; without a raised pool, a new account could receive the same SID and inherit that access, so the pool jumps past anything the lost DC might have issued.krbtgtis the account whose secret signs every Kerberos ticket-granting ticket; it keeps its two most recent passwords, so two resets push the pre-failure password out of history and invalidate tickets forged or issued with it. - Rebuild additional DCs by promotion, then take a fresh backup.
Trade-offs and pitfalls
- Rebuild beats restore when partners exist: it starts from current data, avoids stale SYSVOL, and carries no tombstone risk.
- Virtual DCs: reverting a snapshot rolls the update sequence numbers (USN, the counter each DC uses to tell partners which of its changes they have already received) back and can leave replication quietly inconsistent. Hosts that expose VM-GenerationID let a Windows Server 2012 or later DC detect it, but Microsoft still says snapshots are not a substitute for backups.
- Different server: restoring a system state backup onto a fresh Windows installation is not supported, and restoring it to a different server than it came from is not recommended.
- Global catalog: restoring a global catalog from a backup newer than the other restored domains' backups can create lingering objects, so clear the global catalog flag on that DC when the forest has several domains.
Unlock Full Question Bank
Get access to all 33 Active Directory Architecture and Management interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.