Filesystem Forensics and Data Recovery Questions
Disk and file-system internals for forensic examination and data recovery: NTFS/ext4/APFS/FAT structures (MFT, journaling, inodes), file carving, slack and unallocated space, deleted-file recovery, SSD/TRIM and wear-leveling effects on recoverability, RAID reconstruction, and acquisition/analysis of damaged or encrypted media. Distinct from OS-level artifact analysis (registry, Prefetch, event logs, memory/LSASS, persistence mechanisms) and from timeline construction across sources, which are covered by forensic-artifact-analysis-and-timeline-reconstruction.
A file was deleted three days ago on a system that's still in active use, writing a few gigabytes of new data a day, with roughly 15% free space remaining. How would you reason about the odds that the file's data clusters haven't been overwritten yet, and what would you actually tell the client about the chances of recovery?
Sample Answer
Direct answer
With 15% free space and normal write activity, you cannot promise the client the file is intact. You can give a defensible probability for any ONE of the file's clusters (a cluster is the smallest fixed-size chunk of disk space the filesystem hands out to a file, 4 KiB in the numbers below), and a much smaller, more honest probability that ALL of them survived untouched, and multi-cluster files degrade fast under this kind of write pressure even when each individual cluster's odds look decent.
Structured elaboration
Model new writes as landing uniformly at random across the pool of currently-free clusters (a stated simplification: real allocators have locality and this ignores it, but it is the standard, defensible first-order model to state explicitly as an assumption). If N is the number of free clusters and n write-sized "cluster draws" have happened since deletion, the probability any one specific cluster was never touched is:
P(one cluster survives)≈e−n/N
For a file spanning k clusters, assuming independence between them:
P(all k clusters intact)≈e−nk/N
The exponent for the whole file scales with k, which is exactly why a bigger file's odds fall off a cliff compared to any one of its individual clusters, even though each cluster on its own still looks reasonably likely to have survived.
Worked example
Pinned assumptions: a 1 TB drive with 15% free space, 4 KiB clusters, a middle-of-range write rate of 5 GB/day for "a few gigabytes a day," 3 days elapsed, and a 2 MiB deleted file.
import math
# pinned assumptions, stated explicitly because the conclusion depends on them
disk_free_bytes = 150 * 10**9 # 15% free on a 1 TB drive
cluster_size = 4096 # 4 KiB clusters
write_rate_bytes_per_day = 5 * 10**9 # "a few GB/day" -> middle-of-range estimate
days_elapsed = 3
file_size_bytes = 2 * 1024 * 1024 # a 2 MiB deleted file
N = disk_free_bytes // cluster_size # free clusters in the pool
n = (write_rate_bytes_per_day * days_elapsed) // cluster_size # cluster-writes since deletion
k = file_size_bytes // cluster_size # clusters the file occupied
p_single_survives = math.exp(-n / N)
p_all_k_survive = math.exp(-n * k / N)
print(f"free clusters N = {N:,}")
print(f"cluster-writes so far n = {n:,}")
print(f"file clusters k = {k}")
print(f"n/N = {n/N:.4f}")
print(f"P(one specific cluster untouched) = {p_single_survives:.4f}")
print(f"P(all {k} clusters still intact) = {p_all_k_survive:.3e}")
Output:
free clusters N = 36,621,093
cluster-writes so far n = 3,662,109
file clusters k = 512
n/N = 0.1000
P(one specific cluster untouched) = 0.9048
P(all 512 clusters still intact) = 5.809e-23
With those pinned numbers, 10% of the free pool has been overwritten in the 3 days since deletion, so any ONE cluster has about a 90% chance of being untouched, that is the number to lead with. But a typical multi-megabyte file spans hundreds of clusters, and asking "did ALL of them survive" compounds that 90% hundreds of times over, landing at a probability so small it is effectively zero. That is not a contradiction, it is the mathematically honest reason full recovery of anything but a tiny file gets unlikely fast under active write pressure, even when individual bytes have decent survival odds.
What to tell the client: do not quote either number in isolation. Something like "a meaningful fraction of this file's clusters are probably still there, but the odds of a complete, byte-perfect recovery are low given three days of active use at this write rate, expect partial recovery via carving and fragment reassembly, not a guaranteed whole file" sets an honest expectation. State the assumptions (free-space fraction, average daily write volume, uniform-random write placement) explicitly, since the number changes a lot if any of them is wrong, most of all if the real write pattern is concentrated (a database growing one file, say) rather than spread evenly.
Trade-offs & pitfalls
The uniform-random-placement assumption is the weak link: real allocators often favor previously-free contiguous extents, which could mean the deleted file's own former clusters are somewhat MORE likely to be reused soon, and workloads have hot/cold skew rather than uniform load. Treat this model as a first-order estimate that motivates urgency and a starting number, not a certified figure, and revisit it if the actual write pattern turns out not to be close to uniform. Where possible, validate empirically rather than purely analytically: allocate test files of known size, delete them, run a representative workload for the same elapsed time, and measure the actual survival rate on a comparable test system before quoting the client a number with more confidence than the model deserves.
A suspect used BitLocker and the disk shows signs of wiped container headers. Propose a forensic methodology to attempt recovery of residual header fragments, backup header locations, or potential key material. Include technical steps (search signatures, slack, registry/AD backups, TPM artifacts), legal channels (recovery key requests), and limitations where cryptographic protection precludes recovery.
Sample Answer
Direct answer
BitLocker (Windows' built-in full-disk encryption) stores its configuration and wrapped encryption keys in a Full Volume Encryption (FVE) metadata structure, and it deliberately keeps three identical copies of that metadata at different locations on the volume, specifically so damage to one does not lock out the whole volume. That built-in redundancy is the first and best place to look before turning to broader search-and-carve techniques or legal channels for a recovery key.
Technical recovery steps
- Search for surviving metadata copies first: each FVE metadata block begins with the fixed 8-byte signature "-FVE-FS-", and since a volume normally carries three copies at known, sector-aligned offsets, a wiped "primary" copy still leaves two others to search for at those known locations.
- Search slack and unallocated space more broadly: if all three on-volume copies were deliberately wiped, search the whole image, not just the expected offsets, for the same signature, in case a fragment survives in slack space (the unused tail of a disk cluster, left over when a smaller file occupies a cluster a larger one used to fill) from a prior volume layout or an old shadow copy (Windows' automatic point-in-time snapshot of a volume, kept so files can be rolled back).
- Registry and Windows artifacts: parse the SYSTEM and SECURITY registry hives and BitLocker's own Windows Event Log channel for protector GUIDs (Globally Unique Identifiers), which confirm which recovery mechanism (Trusted Platform Module, PIN, USB key, recovery password) was configured, even without the metadata itself.
- Active Directory and enterprise backups: if the machine was domain-joined with BitLocker recovery escrow enabled, the 48-digit recovery password is stored on the corresponding computer object's msFVE-RecoveryPassword attribute in Active Directory; a domain controller backup or an Azure AD/Intune export can hold this even if the endpoint's own copies are gone.
- Trusted Platform Module (TPM), the hardware chip many BitLocker configurations use to hold or seal the encryption key to a trusted boot state: TPM event logs and Platform Configuration Register values can corroborate that BitLocker was active and how it was configured, though extracting the key itself from a TPM without cooperation is a specialist hardware attack, not a standard forensic step.
- Legal channel: submit a formal recovery-key request through the account or organization that manages the device (a Microsoft account, Azure AD/Intune, or a corporate IT department), backed by appropriate legal authority.
Validating a partially recovered recovery key without a computer
BitLocker's 48-digit recovery password is organized as eight groups of six digits, and each six-digit group is required to be evenly divisible by 11. This is not a security feature, it is a built-in typo check, so an examiner who has recovered a partial or smudged key, say from handwritten notes or a photographed screen, can sanity-check each group by hand before wasting time trying it. The check below is worked by hand, not executable output:
(hand-worked check, not executable output)
Example six-digit group: 594671
594671 / 11 = 54061 (exact, remainder 0) -> passes the check-digit rule
A group that does not divide evenly by 11 is provably wrong, a transcription error, before it is ever tried against the drive, which is a useful, cheap way to validate a recovery key an investigator obtained secondhand rather than pulled directly from the metadata or Active Directory.
Limitations
If all three on-volume FVE metadata copies, every registry trace, every Active Directory or Intune escrow record, and the TPM path are genuinely gone, BitLocker's underlying encryption is not something that can be brute-forced in a forensically useful timeframe; the honest conclusion is that recovery depends entirely on finding some surviving copy of the key material, not on attacking the cryptography itself. Do not conflate "found a recovery password" with "confirmed it unlocks this specific volume"; a device can have multiple protectors or have been re-encrypted, so validate any recovered key against the actual volume before reporting it as the answer. Enterprise escrow and TPM extraction both typically require legal process or vendor and employer cooperation; document each request made and its outcome for the report, including negative results.
You're carving files from unallocated space where the headers and footers have been overwritten, so plain signature matching won't find them. How would you go about detecting and reconstructing files without a header or footer to search for, and how would you keep the false-positive rate under control?
Sample Answer
Direct answer
When headers and footers are gone you cannot search for a signature, so detection has to come from what a chunk of bytes looks like internally: its entropy profile, and format-specific internal structure that survives even without a header (XML tags inside a DOCX, atom boxes inside an MP4, cross-reference patterns inside a PDF). Score every candidate on more than one independent signal rather than trusting any single one, and only accept reconstructions that clear a combined threshold.
Structured elaboration
- Coarse filtering by entropy. Compute Shannon entropy over a sliding window (a few KB, 50% overlap is a common choice). Compressed and encrypted content, which most modern document and media formats effectively are, reads as high entropy, close to the theoretical maximum of 8 bits/byte; plain text reads noticeably lower; an unwritten cluster reads as flat zero. This cannot identify a format on its own, only flag "something structured or compressed lives here" versus "this is empty or plain text."
- Structure-driven anchoring. Within flagged regions, search for internal markers that survive header loss: a ZIP local-file-header signature
PK 03 04or an<?xmlfragment inside DOCX/XLSX/PPTX;ftyp/moov/mdatatom tags inside MP4, each preceded by a 4-byte size field you can sanity-check against the region length;obj/endobjmarkers and cross-reference patterns inside PDF. An anchor that actually parses under the format's own rules, not just matches bytes, is much stronger evidence than the anchor bytes alone. - Combine into one score. Require at least two independent signals to agree before accepting a candidate: entropy in the expected range, a structural anchor that parses, and, where the format has one, an internal integrity check (a chunk length that does not overrun the region, a checksum). Downrank anything that hits only one signal, since a coincidental byte match with nothing else backing it is exactly the shape a false positive takes.
- Resolving ambiguous joins. When two candidate fragments both plausibly extend the same anchor, prefer the one whose declared internal size field actually lines up with where it starts; where that is still a tie, prefer physical proximity on disk over a distant match, since sequential writes are the common case. Where it remains genuinely ambiguous, report both candidates with a confidence score rather than silently picking one.
Worked example
Entropy alone cannot identify a file format, but it is cheap and worth seeing with real numbers. This computes Shannon entropy for three representative 4 KB clusters:
import math, random
from collections import Counter
def shannon_entropy(data):
if not data:
return 0.0
counts = Counter(data)
n = len(data)
return abs(-sum((c / n) * math.log2(c / n) for c in counts.values()))
random.seed(7)
compressed_like = bytes(random.randrange(256) for _ in range(4096)) # stand-in for JPEG/ZIP payload bytes
text_like = (b"the quick brown fox jumps over the lazy dog " * 90)[:4096] # stand-in for a text/log cluster
zeroed = bytes(4096) # stand-in for a never-written cluster
print(f"random/compressed-like cluster: {shannon_entropy(compressed_like):.2f} bits/byte")
print(f"english-text-like cluster: {shannon_entropy(text_like):.2f} bits/byte")
print(f"all-zero (unwritten) cluster: {shannon_entropy(zeroed):.2f} bits/byte")
Output:
random/compressed-like cluster: 7.96 bits/byte
english-text-like cluster: 4.34 bits/byte
all-zero (unwritten) cluster: 0.00 bits/byte
A compressed-media or ZIP-based cluster reads close to the 8-bit ceiling; readable text reads noticeably lower; a never-written cluster reads as flat zero. That 7.96 alone does not say "this is a DOCX" rather than an encrypted volume or random slack, it only says "worth checking for structural anchors," which is exactly why entropy is stage 1, not the whole pipeline.
Trade-offs & pitfalls
- False positives are the central risk here, more than in signature carving, precisely because the strongest signal (an exact byte match) is gone. Multi-signal scoring reduces but never eliminates them; treat borderline scores as candidates for a human examiner to review, not as auto-accepted recoveries.
- Scale: bucket regions by entropy class first, since that is a single pass over the image, and only run the more expensive structural parsers on regions that already look interesting.
- Report a confidence score with every reconstruction, not just a yes/no: "probably a DOCX, medium confidence, joined at an ambiguous boundary" is honest and still useful evidence, a silent wrong guess is not.
- This whole approach is probabilistic by construction, which matters for how you write it up: state the scoring method and thresholds in the report so the reasoning behind an accepted recovery is reproducible, not just asserted.
TRIM is enabled on the SSD you're examining, so the usual 'unallocated space still holds the old data' assumption doesn't hold. You have options ranging from software-level recovery on a logical image up through invasive hardware-level techniques. How would you decide how far up that ladder to go for a given case, and what are you weighing at each step: success likelihood, risk of destroying the evidence, the expertise required, and chain-of-custody?
Sample Answer
Direct answer
Treat this as a staged decision, not a single choice: start with the least invasive, cheapest, most defensible option, probabilistic carving from a logical image, and only climb toward chip-off or vendor/controller-level recovery when there is a specific reason to believe the next rung will find something the current one cannot, weighed explicitly against the risk of destroying evidence, the expertise it demands, and what it does to your chain-of-custody story, the unbroken documented record of who held the evidence and what they did to it that is what lets a court trust the exhibit was not altered.
Structured elaboration: what to weigh at each rung
-
Probabilistic carving from a logical image.
Success likelihood: low to moderate for anything TRIM (the command the operating system sends an SSD to announce that a block is no longer in use) has already reached, since the controller may have already erased the underlying flash; carving can still recover fragments in untrimmed slack or from before the last garbage-collection pass. Risk: minimal, non-destructive, works from a standard forensic image. Expertise: moderate, standard carving tools and format validation. Chain-of-custody: straightforward, a normal write-blocked image and documented process. Escalate past this rung when carving comes up empty or clearly partial AND the case value justifies more. -
Raw NAND acquisition (chip-off).
Why this can beat logical imaging at all: the flash translation layer (FTL) maps host-visible logical block addresses to physical NAND pages, and it is exactly the layer TRIM and garbage collection operate through. A logical image only ever sees what the FTL currently chooses to expose; chip-off reads the physical NAND directly, bypassing the FTL and potentially recovering pages the controller has marked invalid but not yet physically erased, or pages sitting in over-provisioned space the host can never address at all. Risk: high, desoldering and reading chips can permanently damage them, and even a successful read produces raw data that needs reconstructing against that specific controller's wear-leveling and FTL scheme to mean anything. Expertise: high, specialist lab equipment and technique. Chain-of-custody: every physical intervention has to be individually documented; expect scrutiny in court precisely because it is invasive. Escalate past this rung when the case value justifies the cost and risk, an accredited lab is available, and legal authorization covers a destructive technique. -
Controller reprogramming or vendor firmware recovery.
Highest potential payoff: reconstructing FTL mappings via vendor tooling or cooperation can recover data inaccessible any other way, but only when the specific controller and vendor relationship support it, this is not a generic technique applicable to any SSD. Risk: moderate to high, wrong commands can trigger secure-erase routines or accelerate the exact loss being avoided. Expertise: very high, often requiring vendor NDAs and cooperation rather than one examiner working alone. Chain-of-custody and legal/ethical constraints: vendor assistance typically means an NDA and documented external involvement the court needs to understand; non-standard recovery methods invite scrutiny, so the legal authorization for this rung specifically should be settled before starting, not justified after the fact.
Decision framework: what you are actually weighing at each step
Success likelihood given what is already known about this specific case (has TRIM run, how much time has passed, is there other corroborating evidence of the file's existence); the case's evidentiary value, is it worth risking physical damage to the only copy of the evidence; the irreversibility of the next rung, can you still fall back if it fails, or does it consume the option to try anything else; and whether you and your lab actually have the expertise the next rung demands, borrowed expertise under time pressure is where mistakes happen.
Trade-offs & pitfalls
Never treat this as "try the cheap thing, then automatically try the expensive thing"; each escalation should be a deliberate, documented decision with a stated reason, not a default fallback. Vendor or controller-level recovery in particular needs its legal and ethical basis established up front, a warrant or client authorization that specifically covers destructive or vendor-assisted methods, not just the general search authority the case started with.
Explain the TRIM command and SSD garbage collection behavior. Discuss how these features shorten the window for recovering deleted files on flash storage, and describe practical steps an examiner can take during acquisition to maximize recovery chances from an SSD or NVMe device.
Sample Answer
Direct answer
TRIM is the command (ATA TRIM, or the NVMe Dataset Management command's Deallocate attribute) the OS uses to tell an SSD which logical blocks are no longer in use. Garbage collection is the controller's own background housekeeping that compacts still-valid data and erases the freed physical NAND blocks so they can be written again. Together they mean a deleted file's physical bytes can be erased well before you get the drive on your bench, with no reliable way to predict exactly when, which is why the acquisition strategy for SSD/NVMe evidence is built around minimizing further TRIM/GC activity rather than trying to race it.
Structured elaboration
- Why the window shrinks. TRIM marks LBAs (logical block addresses, the numbered slots the OS uses to refer to locations on the drive) invalid immediately; garbage collection then erases the underlying flash independently of new writes, and it can run during any idle period, not only in response to new host activity, so more elapsed time and more OS activity between delete and acquisition both raise the chance the data is already gone by the time you image the drive.
- Immediate actions on seizure. Do not power on or boot a device that is currently off. If it is already running and cannot safely be left running, follow standard live-triage priorities, capturing volatile memory and open file handles first, since those may hold content or mappings TRIM cannot touch, before powering down. Physically isolate the device from any host that could issue further TRIM commands, and never mount it read-write on a live OS.
- Acquisition steps that maximize recovery chances. Use a write-blocker (a device or driver layer that lets read commands reach the drive but stops every write, so the act of imaging cannot change the evidence) or imaging tool specifically verified NOT to pass TRIM/UNMAP/Deallocate commands through to the device, confirmed against the vendor's own documentation rather than assumed. Prefer a full physical/raw image over a logical one where the tool supports it, since logical imaging alone can miss data the controller still holds outside the host-visible address space. Collect SMART data, firmware version, and any available controller/FTL logs, since these help reason about wear-leveling and garbage-collection behavior for that specific model. Where the case value justifies it, escalate to vendor-assisted or chip-off recovery, understanding both are invasive and chip-off specifically requires reconstructing the flash translation layer's logical-to-physical mapping from raw NAND, not something to improvise in the field.
Communicating this to stakeholders
Say plainly, and early, that SSD/NVMe recovery odds are structurally worse than HDD and get worse the more time and OS activity elapses since deletion; that "we cannot promise anything" is not the same as "we will not try"; and that immediate isolation of the device, before it even comes to you, is the single highest-leverage thing a stakeholder can do to help. Avoid quoting a specific recovery percentage without case-specific data (device model, TRIM support and configuration, elapsed time) to back it up: an unsupported number damages credibility later far more than an honest "here is what improves the odds, and here is what we do not control."
Trade-offs & pitfalls
Booting the suspect machine, plugging the SSD into a live OS, or using a consumer-grade cloning tool that passes TRIM through actively accelerates the exact loss you are trying to prevent, this is the single most common and most avoidable mistake. Do not assume SSD behavior mirrors HDD: "deleted but not yet overwritten" is a real, reasoned window on HDD, on SSD it is a much weaker and less predictable one, and both your report and your stakeholder communication should reflect that difference honestly rather than hedge with vague language.
Unlock Full Question Bank
Get access to all 49 Filesystem Forensics and Data Recovery interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.