Log Analysis and Diagnostic Data Gathering Questions
Extracting signal from existing logs and diagnostic output to find a root cause: parsing and querying log data, correlating traces and metrics during an investigation, and gathering the right diagnostic information (including asking clarifying questions) before drawing conclusions. Covers text-processing and query techniques for locating evidence in logs (structured log parsing, ElasticSearch/SQL-style log queries, log aggregation and retention trade-offs) and reconstructing a timeline from the data on hand. This is the analysis-of-existing-data skill used during troubleshooting and investigation across infrastructure and operations roles: distinct from monitoring and observability, which is about instrumenting a system so telemetry exists in the first place (see the observability topics for that), and distinct from SIEM-based security detection and formal digital-forensics practice (chain of custody, artifact/disk/memory analysis), which have their own dedicated coverage elsewhere in the catalog.
Compare different approaches to log aggregation and retention for forensic investigations at enterprise scale: self-hosted ELK, managed cloud logging, and cold object-store. Discuss trade-offs in cost, query latency, retention, compliance, and operational burden for each approach.
Sample Answer
Direct answer
Self-hosted ELK (Elasticsearch, Logstash, Kibana, the classic open-source log-search stack), managed cloud logging, and cold object storage sit at three different points on the same cost-versus-query-latency curve for forensic retention: cheaper, colder, and slower to query as you move from ELK toward object storage, with retention economics and operational burden moving in the opposite direction.
Structured elaboration
| Dimension | Self-hosted ELK | Managed cloud logging | Cold object storage (e.g. an S3-style bucket) |
|---|---|---|---|
| Cost | High fixed cost: infrastructure plus the ops time to run it | Usage-based (ingest and retention volume); shifts cost from headcount to a recurring bill | Lowest cost per gigabyte by a wide margin; the cheapest option for years of retention |
| Query latency | Fast (sub-second to seconds) if sized and indexed correctly | Fast for recent data; heavy ad hoc queries may be rate-limited or cost extra | Slow; minutes, not seconds, unless you've pre-indexed or partitioned the data, since it's not a search engine |
| Retention | Gets expensive fast at scale, since indexed storage costs far more than raw storage | Vendor-tiered; long retention is available but priced accordingly | Cheap and practical for multi-year retention; built for exactly this |
| Compliance | Full control, but you implement immutability and access controls yourself | Strong built-in controls, but possible data-residency constraints depending on vendor regions | Good: object-lock/WORM (write-once-read-many) style immutability features exist, but you still design the access-control and integrity story |
| Operational burden | High: cluster sizing, upgrades, index lifecycle management, backups | Low: the vendor runs the platform; you manage ingestion and IAM (identity and access management) policy | Low for storage itself, moderate for building a reliable ingestion and later query path |
Worked example
A forensic investigation needs to pull specific evidence from 18 months ago. Against cold object storage alone with no pre-built index, that query might take many minutes to hours, since you're effectively scanning partitioned files rather than hitting a search index; the same query against a live ELK cluster holding that same 18 months of data would return in seconds, but only if the team was willing to pay to keep 18 months indexed and hot the whole time, which for most teams is the expensive part of this trade-off, not the storage itself.
Trade-offs & pitfalls
The practical answer for most teams doing forensic investigations at enterprise scale isn't picking exactly one of these three, it's a tiered pattern: keep a shorter, indexed hot/warm window (days to a few weeks) in ELK or a managed logging product for fast, ad hoc investigation, then move older data into cheap, partitioned cold object storage for long-term compliance retention, paired with a serverless or batch query layer for the rare occasions you actually need to search far back. The pitfall to watch for is treating cold storage as a drop-in replacement for a search index: teams sometimes archive to object storage to save cost and only discover during an actual forensic investigation, under time pressure, that "search 18 months of raw files" takes far longer than anyone budgeted for, because nobody tested the retrieval path until it was needed for real.
Write a Logstash/ELK grok pattern (or equivalent) for Nginx 'combined' access logs to extract client_ip, timestamp, method, path, protocol, status, bytes_sent, and user_agent. Explain how you'd handle query strings in path, percent-encoding, and very long user-agent strings to avoid mapping explosion in Elasticsearch.
Sample Answer
Direct answer
Grok is the pattern-matching DSL (domain-specific language, a small purpose-built syntax rather than a general programming language) that Logstash's grok filter uses to name regex fragments and compose them, so you write field names instead of raw regex. The Logstash grok pattern below is illustrative configuration (not something this environment can execute against a real Logstash pipeline); the field boundaries it encodes are verified separately with an equivalent, executed Python regex against a real nginx combined log line.
Approach
Logstash grok pattern
%{IPORHOST:client_ip} - %{USER:ident} \[%{HTTPDATE:timestamp}\] "%{WORD:method} %{DATA:path} %{DATA:protocol}" %{NUMBER:status} %{NUMBER:bytes_sent} "%{DATA:referrer}" "%{DATA:user_agent}"
Each %{PATTERN:field_name} is a named, reusable regex fragment from grok's built-in pattern library: %{IPORHOST} matches an IP address or hostname, %{USER} matches a username token, %{HTTPDATE} matches Apache's bracketed date format, %{WORD} matches one alphabetic token, and %{NUMBER} matches a numeric string, each expanding to a pre-built regex fragment so you don't have to write it by hand. %{DATA} is grok's non-greedy "anything" pattern, used here for path, protocol, referrer, and user_agent because those fields can contain almost any character. Note that grok's own built-in COMBINEDAPACHELOG pattern names the trailing HTTP/1.1 token httpversion and only captures the version number, not the scheme; since the question asks specifically for a field named protocol, this pattern captures the whole HTTP/1.1 token (scheme and version together) under that name instead, which is the more literal reading of "protocol."
Verifying the field boundaries (executed)
import re
# Same field boundaries as the grok pattern
# %{IPORHOST:client_ip} - %{USER:ident} \[%{HTTPDATE:timestamp}\]
# "%{WORD:method} %{DATA:path} %{DATA:protocol}"
# %{NUMBER:status} %{NUMBER:bytes_sent} "%{DATA:referrer}" "%{DATA:user_agent}"
PATTERN = re.compile(
r'^(?P<client_ip>\S+) - (?P<ident>\S+) '
r'\[(?P<timestamp>[^\]]+)\] '
r'"(?P<method>[A-Z]+) (?P<path>\S+) (?P<protocol>[A-Z]+/[\d.]+)" '
r'(?P<status>\d+) (?P<bytes_sent>\d+|-) '
r'"(?P<referrer>[^"]*)" '
r'"(?P<user_agent>[^"]*)"$'
)
line = ('203.0.113.5 - - [10/Oct/2024:13:55:36 -0700] '
'"GET /search?q=test%20query HTTP/1.1" 200 5231 '
'"https://example.com/" '
'"Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36"')
m = PATTERN.match(line)
if m:
for k, v in m.groupdict().items():
print(f"{k}: {v}")
else:
print("no match")
Worked example (executed output)
client_ip: 203.0.113.5
ident: -
timestamp: 10/Oct/2024:13:55:36 -0700
method: GET
path: /search?q=test%20query
protocol: HTTP/1.1
status: 200
bytes_sent: 5231
referrer: https://example.com/
user_agent: Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36
Key points
- Query strings in the path:
%{DATA:path}is non-greedy, so it stops at the first space rather than trying to consume the rest of the line; a query string like?q=test%20querystays part of thepathfield intact, exactly as shown above (/search?q=test%20query). If you needpathandqueryas separate fields, add adissectormutatestep after the grok match to split on the first?. - Percent-encoding: neither grok nor this regex decodes
%20into a space; that's expected; percent-encoding is part of the raw URL and decoding it is a separate, deliberate step (only do it if you actually need the decoded form for display or analysis, since the encoded form is what the server actually received). - Very long user agents and mapping explosion: in Elasticsearch, a
textor default-mappedkeywordfield indexes every unique value it sees; browser and bot user-agent strings are extremely high-cardinality (having a very large number of distinct values) and some are very long, which can blow up index size and even hit Elasticsearch's field-length limits. The practical mitigations are: mapuser_agentaskeywordwithignore_aboveset (so overlong values are stored but not indexed for aggregation), truncate the field before indexing if you only need it for display, and, if you need to group by user agent at all, index a hash of the full string instead of the raw string for aggregation purposes. - Naming
protocolvs. a library's default field name: this is a small but real pitfall in its own right. A team pulling in a stockCOMBINEDAPACHELOG-style pattern will gethttpversion, notprotocol, and a downstream query or dashboard written against "protocol" will silently return nothing until someone notices the field doesn't exist. When a spec (a question, a ticket, a schema doc) names a field explicitly, match that name exactly, or add an explicit rename/alias step, rather than assuming a library's default naming is close enough.
Trade-offs & pitfalls
The main pitfall with grok specifically is that a malformed line (one that doesn't match the pattern at all) produces no match and typically gets tagged as a parse failure rather than partially parsed; a production pipeline needs a fallback path for those lines rather than silently dropping them. The other common mistake is treating query strings as low-cardinality: a path with a highly variable query string (a session token, a timestamp) indexed as a single keyword field will itself cause the same kind of cardinality blowup as an unbounded user-agent field, so the same truncate-or-split discipline usually needs to apply to path as well, not just to user_agent.
Your log-processing pipeline needs to keep up with roughly 200k JSON log lines per second per host, and profiling shows the parsing step itself is the bottleneck, not disk or network I/O. Walk through how you'd diagnose where the time is actually going, what class of tooling you'd consider moving to if a scripting-language parser can't keep up, and how you'd benchmark candidate approaches before committing to a rewrite.
Sample Answer
Direct answer
Before touching the language or the parser, profile the parsing step itself with a sampling profiler to find out whether the cost is really JSON decoding, or something adjacent like object/dict allocation, string encoding conversions, or field extraction logic sitting next to the parse call. Only after that's confirmed as the actual hot path is a rewrite (to a compiled language, or to a SIMD-accelerated JSON parser) worth considering, and even then, the cheapest fix is often to see whether the shell-tooling tier (grep/awk/sed) can filter or pre-shape the data so less JSON ever needs full parsing at all.
Structured elaboration
1. Diagnose where the time is actually going. "Parsing is the bottleneck" from a coarse profile (CPU is pegged, I/O is idle) still leaves several distinct possibilities:
- Use a sampling profiler (
py-spyfor Python,pproffor Go,perfat the OS level) to get a flame graph (a visualization showing which function calls consumed the most CPU time, stacked by call depth, so the widest bars are the hottest code paths) and see whether time is in the actual decode call (e.g.json.loads) or in what your code does with the result (building dict-of-dicts, converting to your internal schema, string encode/decode churn). - Check allocation behavior separately from CPU time (Python's
tracemalloc, Go'spprof -alloc_objects): a lot of "parsing is slow" turns out to be garbage-collector pressure from allocating a new dict and set of strings per line, not the byte-level scanning itself. - Isolate parsing from I/O properly: feed the parser from an in-memory buffer or
/dev/shmfor the profiling run so disk/network jitter can't leak into the measurement, confirming the earlier finding that I/O isn't the bottleneck.
2. The cheaper step before a rewrite. A full language rewrite is a multi-week investment with real risk; before reaching for it, check whether the shell-tooling tier can do enough of the job:
grep/mawk/sedare compiled C programs with tight, allocation-light loops. If most lines can be filtered out (e.g. you only care aboutlevel=ERRORlines, or only a handful of fields) before they ever reach a JSON parser, agrep-then-parse pipeline can cut the number of lines that pay the full JSON-decode cost by an order of magnitude for free.- If you need real field extraction rather than just filtering, tools like
jqor aawk-based fixed-position extractor can pull specific fields without a general-purpose JSON parse, at the cost of being fragile to schema changes. - This step matters because it changes the actual problem: "parse 200k lines/sec" might become "parse the 20k lines/sec that survived filtering," which a scripting-language parser may already handle fine, no rewrite needed at all.
3. If a rewrite is still justified, what class of tooling. Once you've confirmed the decode itself (not allocation, not I/O, not unnecessary work) is the bottleneck, and pre-filtering doesn't reduce the problem enough, the realistic options are:
- A compiled, garbage-collected language (Go): a large step up in throughput with a manageable operational profile; still has GC pauses under sustained allocation, which pooling (reusing buffers/structs via a pool instead of allocating fresh ones per line) mitigates.
- A systems language with manual memory control (Rust): highest achievable throughput and predictable latency (no GC pauses), at the cost of a steeper team ramp-up.
- A SIMD-accelerated JSON parser (e.g. simdjson and its bindings): uses CPU vector instructions (single instruction, multiple data, meaning one instruction processes several bytes in parallel) to scan JSON far faster than a byte-at-a-time parser; typically wrapped in a small Go or Rust service since simdjson itself is a C++ library.
The right choice depends on how much of the rest of your pipeline is already in that language: a Go rewrite that reuses your existing Go services is a smaller organizational lift than introducing Rust or C++ purely for this one hot path.
4. Benchmarking before committing. A microbenchmark that doesn't resemble production traffic will mislead you:
- Build the benchmark corpus from real captured log samples, not synthetic data, including your actual field cardinality, nesting depth, and line-length distribution, not just an average-case line repeated a million times.
- Measure more than one thing: lines/sec throughput, but also p50/p95/p99 latency per line (a parser that's fast on average but spikes on deeply-nested lines will still cause tail-latency problems downstream), and memory allocation rate.
- Run sustained (multi-minute) trials, not single-shot timings; JIT (just-in-time compilation, where the runtime compiles and speeds up hot code paths only after a warm-up period)/GC-based runtimes behave very differently warm vs. cold.
- Shadow the candidate against real production traffic (processing a copy of the live stream without acting on the output) before cutting over, so the decision is validated against real, current data shapes rather than the benchmark corpus alone.
I'm not going to state specific throughput numbers here (lines/sec, milliseconds) because they're entirely dependent on your hardware, log shapes, and even background load at benchmark time; the value is in the benchmark harness's shape (representative data, percentile tracking, sustained runs), which is what you'd actually run to get numbers you can trust for your own environment. A benchmark skeleton looks like this:
import time, statistics
def benchmark(parse_fn, corpus, warmup=5000):
for line in corpus[:warmup]:
parse_fn(line) # warm up JIT/caches, discard timing
latencies = []
start = time.perf_counter()
for line in corpus[warmup:]:
t0 = time.perf_counter()
parse_fn(line)
latencies.append(time.perf_counter() - t0)
total = time.perf_counter() - start
return {
"lines_per_sec": len(corpus[warmup:]) / total,
"p50_us": statistics.median(latencies) * 1e6,
"p99_us": sorted(latencies)[int(len(latencies) * 0.99)] * 1e6,
}
Run this on your own captured corpus against each candidate parser to get numbers specific to your hardware and data, that comparison is the deliverable, not any single absolute number.
Trade-offs & pitfalls
- The most common mistake here is skipping straight to "rewrite in Rust" because 200k/sec sounds intimidating, without ever confirming that decode itself (not allocation, not unnecessary downstream work, not I/O) is actually the cost. That can burn weeks rewriting the wrong bottleneck.
- simdjson-class parsers are fastest at raw scanning but usually parse into a DOM-like view (a tree of in-memory objects you then navigate to pull out fields, the same idea as the DOM a browser builds from parsed HTML) you then have to extract fields from; if your real cost is field extraction and schema mapping, a faster raw parser alone won't fix it.
- A rewrite that's faster in isolation but harder to operate (a language your team doesn't run in production elsewhere) can be a net loss even if the benchmark numbers look good; factor operational cost into the decision, not just throughput.
During an incident retro, you go looking for the application log from four days ago and it's gone; only the last couple of days of rotated files still exist. Walk through how you'd figure out whether that's expected retention behavior or a rotation misconfiguration, and what you'd check or change so it doesn't bite the next investigation.
Sample Answer
Direct answer
Start from the config: read the logrotate (the standard Linux utility that rotates, compresses, and deletes log files on a schedule) stanza for this application and work out how many days of history it's actually supposed to retain, since "only 2 days survive" might be exactly what the config says, just not what anyone expected. If the config genuinely promises more than 2 days, the next question is whether rotation is happening more often than assumed, most commonly because it's triggered by file size, not purely by a daily schedule, and a burst of logging quietly burned through the retention budget faster than usual.
Structured elaboration
- Read the actual config, don't assume it says what people remember. Find the stanza for this app (typically under
/etc/logrotate.d/) and checkrotate N(how many old copies to keep before deleting), and whether rotation isdaily/weeklyorsize-triggered (e.g.size 100M). A config withrotate 14genuinely only promises 14 rotation cycles, not 14 calendar days, those are the same thing only if rotation is purely time-based and never fires early. - Check whether rotation is size-triggered and could have fired more than once a day. This is the most common surprise: if the stanza has a
sizedirective (ormaxsizealongsidedaily), a burst of unusually verbose logging (a bug spamming warnings, a traffic spike) can trigger multiple rotations in a single day. If that happened,rotate 14might only cover 3-4 calendar days instead of 14, and the config was never wrong, the assumption that "rotation count equals days" was. - Check when rotation actually last ran and how often. Logrotate records the last rotation timestamp per config stanza in a status file (commonly
/var/lib/logrotate/statuson newer Debian/Ubuntu packaging, or/var/lib/logrotate.statuson some older systems, check whichever exists on this host). Comparing that against the timestamps on the surviving rotated files (ls -la --time-style=full-iso /var/log/myapp/*.gz) tells you the actual rotation cadence over the retention window, not the cadence the config implies. - Rule out something deleting files outside logrotate entirely. A disk-pressure cleanup cron job, a container image rebuild that wiped a non-persistent volume, or a manual cleanup someone ran and forgot about can all produce the exact same symptom as a rotation misconfiguration. Check for other scheduled jobs touching that directory (
grep -r logrotate /etc/cron*, and anything else withfind/rmagainst/var/log/myapp) before concluding it's purely alogrotateissue. - Dry-run the config to see what it would actually do right now.
logrotate -d /etc/logrotate.d/myappruns a dry run (debug mode: shows what actions would be taken without taking them), which is the fastest way to confirm your reading of the stanza matches what logrotate itself thinks it should do, catching a typo or an unexpected included/overriding config you missed on a manual read.
Worked example
# 1. What does the config actually promise?
cat /etc/logrotate.d/myapp
# e.g.:
# /var/log/myapp/application.log {
# daily
# rotate 14
# size 100M
# compress
# }
# This promises "up to 14 rotations, whichever of daily/100M triggers first"
# -- NOT unconditionally 14 calendar days.
# 2. When did each surviving rotation actually happen, and how big were they?
ls -la --time-style=full-iso /var/log/myapp/application.log*.gz
# 3. What does logrotate's own status file say about last-rotation timing?
cat /var/lib/logrotate/status 2>/dev/null || cat /var/lib/logrotate.status 2>/dev/null
# 4. Confirm the config parses the way you think it does, without applying it.
logrotate -d /etc/logrotate.d/myapp
If step 2's timestamps show several rotations within a single day around the time of a known traffic spike or noisy deploy, that's a strong, checkable signal that size 100M fired early and repeatedly, burning through the 14-rotation budget in far fewer than 14 days, exactly the "expected behavior, unexpected outcome" case rather than a misconfiguration.
What to change so it doesn't bite the next investigation
- If size-triggered rotation is the cause, either raise
rotate Nenough to cover your actual worst-case rotation cadence during a noisy period, or addmaxagesemantics via a longer retention window, or (often the real fix) ship logs off-host to a central store where retention isn't coupled to a single host's disk and a burst of local logging can't shrink your effective history. - If the config is genuinely fine and the real gap is "logrotate's local retention was never meant to serve as your incident-investigation retention," that's a signal the two need to be decoupled: local rotation exists to bound disk usage, not to be the system of record for four-day-old incident evidence.
- Whatever the root cause, document the actual effective retention (in calendar days, under realistic worst-case log volume) somewhere a future investigator will find it before they hit the same surprise.
Trade-offs & pitfalls
rotate Nwith a size trigger is a genuinely reasonable config for bounding disk usage; the mistake isn't the config, it's treating "N rotations" and "N days" as interchangeable when planning how far back an investigation can reach.- Raising
rotate Njust enough to cover today's worst case doesn't protect against tomorrow's worse case (a bigger traffic spike, a noisier bug); if four-day-plus retention actually matters for incident response, local rotation on a single host is the wrong system to depend on at all.
On a typical Linux system, where are system and application logs stored by default and what are the differences between text-based files under /var/log and the systemd journal? Include example commands to:
- View the last 100 lines of a file-based log
- Show logs for a specific systemd unit since yesterday
- Follow a log file in real time
Explain when one source may contain entries the other does not and the implications for incident response.
Sample Answer
Direct answer
/var/log holds traditional, plain-text log files (one file per service or subsystem, rotated by logrotate). The systemd journal is a separate, structured binary log store, written and queried through journald/journalctl, that many modern distributions use alongside or instead of writing directly to /var/log. They can disagree about what happened during an incident, so a thorough investigation checks both rather than assuming one is a superset of the other.
Structured elaboration
- Format:
/var/logfiles are plain text, so any text tool (grep,tail,awk) works directly. The journal stores structured fields (timestamp, unit, PID, priority, and more) in a binary format that requiresjournalctlto read. - Coverage: services that write to standard output/error under systemd have their output captured by the journal automatically. A service that opens its own file and writes to it directly (a legacy daemon with hardcoded file logging) may never reach the journal at all unless something explicitly forwards it.
- Retention:
/var/logfiles are rotated bylogrotateon a schedule and size/count policy you configure. The journal has its own separate size/time limits (SystemMaxUse,SystemMaxFileSize, and similar settings injournald.conf), which can differ from yourlogrotatepolicy and quietly expire data on a different schedule than you expect. - Early boot: the journal can capture kernel and early-boot messages before user-space logging daemons have even started, something a
/var/logfile typically cannot do, since the process that would write it isn't running yet.
Worked example
# View the last 100 lines of a file-based log
tail -n 100 /var/log/syslog
# Show logs for a specific systemd unit since yesterday
journalctl -u nginx.service --since yesterday
# Follow a log file in real time
tail -f /var/log/nginx/access.log
Suppose nginx.service fails to start at all because of a bad config file. Systemd's own record of that failure (the unit failing to activate) lives in the journal, associated with the unit; nginx itself never got far enough to write anything into /var/log/nginx/error.log, since the process never fully started. Grepping only /var/log/nginx/error.log here would find nothing, while journalctl -u nginx.service --since yesterday would show the startup failure.
Trade-offs & pitfalls
The common mistake during incident response is picking one source and treating it as complete. Relying only on /var/log risks missing early-boot events, systemd-level failures, and anything a service only ever sent to standard output. Relying only on the journal risks missing legacy daemons that write directly to files and were never wired into journald, or log lines that already rotated out of the journal's own retention window while an older /var/log archive still has them. When the two disagree about whether something happened, that disagreement is itself a useful signal about how the service is (or isn't) integrated with systemd, not just an annoyance to work around.
Unlock Full Question Bank
Get access to all 28 Log Analysis and Diagnostic Data Gathering interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.