Security Monitoring, SIEM, and Detection Engineering Questions
Building and operating the detection stack: the SOC and detection-engineering practice that answers 'can we see an attack happening.' Covers SIEM platform selection and architecture, use-case and detection-rule query development (for example Splunk SPL, KQL, or Sigma), alert triage and tuning to reduce false positives, detection engineering and closing coverage gaps, mapping detections to the MITRE ATT&CK framework and scoring detection coverage, log analysis and anomaly and baseline development, network and endpoint telemetry sourcing, malware and compromise-indicator recognition, and security operations center (SOC) alert escalation workflows. Distinct from the reactive work of containing, remediating, and communicating during a confirmed incident (incident response and postmortem topics own that ground; this topic stops at 'the alert fired and here is the detection logic', not 'here is how we contained and recovered from it'). Distinct from generic system-reliability monitoring and observability (SLOs, error budgets, uptime dashboards), a separate discipline even when the underlying ingestion mechanics look similar; the anomaly or signal here must be framed as adversarial or security-relevant. Distinct from hardening a software delivery pipeline against supply-chain compromise (SBOM generation, artifact signing, dependency and build-permission controls); this topic only touches the delivery pipeline from the detection side, spotting a compromised build or tainted artifact via telemetry, not the preventive-controls side. Distinct from designing security control architecture and governance (security architecture and cloud security architecture topics own the design-time question of what controls should exist); this topic is the run-time operation of the detection stack once those controls are in place.
Explain the differences between HIDS (Host-based Intrusion Detection System) and NIDS (Network-based Intrusion Detection System). For each, describe the primary telemetry they consume, their deployment location, strengths and weaknesses, and example detections each is best suited for in a SOC environment.
Sample Answer
Direct answer
Host-based Intrusion Detection System (HIDS) monitors activity ON a single host (its file system, processes, logs, registry) and detects threats that are visible only from inside that host, such as a malicious process or an unauthorized file change. Network-based Intrusion Detection System (NIDS) monitors traffic AS IT CROSSES the network and detects threats visible in transit, such as a scan pattern or a command-and-control (C2) callback. They are complementary, not competing: NIDS sees what moves between hosts, HIDS sees what happens on a host regardless of whether it ever touches a monitored network segment.
Structured elaboration
| HIDS | NIDS | |
|---|---|---|
| Primary telemetry | File integrity, process creation, system/security logs, registry (Windows) or auditd (Linux) | Packet captures, NetFlow, DNS queries, protocol metadata |
| Deployment location | Agent installed on each monitored host | Network tap, SPAN port, or inline sensor at a chokepoint (perimeter, segment boundary) |
| Strengths | Sees encrypted-traffic-invisible activity (in-memory actions, local file changes, process lineage); works even for activity that never leaves the host (local privilege escalation, USB-based compromise) | Broad visibility across many hosts from one sensor; sees lateral movement and exfiltration patterns that no single host's local view would reveal |
| Weaknesses | Requires an agent on every host (deployment/maintenance overhead); if the host is compromised at a privileged level, the agent itself can potentially be tampered with or disabled | Blind to encrypted payload content without a decryption point; blind to activity that never crosses the monitored network segment |
| Best-suited detections | Local privilege escalation, persistence mechanism installation (scheduled tasks, registry run keys), unauthorized file modification, credential-dumping tool execution | Port scanning, beaconing/C2 callback patterns, lateral movement between hosts, DNS tunneling, unusual outbound data volume |
Worked example
Consider a compromised workstation used as a pivot point. The attacker installs a scheduled task for persistence and then attempts to move laterally to a file server over Server Message Block (SMB, the Windows file-sharing protocol).
- HIDS visibility: the scheduled-task creation itself (a local Windows event, ID 4698) is visible to HIDS on that specific workstation the moment it happens, since it never has to leave the host to be observed.
- NIDS visibility: the subsequent SMB connection attempt to the file server crosses the network and is visible to a NIDS sensor monitoring that segment, even if the file server itself has no HIDS agent installed or its local logs were tampered with.
Neither alone gives the full picture: HIDS alone would show the persistence mechanism but might miss the lateral movement if the workstation's own network logging is limited; NIDS alone would show the suspicious SMB connection but not WHY it happened. A SOC using both, correlated in a SIEM, gets the complete chain.
Trade-offs and pitfalls
- Coverage gaps compound: an environment with only NIDS is blind to any purely local activity (a local privilege escalation with no network component); an environment with only HIDS is blind to any host that lacks the agent, which in practice is often unmanaged devices, IoT, or legacy systems that cannot run modern endpoint agents, exactly the hosts an attacker prefers to pivot through.
- Common mistake: relying solely on NIDS in cloud environments under the assumption that "cloud traffic is all visible at the network layer." Modern cloud environments use significant east-west traffic inside virtual networks and encrypted service-to-service communication that a traditional network tap may never see, making host/workload-level telemetry (the cloud-native equivalent of HIDS) equally important.
- Deployment cost asymmetry: NIDS has a lower per-host marginal cost (one sensor can cover many hosts on a segment) but a coverage blind spot for anything off that segment; HIDS has a higher per-host marginal cost (an agent to deploy, update, and monitor on every host) but no such blind spot for the hosts it does cover.
- Agent tampering risk: a sufficiently privileged attacker on a compromised host can potentially disable or blind the HIDS agent itself, which is why HIDS telemetry should be forwarded off-host in near-real-time rather than only stored locally; NIDS, being out-of-band on the network path, is inherently harder for a host-level compromise to directly tamper with.
Define false positives and false negatives in the context of detection rules. Provide two realistic examples of each from SIEM/EDR monitoring and discuss the operational impact (analyst time, missed breaches, alert fatigue) of both error types.
Sample Answer
Direct answer
A false positive is an alert that fires when there is no actual malicious activity; a false negative is malicious activity that occurs without any alert firing at all. They are not symmetric risks: a false positive wastes analyst time and, at scale, erodes trust in the detection pipeline, while a false negative means an attacker operated undetected, which can be catastrophic and is far harder to even notice happened.
Structured elaboration
The confusion-matrix framing: for any given detection rule, an event is either genuinely malicious or genuinely benign, and the rule either fires or does not. A false positive is "fired, but benign." A false negative is "did not fire, but malicious." A true positive is "fired, and malicious," the outcome the rule exists to produce. Tuning a detection rule almost always trades off between these two error types: tightening a rule to reduce false positives typically raises the risk of false negatives (the tighter the match criteria, the easier it is for a slightly different attack variant to slip through), and loosening a rule to catch more true positives typically raises false positives.
Two realistic false-positive examples from SIEM/EDR monitoring:
- A rule alerting on "PowerShell invoked with an encoded command" fires on a legitimate IT automation script that happens to use
-EncodedCommandto pass a complex argument, a common and benign administrative pattern, not just an attacker technique. - A rule alerting on "large outbound data transfer" fires when an employee legitimately uploads a large video file to an approved cloud storage service for a work project.
Two realistic false-negative examples:
- A brute-force detection rule with a fixed threshold of "10 failed logins in 5 minutes" misses a "low-and-slow" attacker deliberately spacing attempts at 1 every 45 seconds, staying just under the threshold indefinitely.
- A signature-based malware detector misses a fileless attack that lives entirely in memory and abuses a legitimate signed system tool, since there is no malicious file on disk for a file-hash-based signature to ever match.
Worked example
Consider the brute-force rule from false-negative example 1 above, with threshold "10 failures in 5 minutes." An attacker sending exactly 1 attempt every 45 seconds produces 6-7 attempts in any 5-minute window, always below the threshold of 10. Over an hour, that is still only about 80 attempts, spread thinly enough that no single 5-minute window ever crosses the line, even though the account is under sustained attack the entire time. This demonstrates the operational impact directly: the false negative here is not a bug in the rule's logic, it is a predictable consequence of a fixed, publicly-inferable threshold, and it means the SOC has zero visibility into an active, ongoing attack against that account.
Trade-offs and pitfalls
- Operational impact of false positives: analyst time is the direct cost (each dismissed alert still takes minutes to review), but the larger cost is alert fatigue, when the false-positive rate is high enough, analysts start pattern-matching "this rule always fires for nothing" and disposition alerts from that rule faster and less carefully, which increases the risk of missing the rare TRUE positive buried among the noise.
- Operational impact of false negatives: the cost is invisible until much later, an undetected compromise continues to escalate (persistence, lateral movement, exfiltration) with no SOC awareness, and the eventual discovery (often via an external notification, a ransomware note, or a downstream anomaly) arrives with a much larger scope to investigate and remediate than if it had been caught early.
- Common mistake: optimizing purely for a low false-positive rate without tracking false negatives at all, because false negatives are much harder to measure (you generally cannot count what you did not detect) so an under-resourced program can easily drift toward "quiet dashboards" that actually reflect blind spots, not real safety.
- Practical mitigation for the false-negative example above: layering a SECOND, complementary detection (a longer-window, lower-threshold rule specifically for slow/distributed brute force, or account-lockout-policy-driven signals) rather than relying on one rule to catch every variant of the same underlying attack pattern.
Write a Sigma rule (YAML-style) that detects suspicious usage of certutil or powershell when invoked with command-line patterns indicating encoding or decoding of content (for example '-EncodedCommand' or 'certutil -decode'). Target Windows ProcessCreate events and include reasonable fields (process_name, command_line, parent_process). Keep the rule generic and explain rationale for key fields.
Sample Answer
Direct answer
Certutil and PowerShell both ship as trusted, signed native Windows binaries, which is exactly why attackers abuse them for encoding/decoding staged payloads (a living-off-the-land technique, MITRE ATT&CK T1027 Obfuscated Files or Information, and T1140 Deobfuscate/Decode Files or Information): the activity blends in with legitimate administrative tool usage rather than requiring the attacker to bring in a separate, more easily-flagged decoder utility.
Structured elaboration
title: Suspicious Certutil or PowerShell Encode/Decode Usage
id: 9e1f2a3b-4c5d-4e6f-8a9b-0c1d2e3f4a5b
status: experimental
description: Detects certutil or PowerShell invoked with command-line patterns
indicating encoding or decoding of content, a common living-off-the-land
technique for staging or deobfuscating a payload without downloading a
separate decoder tool.
logsource:
category: process_creation
product: windows
detection:
selection_certutil:
Image|endswith: '\certutil.exe'
CommandLine|contains:
- '-decode'
- '-encode'
- '/decode'
- '/encode'
selection_powershell:
Image|endswith: '\powershell.exe'
CommandLine|contains:
- '-EncodedCommand'
- 'FromBase64String'
- 'ToBase64String'
filter_known_admin_tooling:
ParentImage|endswith: '\sccm_agent.exe'
condition: (selection_certutil or selection_powershell) and not filter_known_admin_tooling
level: medium
fields:
- Image
- CommandLine
- ParentImage
- User
tags:
- attack.defense_evasion
- attack.t1140
- attack.t1027
Rationale for key fields: Image narrows each selector to the exact binary (certutil.exe or powershell.exe) rather than matching on process NAME alone (which an attacker could rename a different binary to spoof, hence the more specific full-path/endswith match); CommandLine carries the actual evidence of encode/decode intent, since neither binary's mere presence is suspicious on its own, both are legitimate, common administrative tools; ParentImage supports both the false-positive filter and, more broadly, helps an analyst judge whether the invocation chain looks like normal tooling (a management agent) or something else (a document application spawning certutil.exe, for instance, which would be a much stronger standalone signal even without matching this specific rule).
Kept the rule GENERIC, per the question's own instruction, rather than narrowly targeting one specific known malicious command line: the two selectors cover the certutil and PowerShell encode/decode SURFACE broadly (multiple flag spellings, multiple relevant cmdlets/methods) rather than a single exact string, since a rule keyed to one specific observed command line is trivially evaded by the next minor variation an attacker tries.
Worked example
Parsed and converted to SPL using pySigma with the Splunk backend, executed locally:
Output (actual output):
(Image="*\\certutil.exe" CommandLine IN ("*-decode*", "*-encode*", "*/decode*", "*/encode*")) OR (Image="*\\powershell.exe" CommandLine IN ("*-EncodedCommand*", "*FromBase64String*", "*ToBase64String*")) NOT ParentImage="*\\sccm_agent.exe" | table Image,CommandLine,ParentImage,User
The backend doubles each backslash in the Image/ParentImage path literals because SPL treats a single backslash as a string-escape character inside quoted values, so this is correct SPL syntax, not a formatting artifact. This confirms the rule parses as valid Sigma and translates into a runnable SPL search combining both selectors with the false-positive filter applied to either branch, matching the intended (A or B) and not C logic.
Trade-offs and pitfalls
- Certutil's
-decode/-encodeis a genuinely dual-use capability: legitimate IT operations occasionally use certutil for exactly this purpose (base64-encoding/decoding a certificate or small file as part of a manual troubleshooting step), so this rule is deliberately shipped atlevel: mediumrather than high, reflecting a real, non-trivial false-positive rate that should be expected and tuned against actual fired-alert data, following the same evidence-driven approach used for any noisy rule. - Common mistake: matching only the flag spelling an analyst happened to observe in one incident (for example only
-decode) and missing the sibling spellings (/decode,-encode,/encode); Windows utilities frequently accept both dash- and slash-prefixed flags, and a rule that only covers one form is trivially evaded by an attacker (or, just as often, missed against entirely benign administrative usage that happens to use the other form). - This rule catches the INVOCATION pattern only: a fuller detection pipeline would pair this with a follow-up step that actually decodes and re-scans Base64 content flagged by this rule, since the encode/decode ACT is the signal here, not an assessment of what was encoded or decoded.
- MITRE mapping precision: T1140 (Deobfuscate/Decode Files or Information) is the more precise tag for the DECODE direction specifically; T1027 (Obfuscated Files or Information) is the broader parent concept covering obfuscation generally. Tagging both, as this rule does, is defensible since the rule covers both encode and decode invocations, but a rule scoped to decode-only activity specifically would be more precisely tagged with T1140 alone.
System design (medium): Design a detection coverage matrix that maps existing detections to MITRE ATT&CK techniques for a mid-size organization. Explain how you would populate the matrix, calculate a coverage score per tactic/technique, prioritize gaps for remediation, and what dashboards or KPIs you'd provide to leadership to show progress over time.
Sample Answer
Direct answer
A detection coverage matrix maps every MITRE ATT&CK technique relevant to the organization's threat model against the detections that actually cover it, scored by confidence, not just presence; populated honestly and validated (not just self-reported), it turns "are we covered" from a vague sense into a specific, prioritizable list of gaps, and weighting by asset criticality is what keeps that prioritization pointed at what actually matters rather than at whatever gap happens to look largest on paper.
Structured elaboration
Populating the matrix: start from the subset of the full ATT&CK matrix genuinely relevant to the organization (informed by its own threat model, industry, and known attacker interest, not the entire framework indiscriminately); for each relevant technique, record every mapped detection rule, and score each mapping's CONFIDENCE (not just its existence): a validated detection (confirmed via red-team or purple-team testing to actually fire against a realistic instance of the technique) scores highest; a rule mapped but never validated scores lower; a technique with no mapped detection at all scores zero.
Calculating a coverage score per tactic/technique: a simple per-technique score (0 for no coverage, a partial value for unvalidated coverage, 1 for validated coverage) rolls up to a per-tactic average, but a RAW, unweighted average treats a gap in a low-value technique identically to a gap in a technique targeting the organization's most critical assets, which is rarely the right prioritization signal.
Weighting by asset criticality: assign each technique a weight reflecting how much it threatens the organization's actual crown-jewel assets (informed by which techniques are typically used against the specific systems/data the organization most needs to protect), and compute a criticality-WEIGHTED coverage score alongside the raw average, since the two can diverge meaningfully and the weighted number is the one that should actually drive prioritization.
Prioritizing gaps for remediation: rank uncovered or weakly-covered techniques by their weight (criticality) first, then by how central that technique is to a realistic attack path the organization is likely to face, rather than by raw technique count or alphabetical/ID order.
Dashboards/KPIs for leadership: an overall weighted coverage percentage and its trend over time (the single number leadership will most often ask for); a breakdown by tactic (showing whether coverage is unevenly distributed, strong on some tactics and weak on others); and a short, named list of the highest-weighted current gaps with a remediation timeline, giving leadership something concrete and actionable rather than an abstract score alone.
Worked example
A simplified 5-technique matrix (illustrative subset, a real matrix would span the organization's full relevant technique set), each scored 0 (no coverage), 0.5 (partial/unvalidated), or 1 (validated), with a criticality weight reflecting how central that technique is to the organization's highest-value assets:
| Technique | Coverage score | Criticality weight |
|---|---|---|
| T1078 Valid Accounts | 1.0 | 3 |
| T1110.001 Password Guessing | 0.0 | 3 |
| T1558.003 Kerberoasting | 0.5 | 1 |
| T1550.002 Pass the Hash | 1.0 | 1 |
| T1003.001 LSASS Memory | 0.0 | 2 |
Raw average coverage: (1.0+0.0+0.5+1.0+0.0)/5=0.5, 50%. Criticality-weighted coverage: 3+3+1+1+21.0×3+0.0×3+0.5×1+1.0×1+0.0×2=104.5=0.45, 45%.
The two numbers tell a materially different prioritization story: a raw 50% might read as "roughly half covered, no urgent single gap," but the weighted 45% and, more importantly, the WEIGHT DISTRIBUTION behind it point specifically at T1110.001 (Password Guessing), a zero-coverage technique carrying the SAME highest weight (3) as the organization's best-covered technique. That is the gap a leadership dashboard should surface as the top remediation priority, not because it is alphabetically first or numerically largest as a raw count, but because it combines total absence of coverage with the organization's own stated highest asset-criticality weighting.
Trade-offs and pitfalls
- Common mistake: scoring coverage as a binary (mapped or not mapped) rather than by validated confidence; a rule that is mapped to a technique but has never been tested against a realistic instance of it can create false confidence identical in the matrix to a rule that has been rigorously validated.
- Weighting itself needs a defensible, documented basis: an arbitrary or intuition-only weighting scheme is vulnerable to the same bias it is meant to correct for; weights should be traceable to a real criticality assessment (which assets the organization has already identified as highest-value) rather than assigned ad hoc by whoever built the matrix.
- Common mistake: presenting only the single overall percentage to leadership without the per-tactic breakdown or the named gap list; a single number invites a false sense of either complacency ("we're at 80%, good enough") or alarm ("we're at 45%, we're in trouble") without the context needed to act on it specifically.
- The matrix needs periodic revalidation, not just periodic re-scoring: an environment change (a new tool deployed, a detection quietly breaking due to a schema change upstream) can silently invalidate a previously-validated mapping without the coverage score itself ever being manually updated to reflect that; a stale, unrevalidated "validated" score is arguably worse than an honest "unvalidated" one, since it hides the actual gap.
Given a relational table login_attempts(user_id TEXT, timestamp TIMESTAMP, status TEXT, source_ip TEXT), write a standard SQL query that returns user_id and source_ip pairs which had 5 or more failed login attempts within any five-minute window. State assumptions about timestamp precision and indexing.
Sample Answer
Direct answer
Finding (user_id, source_ip) pairs with 5 or more failed logins in any 5-minute WINDOW, not a fixed calendar bucket, is a self-join problem: join the table against itself on matching user and IP, count how many failures fall within 5 minutes trailing each individual failed attempt, and keep only the pairs where that trailing count reaches the threshold.
Structured elaboration
SELECT a.user_id, a.source_ip, COUNT(*) AS fail_count_in_window
FROM login_attempts a
JOIN login_attempts b
ON a.user_id = b.user_id
AND a.source_ip = b.source_ip
AND b.status = 'fail'
AND b.timestamp BETWEEN datetime(a.timestamp, '-5 minutes') AND a.timestamp
WHERE a.status = 'fail'
GROUP BY a.user_id, a.source_ip, a.timestamp
HAVING COUNT(*) >= 5
ORDER BY a.user_id, a.timestamp;
Assumptions about timestamp precision and indexing: timestamp is stored with at least second-level precision in a format the database can compare and offset arithmetically (ISO 8601 text works directly in SQLite via the datetime() function used above; a genuine TIMESTAMP column type in Postgres/MySQL would use the equivalent native interval arithmetic instead of datetime(..., '-5 minutes')). For this self-join to perform acceptably at real data volume, a composite index on (user_id, source_ip, timestamp) is essential, since the join condition filters on exactly those three columns together, without it, the database would need a full scan for every row's window comparison, an O(n2)-shaped cost at scale that the index converts into a much cheaper indexed range lookup per row.
Worked example
Executed against a real, populated in-memory SQLite database with four constructed users, each testing a different edge of the specification:
import sqlite3
conn = sqlite3.connect(":memory:")
cur = conn.cursor()
cur.execute("CREATE TABLE login_attempts (user_id TEXT, timestamp TEXT, status TEXT, source_ip TEXT)")
rows = [
# alice: 5 failures within 4 minutes from ONE IP -> should be flagged
("alice", "2026-07-30 09:00:00", "fail", "203.0.113.5"),
("alice", "2026-07-30 09:01:00", "fail", "203.0.113.5"),
("alice", "2026-07-30 09:02:00", "fail", "203.0.113.5"),
("alice", "2026-07-30 09:03:00", "fail", "203.0.113.5"),
("alice", "2026-07-30 09:04:00", "fail", "203.0.113.5"),
# bob: only 4 failures within 5 minutes -> below threshold
("bob", "2026-07-30 09:10:00", "fail", "198.51.100.9"),
("bob", "2026-07-30 09:11:00", "fail", "198.51.100.9"),
("bob", "2026-07-30 09:12:00", "fail", "198.51.100.9"),
("bob", "2026-07-30 09:13:00", "fail", "198.51.100.9"),
# carol: 5 failures but spread across 20 minutes -> never 5 within any single 5-min window
("carol", "2026-07-30 09:20:00", "fail", "192.0.2.7"),
("carol", "2026-07-30 09:25:00", "fail", "192.0.2.7"),
("carol", "2026-07-30 09:30:00", "fail", "192.0.2.7"),
("carol", "2026-07-30 09:35:00", "fail", "192.0.2.7"),
("carol", "2026-07-30 09:40:00", "fail", "192.0.2.7"),
# dave: 5 failures total but split across TWO different source IPs, 3+2, neither IP alone reaches 5
("dave", "2026-07-30 09:50:00", "fail", "10.0.0.1"),
("dave", "2026-07-30 09:50:30", "fail", "10.0.0.1"),
("dave", "2026-07-30 09:51:00", "fail", "10.0.0.1"),
("dave", "2026-07-30 09:51:30", "fail", "10.0.0.2"),
("dave", "2026-07-30 09:52:00", "fail", "10.0.0.2"),
]
cur.executemany("INSERT INTO login_attempts VALUES (?, ?, ?, ?)", rows)
conn.commit()
# ... (query as shown above) ...
print(sorted({(r[0], r[1]) for r in cur.execute(query).fetchall()}))
Output (actually executed with python3's built-in sqlite3):
[('alice', '203.0.113.5')]
Only alice's (user_id, source_ip) pair is flagged. Bob's 4 failures correctly stay below the >= 5 threshold. Carol's 5 failures, spread across 20 minutes, correctly never accumulate 5 within any single trailing 5-minute window, confirming the query genuinely enforces a SLIDING window, not just a total-count-within-the-data-range check. Dave's 5 total failures, split 3-and-2 across two different source IPs, correctly stay unflagged for EITHER individual (user_id, source_ip) pair, confirming the join's source_ip equality condition genuinely scopes the count per-IP, not per-user-across-all-IPs.
Trade-offs and pitfalls
- The Carol and Dave test cases are the two edge cases most likely to reveal a subtly wrong query, and both are directly, deliberately exercised above: a query that mistakenly counted failures within a fixed CALENDAR bucket (rather than a genuinely sliding window anchored to each individual attempt) would have failed Carol's case in one direction or another depending on bucket alignment; a query that grouped only by
user_id(forgetting to also join and group onsource_ip) would have incorrectly flagged Dave, since his 5 total failures across two IPs would appear to cross the threshold if IP were ignored. - Common mistake: using a plain
GROUP BYwith a fixed time-bucket function (like truncating to the nearest 5-minute mark) instead of the self-join sliding-window approach; a fixed-bucket approach can miss a genuine cluster of 5 failures that happens to straddle a bucket boundary (2 failures in one bucket, 3 in the next, never grouped together despite genuinely occurring within 5 minutes of each other), exactly the kind of boundary artifact a sliding self-join avoids by anchoring the window to each individual row rather than a fixed grid. - This self-join pattern's cost grows with the SQUARE of matching rows for a very chatty (user, IP) pair without the composite index named above: at real production volume, this specific query shape is exactly why the index is not an optional performance nicety but a genuine correctness-adjacent requirement, an unindexed version of this query could time out or degrade the whole database's performance under real load, a materially worse outcome than merely running slowly.
- Output includes the RAW per-attempt window counts, not just the distinct flagged pairs: a caller wanting just the pairs (as shown in the worked example's final printed result) should deduplicate the raw output, since the base query as written returns one row per QUALIFYING trailing window, which can include multiple rows for the same pair if the pattern persists across several consecutive attempts.
Unlock Full Question Bank
Get access to all Security Monitoring, SIEM, and Detection Engineering interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.