# Where to Learn Applied ML for Incident Response: Start at Scoping

By Curtis Lambert · 2026-10-07

> Applied ML for incident response pays off in scoping: finding each host that looks like patient zero. What to practice, which datasets to use, and where to learn it.

The incident has one confirmed host. The CISO's first question is not "how did they get in." It is "how many more are there," and every hour you spend answering it is an hour the attacker keeps their access. Scoping is where machine learning earns its place in incident response, and it is the skill most training skips.

Most "ML for IR" material teaches classifiers: label malware, label phishing. Responders rarely need a classifier mid-incident. They need two answers fast: which hosts could the attacker have reached, and which hosts are behaving like the one we know is bad. Both are a few dozen lines of Python on logs you already collect.

## Question one: who could they have reached?

Lateral movement over SMB, RDP, or WMI ([T1021.002](/mitre/T1021.002), [T1021.001](/mitre/T1021.001), [T1047](/mitre/T1047)) with valid accounts ([T1078](/mitre/T1078)) leaves Windows Event ID 4624 on the destination host. Logon types 3 (network) and 10 (RemoteInteractive) are the ones to keep. Treat each logon as a directed, timestamped edge from source host to destination host. A host is in scope only if a path reaches it from patient zero *in time order*: a logon to host B at 01:00 cannot carry an attacker who reached host A at 02:10.

```python
import pandas as pd

logons = pd.read_json("4624.jsonl", lines=True)
logons["ts"] = pd.to_datetime(logons["TimeCreated"], utc=True)
logons = logons[logons["LogonType"].isin([3, 10])
                & ~logons["TargetUserName"].str.endswith("$", na=False)]

# IpAddress is the source; Computer is the host that logged the event.
# ip_to_host comes from DHCP leases or DNS for the incident window.
logons["src"] = logons["IpAddress"].map(ip_to_host)
edges = (logons.dropna(subset=["src"])
               .query("src != Computer")
               .sort_values("ts"))

def reachable(edges, patient_zero, t0):
    reached = {patient_zero: t0}
    for e in edges[edges["ts"] >= t0].itertuples():
        if e.src in reached and e.Computer not in reached:
            reached[e.Computer] = e.ts
    return pd.Series(reached, name="earliest_possible").sort_values()

t0 = pd.Timestamp("2026-09-14 02:10", tz="UTC")
in_reach = reachable(edges, "WS-0412", t0)
```

Because the edges are processed in time order, a single pass gives every host's earliest possible compromise time. Field names depend on your export (Winlogbeat nests them under `winlog.event_data`), so adjust the column names.

Run it unfiltered and the answer is usually "everything within an hour." A backup server or a vulnerability scanner authenticates to the whole fleet, and an attacker who reaches one of those inherits all its paths. That result is honest, but it is useless for triage. Restrict `edges` to the accounts you have evidence were compromised, rerun, and the set shrinks to something you can work. The output is an upper bound on scope, not a list of findings.

## Question two: who looks like patient zero?

Reachability says who *could* be compromised. Behavior says who probably *is*. Sysmon Event ID 1 records every process with its parent. Reduce each to a `parent>child` token, keep only tokens that are new to each host since the intrusion started, and you have a document per host. [scikit-learn's](https://scikit-learn.org/) `TfidfVectorizer` weights those documents so a pair present on every machine (`explorer.exe>chrome.exe`) counts for almost nothing and a rare one (`wmiprvse.exe>powershell.exe`, `services.exe>` a random eight-character binary from a PsExec-style install, [T1569.002](/mitre/T1569.002)) dominates. [`NearestNeighbors`](https://scikit-learn.org/stable/modules/generated/sklearn.neighbors.NearestNeighbors.html) with cosine distance then ranks hosts by similarity to patient zero.

```python
from pathlib import PureWindowsPath
from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.neighbors import NearestNeighbors

procs = pd.read_json("sysmon_eid1.jsonl", lines=True)
procs["ts"] = pd.to_datetime(procs["UtcTime"], utc=True)
exe = lambda p: PureWindowsPath(p).name.lower() if isinstance(p, str) else "?"
procs["pair"] = procs["ParentImage"].map(exe) + ">" + procs["Image"].map(exe)

base = procs[procs["ts"].between(t0 - pd.Timedelta(days=30), t0)]
new = (procs[procs["ts"] >= t0]
       .merge(base[["Computer", "pair"]].drop_duplicates(),
              how="left", indicator=True)
       .query("_merge == 'left_only'"))

docs = new.groupby("Computer")["pair"].apply(list)
vec = TfidfVectorizer(analyzer=lambda pairs: pairs, sublinear_tf=True)
X = vec.fit_transform(docs)

nn = NearestNeighbors(metric="cosine").fit(X)
i = docs.index.get_loc("WS-0412")
dist, idx = nn.kneighbors(X[i], n_neighbors=min(20, len(docs)))
lookalikes = pd.Series(1 - dist[0], index=docs.index[idx[0]], name="similarity")
```

Always print the terms behind the score before you act on it: `sorted(zip(X[i].toarray()[0], vec.get_feature_names_out()), reverse=True)[:10]`. If the top terms for patient zero are attacker tradecraft, the neighbors sharing them are your next images. If the top terms are `ccmexec.exe>` something, the similarity is a software deployment, not an intrusion.

Hosts that appear in both `in_reach` and the top of `lookalikes` go to the front of the queue. That intersection is the deliverable: a ranked list a lead responder can hand to the people pulling disk images.

## Where this breaks

- **VDI and freshly imaged hosts** have no 30-day baseline, so every pair is "new" and they float to the top of every similarity list. Baseline them against a golden-image host instead.
- **Change windows.** An SCCM or Intune push during the incident window produces a cluster of nearly identical new pairs on hundreds of hosts. Check the change calendar before you trust a large cluster.
- **Hands-on-keyboard attackers using RDP and built-in tools** may leave few distinctive process pairs. The logon graph carries more weight then, and the similarity list carries less.
- **Identity outside Windows.** Entra ID sign-ins, VPN sessions, and SSH to Linux hosts are not in 4624. If the attacker moved through the cloud tenant, this graph misses that hop entirely.

None of it replaces forensic confirmation. It decides the order you collect evidence in, which matters because Mandiant's [M-Trends 2025](https://cloud.google.com/security/resources/m-trends) still puts global median dwell time at 11 days: an attacker who has been inside for a week and a half has had time to spread. NIST's [SP 800-61 Rev. 3](https://csrc.nist.gov/pubs/sp/800/61/r3/final) frames scoping as part of continuous detection and response rather than a one-time step, and this is the kind of analysis worth rerunning as new logs arrive.

## Where to learn it

You can teach yourself most of this with public data. Three sources are worth your time:

- **[OTRF Security-Datasets](https://github.com/OTRF/Security-Datasets):** recorded attack simulations with Sysmon and Security logs, including lateral movement over WMI, PsExec, and remote services.
- **[Splunk BOTS v3](https://github.com/splunk/botsv3):** a multi-host environment with a full scenario, which is what you need to practice scoping rather than single-host triage.
- **[EVTX-ATTACK-SAMPLES](https://github.com/sbousseaden/EVTX-ATTACK-SAMPLES):** raw `.evtx` files organized by ATT&CK tactic. Good for checking that your parser handles the real fields.

Run [Chainsaw](https://github.com/WithSecureLabs/chainsaw) or [Hayabusa](https://github.com/Yamato-Security/hayabusa) over the same files first. Sigma-rule hits make good labels for checking whether your similarity ranking surfaces the hosts the rules flagged.

If you want instruction, judge a course by its lab data. A course that teaches IR analytics on one host's logs, or on a Kaggle intrusion dataset with no hostnames and no timestamps, cannot teach scoping. Ask whether the labs include many hosts, real Windows event fields, and a question with a time dimension. For the timeline-building and command-line clustering half of the job, see [data science for incident responders](/blog/data-science-for-incident-responders).

We teach the building blocks (pandas on Windows and network logs, TF-IDF, nearest neighbors, anomaly detection) in [Threat Hunting with Data Science](/courses/threat-hunting-data-science) and on day four of [Applied Data Science & AI](/courses/applied-data-science-ai), in Jupyter in the AI Training Dojo. The same notebooks you write for a hunt are the ones you open at 3 a.m. when there is a patient zero and a CISO waiting on a number.

## FAQ

### Where can I learn applied machine learning for incident response?

Look for training where the labs are multi-host incident data (Windows Security and Sysmon logs from a whole network, not a single CSV) and where the exercises answer an IR question: which hosts are in scope, when the intrusion started, and which accounts were used. Practice on public data such as OTRF Security-Datasets, Splunk's BOTS v3, and EVTX-ATTACK-SAMPLES. GTK Cyber's Threat Hunting with Data Science and Applied Data Science & AI courses teach the underlying methods (pandas, scikit-learn, anomaly detection) in Jupyter on security data.

### Can machine learning tell me which hosts are compromised during an incident?

No. It ranks hosts by how closely their behavior after the intrusion resembles the known-bad host, and by whether an attacker could have reached them through logons in the right time order. That ranking tells you which machines to image first. Confirmation still comes from forensic artifacts on each host: Prefetch, Amcache, service installs, memory.

### Why use TF-IDF on parent-child process pairs instead of a list of IOCs?

IOCs only match what you already know: a hash, a file name, an IP. Attackers rename binaries and rotate infrastructure between hosts. Parent-child pairs such as wmiprvse.exe launching powershell.exe describe the execution pattern, which tends to stay the same across hosts even when the payload changes. TF-IDF down-weights pairs that appear on every host, so the rare ones drive the similarity score.


---

Canonical: https://gtkcyber.com/blog/applied-ml-incident-response-scoping/