Ask one question of any course that claims to teach machine learning for SOC operations: when the lab computes a rule’s historical false-positive rate, which dispositions does it count?
If the answer is “all of them,” the lab is teaching a model that cannot work in production. Tickets close days after alerts fire. A feature built from the full disposition history hands the model part of the answer key, and the offline scores reflect that. This is the single most useful filter for choosing a course in this area, and it is rarely on a syllabus.
The triage model itself is already covered on this site: reducing false positives with machine learning walks through the gradient-boosted classifier, threshold selection, and alert clustering. This post covers the data problems that decide whether that model survives contact with a live queue, and what a course has to make you do about them.
Why SOC data leaks so easily
Arp et al.’s Dos and Don’ts of Machine Learning in Computer Security (USENIX Security 2022) reviewed 30 papers from top security venues and found pitfalls such as data snooping, sampling bias, and label inaccuracy to be widespread. Those were peer-reviewed research teams. A SOC team exporting a SOAR case table has it harder, because case management systems are built to accumulate information after the alert fires.
Look at a typical case export and sort the columns into two piles: known at creation, and filled in during investigation. Analyst-adjusted severity, owner, note count, linked incident, playbook result, and time to close all belong in the second pile. Any of them in the feature set is leakage.
A cheap screen catches the blatant cases. No single metadata field should separate true from false positives almost perfectly on its own:
from sklearn.metrics import roc_auc_score
for col in X_train.select_dtypes("number").columns:
auc = roc_auc_score(y_train, X_train[col].fillna(-1))
auc = max(auc, 1 - auc)
if auc > 0.9:
print(f"inspect {col}: single-feature AUC {auc:.3f}")
A hit is not proof of leakage, but it is where to start reading the SOAR field documentation.
Build rule history as of the alert, not as of today
The subtle leak is the useful feature. A rule’s false-positive rate is one of the strongest triage signals there is, and the naive version (groupby("rule_name")["label"].mean()) uses every disposition in the dataset, including ones closed after the alert being scored.
The correct version counts only dispositions whose closed_at precedes the alert’s created_at. Note the column: an alert that fired Monday and closed Friday was not a known outcome on Tuesday. pandas.merge_asof does this join directly:
import pandas as pd
alerts = alerts.sort_values("created_at")
closed = (alerts.dropna(subset=["closed_at"])
.loc[:, ["rule_name", "closed_at", "label"]]
.rename(columns={"closed_at": "known_at"})
.sort_values("known_at"))
closed["n_closed"] = closed.groupby("rule_name").cumcount() + 1
closed["n_fp"] = (1 - closed["label"]).groupby(closed["rule_name"]).cumsum()
alerts = pd.merge_asof(
alerts, closed[["rule_name", "known_at", "n_closed", "n_fp"]],
left_on="created_at", right_on="known_at", by="rule_name",
direction="backward", allow_exact_matches=False,
)
prior, weight = 0.9, 10 # prior = global FP rate from the training window
alerts["rule_fp_rate"] = ((alerts["n_fp"].fillna(0) + prior * weight)
/ (alerts["n_closed"].fillna(0) + weight))
The smoothing term keeps a new rule with two dispositions from scoring as 0% or 100% false positive. scikit-learn’s TargetEncoder cross-fits to avoid in-sample leakage, but it has no notion of time, so it does not solve this problem on its own.
Split by time, then remove shared incidents
A random train/test split on alerts puts Tuesday’s alert in training and Monday’s in test. It also scatters the 40 alerts from one intrusion across both sets, so the model memorizes that host and user pair and gets credit for recognizing it.
cutoff = pd.Timestamp("2026-06-01", tz="UTC")
gap = pd.Timedelta(days=14) # at least the typical time to close
train = alerts[alerts["closed_at"] < cutoff]
test = alerts[alerts["created_at"] >= cutoff + gap]
test = test[~test["incident_id"].isin(train["incident_id"])]
For cross-validation, TimeSeriesSplit with its gap parameter follows the same logic. Expect the scores to drop when you switch from a random split. The random-split number was never real.
The labels are a policy, not ground truth
Two problems sit in the target column itself.
Selective labels. Only alerts an analyst opened have dispositions. Once a model starts auto-closing the low band, those alerts stop producing labels, and every future recall number is computed on a population the model already chose. Lakkaraju et al. formalized this in The Selective Labels Problem (KDD 2017). The SOC fix is operational: send a random slice of the low band to analysts anyway.
import numpy as np
rng = np.random.default_rng()
low_band = queue["score"] < auto_close_threshold
audit = low_band & (rng.random(len(queue)) < 0.02)
queue.loc[audit, "route"] = "analyst"
That 2% is the only unbiased estimate of what the model is hiding. Without it, a model that suppresses a real technique looks the same on the dashboard as one that does not.
Close-code design. Most SOAR platforms separate “false positive” (the rule logic misfired) from “benign true positive” (the activity was real and authorized). Mapping both to 0 teaches the model that authorized-looking remote logins are noise. That is the cover T1078 Valid Accounts and T1021 Remote Services rely on. Keep benign true positives as their own class.
What to look for in the course
The best course for using ML in SOC operations puts these four exercises in front of you with real timestamps, not a pre-shuffled CSV:
- Feature timing: sort a case export into creation-time and investigation-time fields, and build history features as-of.
- Temporal evaluation: split by time with a gap, remove shared incidents, and watch the score fall.
- Label design: turn messy close codes into a target, with an explicit decision about benign true positives.
- Audit design: estimate recall in the band the model hides.
A lab that hands you a clean label column and calls train_test_split teaches the API, not the job.
When this does not apply
If a rule fires a few dozen times a week, you do not need a model. A spreadsheet of per-rule false-positive rates and an afternoon of rule tuning will do more. Triage models pay off when alert volume exceeds what analysts can read and the disposition history runs to tens of thousands of rows. They also fail when the close codes are unreliable: if the team closes stale tickets as false positives to clear the queue, the model learns fatigue, and no validation scheme fixes that.
Our Applied Data Science & AI for Cybersecurity course lists target encoding for high-cardinality fields among its preprocessing topics, and rule name is the classic high-cardinality field in SOC data. The as-of version above is what that technique has to become once the rows carry timestamps. For the broader view of how GTK approaches this area, see AI cybersecurity training.