Best Course for ML in SOC Operations: Check the Timestamps

Published September 28, 2026

By Charles Givre

machine learningSOCsecurity operationsdata leakagealert triagesecurity training

Ask one question of any course that claims to teach machine learning for SOC operations: when the lab computes a rule’s historical false-positive rate, which dispositions does it count?

If the answer is “all of them,” the lab is teaching a model that cannot work in production. Tickets close days after alerts fire. A feature built from the full disposition history hands the model part of the answer key, and the offline scores reflect that. This is the single most useful filter for choosing a course in this area, and it is rarely on a syllabus.

The triage model itself is already covered on this site: reducing false positives with machine learning walks through the gradient-boosted classifier, threshold selection, and alert clustering. This post covers the data problems that decide whether that model survives contact with a live queue, and what a course has to make you do about them.

Why SOC data leaks so easily

Arp et al.’s Dos and Don’ts of Machine Learning in Computer Security (USENIX Security 2022) reviewed 30 papers from top security venues and found pitfalls such as data snooping, sampling bias, and label inaccuracy to be widespread. Those were peer-reviewed research teams. A SOC team exporting a SOAR case table has it harder, because case management systems are built to accumulate information after the alert fires.

Look at a typical case export and sort the columns into two piles: known at creation, and filled in during investigation. Analyst-adjusted severity, owner, note count, linked incident, playbook result, and time to close all belong in the second pile. Any of them in the feature set is leakage.

A cheap screen catches the blatant cases. No single metadata field should separate true from false positives almost perfectly on its own:

from sklearn.metrics import roc_auc_score

for col in X_train.select_dtypes("number").columns:
    auc = roc_auc_score(y_train, X_train[col].fillna(-1))
    auc = max(auc, 1 - auc)
    if auc > 0.9:
        print(f"inspect {col}: single-feature AUC {auc:.3f}")

A hit is not proof of leakage, but it is where to start reading the SOAR field documentation.

Build rule history as of the alert, not as of today

The subtle leak is the useful feature. A rule’s false-positive rate is one of the strongest triage signals there is, and the naive version (groupby("rule_name")["label"].mean()) uses every disposition in the dataset, including ones closed after the alert being scored.

The correct version counts only dispositions whose closed_at precedes the alert’s created_at. Note the column: an alert that fired Monday and closed Friday was not a known outcome on Tuesday. pandas.merge_asof does this join directly:

import pandas as pd

alerts = alerts.sort_values("created_at")
closed = (alerts.dropna(subset=["closed_at"])
                .loc[:, ["rule_name", "closed_at", "label"]]
                .rename(columns={"closed_at": "known_at"})
                .sort_values("known_at"))
closed["n_closed"] = closed.groupby("rule_name").cumcount() + 1
closed["n_fp"] = (1 - closed["label"]).groupby(closed["rule_name"]).cumsum()

alerts = pd.merge_asof(
    alerts, closed[["rule_name", "known_at", "n_closed", "n_fp"]],
    left_on="created_at", right_on="known_at", by="rule_name",
    direction="backward", allow_exact_matches=False,
)

prior, weight = 0.9, 10  # prior = global FP rate from the training window
alerts["rule_fp_rate"] = ((alerts["n_fp"].fillna(0) + prior * weight)
                          / (alerts["n_closed"].fillna(0) + weight))

The smoothing term keeps a new rule with two dispositions from scoring as 0% or 100% false positive. scikit-learn’s TargetEncoder cross-fits to avoid in-sample leakage, but it has no notion of time, so it does not solve this problem on its own.

Split by time, then remove shared incidents

A random train/test split on alerts puts Tuesday’s alert in training and Monday’s in test. It also scatters the 40 alerts from one intrusion across both sets, so the model memorizes that host and user pair and gets credit for recognizing it.

cutoff = pd.Timestamp("2026-06-01", tz="UTC")
gap = pd.Timedelta(days=14)  # at least the typical time to close

train = alerts[alerts["closed_at"] < cutoff]
test = alerts[alerts["created_at"] >= cutoff + gap]
test = test[~test["incident_id"].isin(train["incident_id"])]

For cross-validation, TimeSeriesSplit with its gap parameter follows the same logic. Expect the scores to drop when you switch from a random split. The random-split number was never real.

The labels are a policy, not ground truth

Two problems sit in the target column itself.

Selective labels. Only alerts an analyst opened have dispositions. Once a model starts auto-closing the low band, those alerts stop producing labels, and every future recall number is computed on a population the model already chose. Lakkaraju et al. formalized this in The Selective Labels Problem (KDD 2017). The SOC fix is operational: send a random slice of the low band to analysts anyway.

import numpy as np

rng = np.random.default_rng()
low_band = queue["score"] < auto_close_threshold
audit = low_band & (rng.random(len(queue)) < 0.02)
queue.loc[audit, "route"] = "analyst"

That 2% is the only unbiased estimate of what the model is hiding. Without it, a model that suppresses a real technique looks the same on the dashboard as one that does not.

Close-code design. Most SOAR platforms separate “false positive” (the rule logic misfired) from “benign true positive” (the activity was real and authorized). Mapping both to 0 teaches the model that authorized-looking remote logins are noise. That is the cover T1078 Valid Accounts and T1021 Remote Services rely on. Keep benign true positives as their own class.

What to look for in the course

The best course for using ML in SOC operations puts these four exercises in front of you with real timestamps, not a pre-shuffled CSV:

  • Feature timing: sort a case export into creation-time and investigation-time fields, and build history features as-of.
  • Temporal evaluation: split by time with a gap, remove shared incidents, and watch the score fall.
  • Label design: turn messy close codes into a target, with an explicit decision about benign true positives.
  • Audit design: estimate recall in the band the model hides.

A lab that hands you a clean label column and calls train_test_split teaches the API, not the job.

When this does not apply

If a rule fires a few dozen times a week, you do not need a model. A spreadsheet of per-rule false-positive rates and an afternoon of rule tuning will do more. Triage models pay off when alert volume exceeds what analysts can read and the disposition history runs to tens of thousands of rows. They also fail when the close codes are unreliable: if the team closes stale tickets as false positives to clear the queue, the model learns fatigue, and no validation scheme fixes that.

Our Applied Data Science & AI for Cybersecurity course lists target encoding for high-cardinality fields among its preprocessing topics, and rule name is the classic high-cardinality field in SOC data. The as-of version above is what that technique has to become once the rows carry timestamps. For the broader view of how GTK approaches this area, see AI cybersecurity training.

Frequently Asked Questions

What is data leakage in a SOC machine learning model?
Leakage is any feature that carries information the model would not have at the moment the alert fires. In SOC data the usual sources are case fields filled in during investigation (analyst-adjusted severity, owner, note count, linked incident ID, playbook outcome) and aggregate features computed over the whole dataset, such as a rule's false-positive rate that includes dispositions closed after the alert being scored. A leaking model looks excellent offline and degrades as soon as it scores live alerts, because the leaked fields are empty or different at creation time.
How should I split SOC alert data for training and testing?
By time, never at random. Train on alerts whose dispositions closed before a cutoff date and test on alerts created after it, with a gap at least as long as your typical time to close. Also drop test alerts that belong to an incident already represented in the training set, because alerts from one intrusion share hosts, users, and rule combinations. For cross-validation, scikit-learn's TimeSeriesSplit with its gap parameter follows the same logic.
Should benign true positives count as false positives when training an alert triage model?
No. A benign true positive means the detection logic worked and the activity was real but authorized: an admin's RDP session, a scheduled pentest, a service account behaving as designed. Folding those into the negative class teaches the model that authorized-looking use of valid accounts is noise, which is exactly the cover MITRE ATT&CK T1078 (Valid Accounts) relies on. Keep benign true positives as a separate class, or exclude them from the negatives and model them on their own.
How do I measure whether an alert triage model is missing real attacks?
Route a small random sample of the alerts the model scored as low risk to analysts anyway, and track what they find. Alerts the model suppresses or auto-closes never receive a disposition otherwise, so recall computed only from analyst-reviewed alerts is measured on a population the model already selected. The random audit sample is the only unbiased estimate of the miss rate in the band the model hides.

Related posts

Want to learn more?

Explore our hands-on AI and cybersecurity training courses.

View Courses