What Bank Security Teams Need From AI Security Training

Published August 19, 2026

By Summer Rankin

financial servicesadversarial AImodel evaluationAI securitycybersecurity training

A fraud model is the only detection system in a bank that the adversary gets to query all day, at will, with a clean answer on every attempt. Approve or decline is a label. Card testing with a run of small transactions is not reconnaissance in any loose sense; it is a labeled query campaign against a classifier, and it is how the attacker learns the decision boundary without ever seeing the model.

Financial institutions have more production ML in the path of real money than almost any other sector, and the security teams asked to defend it were trained on networks, endpoints, and web applications. The gap is narrow and specific, and closing it does not require a data science curriculum.

Query Access Is the Exposure

Adversaries do not need gradients or weights to attack a deployed classifier. Black-box optimization (MITRE ATLAS AML.T0043.001) reconstructs enough of a decision boundary from output labels alone to find inputs that cross it, and the goal, evading the model (ATLAS AML.T0015), needs nothing else.

What limits the attack is query budget. Every probe costs the adversary a transaction, a card, or an account. Which reframes controls you already own: velocity limits, device reputation, and account-age gates are not only fraud controls, they are rate limits on the adversary’s learning loop. A team that understands the model as a queryable oracle will argue for those limits differently than a team that treats the model as a black box the vendor tuned.

Constrain the Attack or the Number Is Fiction

Here is where most first attempts go wrong. Adversarial ML libraries were built for images, where any pixel can take any value and an unconstrained perturbation is still a picture. Tabular financial features do not work that way. Turn a generic attack loose on a transaction record and it returns an evasive example with a negative transfer amount, an account age that decreased since last month, and a device first seen next Tuesday. The attack succeeded against the model and describes nothing an adversary can do.

Write down the action space first: per feature, does the adversary control it, in which direction, and at what cost?

import numpy as np

FEATURES = ["amount", "hour_of_day", "device_age_days", "velocity_1h", "acct_age_days"]

# Attacker-reachable moves, in scaled feature units, with the sign of the
# available direction. acct_age_days is absent because it cannot be moved.
ACTIONS = {
    "amount":          (-0.9, 0.0),    # smaller transfers are always available
    "hour_of_day":     (-8.0, 8.0),    # free
    "device_age_days": (0.0, 30.0),    # only increases, and only by waiting
    "velocity_1h":     (-4.0, 0.0),    # slowing down is free, speeding up is not
}

def cost_to_evade(model, x, trials=500, rng=np.random.default_rng(0)):
    """Fewest attacker actions that turn a decline into an approve."""
    cheapest = None
    for _ in range(trials):
        cand, moves = x.copy(), 0
        for feat, (lo, hi) in ACTIONS.items():
            step = rng.uniform(lo, hi)
            if abs(step) < 0.05:
                continue
            cand[FEATURES.index(feat)] += step
            moves += 1
        if model.predict(cand.reshape(1, -1))[0] == 0:      # scored legitimate
            cheapest = moves if cheapest is None else min(cheapest, moves)
    return cheapest                                          # None = no evasion found

Run that over held-out confirmed-fraud records and you get a distribution instead of a score. If 60 percent of declined transactions become approvals after two reachable moves, the model’s AUC is not the number that describes your risk. Cost-to-evade is, and it is denominated in things a fraud team already reasons about: attempts, cards, waiting time.

Use Adversarial Robustness Toolbox for the real version. HopSkipJump is decision-based, so it works against a model that returns only approve or decline, which is the access an external adversary actually has. Apply the same feature mask you defined above, because ART will otherwise perturb whatever you hand it.

The Governance Hook Already Exists

Security teams in banks tend to pitch this work as new risk requiring new budget. It is easier than that. SR 11-7, the 2011 Federal Reserve and OCC guidance on model risk management, already requires effective challenge and ongoing monitoring for models in use. Adversarial robustness is validation evidence under that standard: performance on adversary-chosen inputs rather than on inputs the historical sample happened to contain.

That makes the deliverable format the decision that matters. An evasion writeup filed as a red-team finding gets queued behind everything else in the security backlog. The same result filed as a validation finding against a model identifier in the model inventory has a remediation owner, a due date, and a validator who is obligated to look at it.

On the European side, DORA Article 26 requires threat-led penetration testing every three years for identified entities, scoped to systems supporting critical or important functions. It never says “model,” which is why the scoping conversation is worth having early: if payment fraud scoring supports a critical function, the classifier is in scope, and a test of the API in front of it is not a test of it.

Where This Training Does Not Pay Off

If your security team cannot get either the production model, a surrogate, or a scoring endpoint, none of the above runs, and in plenty of institutions that access takes longer to negotiate than the training takes to deliver. Start the request first.

When model access is genuinely blocked, test the pipeline instead, which is often the softer target anyway. Fraud and AML models retrain on analyst dispositions, so the label feedback loop is a poisoning path (ATLAS AML.T0020): an adversary who can influence which cases get marked legitimate, through mule accounts that generate clean history or by exhausting a review queue, is editing next month’s training set. That attack needs no model access at all.

And treat a failed evasion search as a weak result rather than a clean bill of health. A random search inside the action space gives a floor on the attacker’s cost, not a bound. Finding nothing means your search was not strong enough, which is exactly the honest sentence to put in the validation writeup.

We teach evasion, poisoning, and model extraction as labs rather than lecture in Applied Data Science and AI for Cybersecurity, half of which is hands-on notebook work, and financial services teams usually want the fraud-model version of those labs against their own feature set, which is what a custom engagement is for. Details on delivery inside a regulated environment are on the financial services training page, and the model-agnostic methodology is in how to evaluate ML model robustness for security use cases.

Frequently Asked Questions

What AI security training do financial services security teams need?
Three capabilities that generic AI courses skip. First, constrained adversarial evaluation of the models the institution already runs in production: fraud scoring, AML transaction monitoring, KYC document classification. Second, testing the pipeline around the model, since label feedback from analyst dispositions is a poisoning path (MITRE ATLAS AML.T0020) that most threat models never mention. Third, writing the results as model validation findings rather than penetration test findings, because model risk management is the function with the authority to make a modeling team act. The prerequisite is working Python and security testing experience, not a statistics degree.
How do you test a fraud model for adversarial evasion without producing unrealistic results?
Define the attacker's action space before you touch a library. For each feature, record whether the adversary controls it, which direction they can move it, and what moving it costs. Transaction amount usually goes down freely, account age cannot go down at all, device age only increases with waiting, and velocity features are bounded by the attacker's own patience. Then search only inside that space and report the number of actions needed to flip a decline into an approve. An unconstrained perturbation search will happily hand you a negative transfer amount from an account created in the future, which tells you nothing about your exposure.
How does adversarial testing of machine learning models fit under SR 11-7?
SR 11-7, the 2011 Federal Reserve and OCC guidance on model risk management, already requires effective challenge and ongoing monitoring of models in use. Adversarial robustness results are validation evidence: they measure whether the model holds up on inputs an adversary chooses rather than inputs the historical sample happened to contain. The practical consequence is a formatting decision. A red-team report describing an evasion technique gets triaged as a security finding and queued. The same result written as a validation finding against a specific model identifier lands in the model inventory, where remediation has an owner and a date.
Does DORA require testing machine learning models?
DORA (Regulation (EU) 2022/2554) does not name machine learning, and that is the point. Article 26 requires threat-led penetration testing at least every three years for identified financial entities, scoped to the ICT systems supporting critical or important functions. If transaction monitoring or payment fraud scoring is a critical function at your institution, the model making those decisions sits inside the scope, and a test that covered only the surrounding web and API layers has left the decision logic untested. Read the scoping question as an inventory question first: which models sit in the path of a critical function?

Related posts

Want to learn more?

Explore our hands-on AI and cybersecurity training courses.

View Courses