# AI Red Team Training for Security Engineers: Count the Trials

By Summer Rankin · 2026-09-25

> AI red team training for security engineers should teach findings as rates: sample sizes, confidence bounds, and retest math for nondeterministic LLM targets.

A security engineer gets a ticket: the support chatbot leaked its system prompt to a role-play jailbreak. The model team pushes a revised system prompt and a new input filter. The engineer replays the payload ten times, gets ten refusals, and closes the ticket.

That retest proves very little. Ten clean runs against a nondeterministic target put the one-sided 95% upper bound on the true success rate at about 26%. The fix might work. It might also leave a payload that lands one time in five, and ten trials cannot tell those two apart.

This is the gap in most AI red team training aimed at security engineers. Courses teach payloads: [LLM Jailbreak (AML.T0054)](/atlas/AML.T0054), [LLM Prompt Injection (AML.T0051)](/atlas/AML.T0051), and the entries under [OWASP LLM01:2025](https://genai.owasp.org/llmrisk/llm01-prompt-injection/). They rarely teach what an engineer needs to sign off on a fix, which is how to turn a pile of pass/fail runs into a claim that holds up.

## Temperature 0 Does Not Save You

The usual first reaction is to pin `temperature=0` and treat the model like a deterministic function. It is not one. Thinking Machines Lab's write-up [Defeating Nondeterminism in LLM Inference](https://thinkingmachines.ai/blog/defeating-nondeterminism-in-llm-inference/) sent the same prompt a thousand times at temperature 0 and got dozens of distinct completions. The cause is ordinary: inference servers batch requests together, and the floating-point results of the kernels change with batch size. Your prompt's output depends on who else was hitting the endpoint at that moment.

Two consequences for testing. First, a single reproduction is an anecdote even at temperature 0. Second, the production app almost never runs at temperature 0, so the configuration you should test is the deployed one, with its real sampling settings, system prompt, and retrieval context. The per-attempt outcome is a coin flip with an unknown bias. Your job is to estimate the bias.

## Report an Interval, Not a Screenshot

Every red-team case becomes a count: `k` successes in `n` attempts. The number that belongs in the finding is the interval around `k/n`, and [SciPy](https://docs.scipy.org/doc/scipy/reference/generated/scipy.stats.binomtest.html) computes it in one line:

```python
from scipy.stats import binomtest

def asr(k, n):
    ci = binomtest(k, n).proportion_ci(confidence_level=0.95, method="wilson")
    return f"{k}/{n} = {k/n:.0%} (95% CI {ci.low:.0%} to {ci.high:.0%})"

print(asr(3, 20))   # 3/20 = 15% (95% CI 5% to 36%)
```

Three hits in twenty is not "15%." It is "somewhere between about 5% and 36%," and the upper end is what a risk owner should plan around. Use the Wilson interval rather than the textbook normal approximation, which produces nonsense (negative lower bounds, zero-width intervals) at the small counts red teams actually collect.

Zero successes deserves its own rule. With `0/n`, the one-sided 95% upper bound is `1 - 0.05**(1/n)`, which is close to `3/n`. Work it backward to plan the retest:

| Claim you want to make | Clean attempts required |
|---|---|
| ASR below 25% | 11 |
| ASR below 10% | 29 |
| ASR below 5% | 59 |
| ASR below 1% | 299 |

Pick the threshold before running anything. "We need the leak rate under 5%" is a requirement an engineer can test. "It seems fixed" is not.

## Before and After Is a Two-Sample Test

Fix verification compares two rates. Suppose the payload landed 6 of 40 times before the patch and 1 of 40 after. That looks like a clear win. [Fisher's exact test](https://docs.scipy.org/doc/scipy/reference/generated/scipy.stats.fisher_exact.html) disagrees:

```python
from scipy.stats import fisher_exact

before, after, n = 6, 1, 40
_, p = fisher_exact([[before, n - before], [after, n - after]], alternative="greater")
print(f"p = {p:.3f}")   # p = 0.054
```

At the conventional 0.05 threshold, 40 runs per arm cannot distinguish that patch from no change. Run 100 per arm, or accept that the ticket closes on judgment instead of evidence and write that down in the finding.

The tooling handles the repetition. [promptfoo](https://www.promptfoo.dev/docs/usage/command-line/) reruns every test case with `promptfoo eval --repeat 40`, and its assertions give you a countable pass/fail per run. [garak](https://github.com/NVIDIA/garak) sets outputs per probe prompt with `--generations`, and its report counts how many tripped a detector. Neither tool picks `n`. That part is on you.

## The Rate Is Not the Risk

A per-attempt ASR describes one try. Attackers get many. At 4% per attempt, the chance of at least one success over 50 attempts is `1 - 0.96**50`, about 87%. For a public chatbot with no per-user rate limit, a "low" ASR is a working exploit with a short wait.

This is where security engineers earn their place on an AI red team. The compensating controls are ones they already build: per-identity rate limits on the model endpoint, alerting on repeated refusals from one session (a refusal burst is a decent jailbreak-in-progress signal), and output-side checks that do not depend on the model cooperating, such as a canary string planted in the system prompt and blocked at the response filter. A finding that says "2/50 attempts, ASR 4% (95% CI 1% to 13%), no rate limit, no refusal alerting" tells the owner what to fix. A finding with one screenshot starts an argument.

## Where Counting Does Not Apply

Rates are the wrong frame when one success is enough. If an injected instruction can make an agent call a tool that runs without an authorization check, the fix is the authorization check, and it does not matter whether the injection lands in 2% or 90% of attempts. The same holds for anything that writes to a durable store: a poisoned document that enters a RAG index once stays there.

The binomial math also assumes independent trials, and some stacks quietly break that. An application-level semantic cache that returns stored answers for similar prompts turns forty "attempts" into one attempt and thirty-nine replays. Check for a cache before trusting any count, and vary a nonce in the payload if you need to bypass it.

Most of the statistics here is first-year material, but it changes what a red-team report is allowed to claim. The [AI Red-Teaming](/courses/ai-red-teaming) course at GTK Cyber puts robustness evaluation and red-team reporting in the same two days as the payload work, because a finding a model owner cannot act on is not finished. For the payload side of the job, start with [how to red team an LLM-powered application](/blog/red-teaming-llm-powered-applications).

## FAQ

### How many attempts do I need before I can say a jailbreak fix works?

It depends on the residual rate you are willing to accept. If you run n independent attempts and see zero successes, the one-sided 95% upper bound on the true success rate is 1 - 0.05^(1/n), which is close to 3/n (the statistician's rule of three). Ten clean attempts only bound the rate below about 26%. Showing it is under 5% takes 59 clean attempts, and under 1% takes about 300. Decide the threshold first, then compute the trial count, then run the retest.

### Does setting temperature to 0 make LLM red-team results reproducible?

No. Temperature 0 removes sampling randomness, but hosted and self-hosted inference servers still produce different outputs for the same prompt because floating-point results change with batch composition on the GPU. Thinking Machines Lab documented this in 2025 by sending one prompt a thousand times at temperature 0 and getting dozens of distinct completions. Your production application also probably does not run at temperature 0, so testing there measures a configuration nobody attacks.

### What is the difference between attack success rate and whether a system is exploitable?

Attack success rate (ASR) is the fraction of attempts that succeed for a given payload and configuration. Exploitability depends on how many attempts an attacker gets. A 4% per-attempt rate becomes roughly an 87% chance of at least one success over 50 attempts. If the endpoint has no per-user rate limit and no detection on repeated refusals, a low ASR is still a working exploit. Report the rate and the retry budget together.

### Which tools support repeated trials for LLM red-teaming?

Promptfoo's eval command takes a --repeat flag that reruns every test case, and its assertions give you a pass/fail per run you can count. Garak's --generations option sets how many outputs are requested for each probe prompt, and its report counts how many of those tripped a detector. For the statistics, scipy.stats.binomtest gives confidence intervals on a success rate and scipy.stats.fisher_exact compares a before and after run. None of these tools pick the sample size for you.


---

Canonical: https://gtkcyber.com/blog/ai-red-team-training-security-engineers/