A security engineer gets a ticket: the support chatbot leaked its system prompt to a role-play jailbreak. The model team pushes a revised system prompt and a new input filter. The engineer replays the payload ten times, gets ten refusals, and closes the ticket.
That retest proves very little. Ten clean runs against a nondeterministic target put the one-sided 95% upper bound on the true success rate at about 26%. The fix might work. It might also leave a payload that lands one time in five, and ten trials cannot tell those two apart.
This is the gap in most AI red team training aimed at security engineers. Courses teach payloads: LLM Jailbreak (AML.T0054), LLM Prompt Injection (AML.T0051), and the entries under OWASP LLM01:2025. They rarely teach what an engineer needs to sign off on a fix, which is how to turn a pile of pass/fail runs into a claim that holds up.
Temperature 0 Does Not Save You
The usual first reaction is to pin temperature=0 and treat the model like a deterministic function. It is not one. Thinking Machines Lab’s write-up Defeating Nondeterminism in LLM Inference sent the same prompt a thousand times at temperature 0 and got dozens of distinct completions. The cause is ordinary: inference servers batch requests together, and the floating-point results of the kernels change with batch size. Your prompt’s output depends on who else was hitting the endpoint at that moment.
Two consequences for testing. First, a single reproduction is an anecdote even at temperature 0. Second, the production app almost never runs at temperature 0, so the configuration you should test is the deployed one, with its real sampling settings, system prompt, and retrieval context. The per-attempt outcome is a coin flip with an unknown bias. Your job is to estimate the bias.
Report an Interval, Not a Screenshot
Every red-team case becomes a count: k successes in n attempts. The number that belongs in the finding is the interval around k/n, and SciPy computes it in one line:
from scipy.stats import binomtest
def asr(k, n):
ci = binomtest(k, n).proportion_ci(confidence_level=0.95, method="wilson")
return f"{k}/{n} = {k/n:.0%} (95% CI {ci.low:.0%} to {ci.high:.0%})"
print(asr(3, 20)) # 3/20 = 15% (95% CI 5% to 36%)
Three hits in twenty is not “15%.” It is “somewhere between about 5% and 36%,” and the upper end is what a risk owner should plan around. Use the Wilson interval rather than the textbook normal approximation, which produces nonsense (negative lower bounds, zero-width intervals) at the small counts red teams actually collect.
Zero successes deserves its own rule. With 0/n, the one-sided 95% upper bound is 1 - 0.05**(1/n), which is close to 3/n. Work it backward to plan the retest:
| Claim you want to make | Clean attempts required |
|---|---|
| ASR below 25% | 11 |
| ASR below 10% | 29 |
| ASR below 5% | 59 |
| ASR below 1% | 299 |
Pick the threshold before running anything. “We need the leak rate under 5%” is a requirement an engineer can test. “It seems fixed” is not.
Before and After Is a Two-Sample Test
Fix verification compares two rates. Suppose the payload landed 6 of 40 times before the patch and 1 of 40 after. That looks like a clear win. Fisher’s exact test disagrees:
from scipy.stats import fisher_exact
before, after, n = 6, 1, 40
_, p = fisher_exact([[before, n - before], [after, n - after]], alternative="greater")
print(f"p = {p:.3f}") # p = 0.054
At the conventional 0.05 threshold, 40 runs per arm cannot distinguish that patch from no change. Run 100 per arm, or accept that the ticket closes on judgment instead of evidence and write that down in the finding.
The tooling handles the repetition. promptfoo reruns every test case with promptfoo eval --repeat 40, and its assertions give you a countable pass/fail per run. garak sets outputs per probe prompt with --generations, and its report counts how many tripped a detector. Neither tool picks n. That part is on you.
The Rate Is Not the Risk
A per-attempt ASR describes one try. Attackers get many. At 4% per attempt, the chance of at least one success over 50 attempts is 1 - 0.96**50, about 87%. For a public chatbot with no per-user rate limit, a “low” ASR is a working exploit with a short wait.
This is where security engineers earn their place on an AI red team. The compensating controls are ones they already build: per-identity rate limits on the model endpoint, alerting on repeated refusals from one session (a refusal burst is a decent jailbreak-in-progress signal), and output-side checks that do not depend on the model cooperating, such as a canary string planted in the system prompt and blocked at the response filter. A finding that says “2/50 attempts, ASR 4% (95% CI 1% to 13%), no rate limit, no refusal alerting” tells the owner what to fix. A finding with one screenshot starts an argument.
Where Counting Does Not Apply
Rates are the wrong frame when one success is enough. If an injected instruction can make an agent call a tool that runs without an authorization check, the fix is the authorization check, and it does not matter whether the injection lands in 2% or 90% of attempts. The same holds for anything that writes to a durable store: a poisoned document that enters a RAG index once stays there.
The binomial math also assumes independent trials, and some stacks quietly break that. An application-level semantic cache that returns stored answers for similar prompts turns forty “attempts” into one attempt and thirty-nine replays. Check for a cache before trusting any count, and vary a nonce in the payload if you need to bypass it.
Most of the statistics here is first-year material, but it changes what a red-team report is allowed to claim. The AI Red-Teaming course at GTK Cyber puts robustness evaluation and red-team reporting in the same two days as the payload work, because a finding a model owner cannot act on is not finished. For the payload side of the job, start with how to red team an LLM-powered application.