During the Second World War, radar operators faced a genuinely difficult problem: a blip on the screen might be an incoming aircraft, or it might be noise, weather, a flock of birds, an equipment artefact. Two operators watching the exact same screen, with the exact same ability to perceive a signal, could still call it differently, one erring toward "report it" to avoid missing a real threat, another erring toward "ignore it" to avoid crying wolf. Psychophysicists later formalised this into signal detection theory, and one of its central insights is worth borrowing directly for CRO.

The theory splits what looks like a single skill, detecting a signal, into two genuinely separate things: sensitivity, how well you can actually distinguish signal from noise, and criterion, how much evidence you personally require before you're willing to call something a signal. Two people with identical sensitivity can reach opposite conclusions purely because their criteria differ. That distinction maps onto A/B testing almost exactly, and it explains a specific, recurring source of disagreement that "check your sample size" advice doesn't touch.

What signal detection theory actually measures

Sensitivity is a property of the test itself: given a genuine underlying effect, how reliably can it be distinguished from random noise. In A/B testing terms, this is what proper test design controls for, sample size, minimum detectable effect, test duration. Get these right and sensitivity is high, a real difference is more likely to actually show up as one.

Criterion is a completely separate decision: given the evidence a test has produced, how much of it is required before you're willing to declare a winner. This is what a significance threshold actually is. A team requiring p<0.05 has set a more permissive criterion than a team requiring p<0.01, and neither choice is simply "correct," they represent different trade-offs between catching real effects quickly and avoiding false ones.

The one-sentence distinction

Sensitivity is about the test. Criterion is about the person reading it. Most disagreements over a test result are actually disagreements about criterion, dressed up as disagreements about data.

The hit, miss, false-alarm, correct-rejection matrix

Signal detection theory organises every possible outcome into a simple 2x2 matrix, and it translates directly onto a CRO test.

SDT term
A/B testing equivalent
Hit
A genuine winning variant is correctly identified as a winner. The outcome every test is trying to produce.
Miss
A genuine winning variant fails to reach significance and gets discarded, usually from an underpowered test or one stopped too early. A real improvement is lost.
False alarm
No genuine difference exists, but the test is declared a winner anyway, commonly from peeking at results early or running many simultaneous tests without correcting for it.
Correct rejection
No genuine difference exists, and the test correctly concludes there isn't one. Unglamorous, but exactly what a well-calibrated test should do most of the time.

Why "criterion" explains disagreements the data can't

Most CRO content treats a false positive as a data problem, not enough sample size, a bug in the tracking, a novelty effect. Those are real causes. But a significant proportion of disputed test results come from a criterion mismatch that nobody named explicitly.

1
Peeking shifts the criterion without anyone deciding to

Checking results daily and stopping the moment p crosses 0.05 feels like following the rules, but it silently inflates the real false positive rate far above the stated 5%, because the test gets multiple chances to cross the line by chance alone. The criterion everyone thinks they're using isn't the one actually in effect.

2
Different teams use different thresholds without a house standard

One analyst calls a test significant at 90% confidence, another insists on 99%. Both are internally consistent choices. Without an agreed house standard, the same underlying data will produce genuine, defensible disagreement between them every time.

3
One-tailed versus two-tailed testing changes the criterion invisibly

A one-tailed test, only checking whether the variant beat control, not whether it also lost, effectively lowers the bar for significance in the direction you're hoping for. It's a legitimate choice in the right context, but it's a criterion decision, not a neutral default.

Setting a criterion deliberately, instead of by accident

None of this argues for a single universally correct criterion. It argues for choosing one on purpose, before a test runs, rather than discovering it retroactively once results are already in front of a stakeholder who wants an answer.

How to set criterion deliberately
  • Fix the significance threshold and stopping rule before launch: Decide the confidence level and the sample size or duration that ends the test, and write both down before results start arriving. This is the single highest-leverage fix for accidental criterion drift.
  • Separate test design conversations from result-reading conversations: Sensitivity, sample size, minimum detectable effect, gets decided at design time. Criterion, the significance bar, gets decided at the same time, not renegotiated once the data looks a certain way.
  • Use a sequential testing method if speed genuinely matters: These are built specifically to allow legitimate early stopping without inflating the false positive rate the way informal peeking does, giving you speed without quietly lowering your criterion.
  • Document a house standard so criterion isn't set person by person: An agreed default confidence level, applied consistently, removes the specific disagreement this piece describes before it starts.
Key takeaways
  • Sensitivity, how reliably a test can detect a real effect, and criterion, how much evidence is required to call it, are genuinely separate decisions
  • The hit/miss/false-alarm/correct-rejection matrix maps directly onto test outcomes and clarifies what each type of error actually costs
  • Peeking, inconsistent thresholds across teams, and one-tailed testing all shift the effective criterion without anyone deciding to
  • Fix the significance threshold and stopping rule before a test launches, and document a house standard so criterion isn't set person by person
Set a testing criterion your whole team actually agrees on

Our A/B testing programmes fix the significance threshold and stopping rule before a test launches, so a result is a genuine hit, not a criterion mismatch dressed up as a debate about data.

Explore A/B Testing →