#5209: Sensitivity, Specificity & the Numbers Doctors Use

A field guide to the numbers in a medical paper — what sensitivity, effect size, and risk ratio actually answer, and when they mislead.

Featuring
Listen
0:00
0:00
Episode Details
Episode ID
MWP-5391
Published
Duration
22:21
Audio
Direct link
Pipeline
V5.2
TTS Engine
chatterbox-regular
Script Writing Agent
deepseek-v4-pro

AI-Generated Content: This podcast is created using AI personas. Please verify any important information independently.

Most people who read clinical trials hit the same wall: the prose is fine, then the results section starts throwing around sensitivity, specificity, standard deviation, effect size, and risk ratio, and you're expected to just know what they mean. The problem isn't that the terms are complicated. It's that each one answers a different question, and papers never label which question they're answering.

Sensitivity and specificity are properties of the test, not the patient. Sensitivity is the share of truly sick people a test correctly catches; specificity is the share of truly healthy people it correctly clears. Both are calculated within a single disease-status group, so they don't change when a condition gets more or less common. Positive and negative predictive value do change with prevalence — and they're the numbers a patient actually wants. A test that's 90% sensitive and 90% specific yields a 90% chance of true disease in a clinic where half the patients are sick, but only 8.3% in general screening where 1% have the condition. Same test. Different crowd.

The how-big-and-how-certain trio covers standard deviation (spread around the average), control group size (statistical power, conventionally at least 80%), and effect size (magnitude, separate from significance). Cohen's d thresholds of 0.2, 0.5, and 0.8 are rules of thumb that hardened into dogma — Cohen himself called them arbitrary, and a d of 0.5 means something very different in a medical trial than in a personality study.

Finally, the risk family. Risk ratio compares bad outcomes between treated and control groups, and relative risk reduction routinely makes tiny absolute changes sound enormous: the ASCOT trial's 36% relative reduction came with a 1.1% absolute reduction. Absolute risk reduction and number needed to treat tell the real story. Composite outcomes are their own trap — a drug can "reduce death, stroke, heart attack, and hospitalization by 15%" while only moving one of those four.

Downloads

Episode Audio

Download the full episode as an MP3 file

Download MP3
Transcript (TXT)

Plain text transcript file

Transcript (PDF)

Formatted PDF with styling

#5209: Sensitivity, Specificity & the Numbers Doctors Use

Corn
Daniel's been reading clinical trial write-ups again, and he hit the wall every smart layperson hits. You're following the prose fine, then the results section starts throwing around sensitivity, specificity, standard deviation, effect size, risk ratio, and you're supposed to just... know what those mean. He wants a field guide. Not a statistics course, but enough to actually read a paper and know whether the finding is real, how big it is, and whether it applies to you.
Herman
And the thing is, these terms aren't hard because they're complicated. They're hard because each one answers a different question, and papers never label which question they're answering. You're just expected to notice that sensitivity and specificity are about the test, not the patient, and that a risk ratio of point six means something completely different depending on how common the bad outcome was to begin with.
Corn
So let's do what Daniel actually asked. Walk through the diagnostic pair first, then the how-big-and-how-certain trio, then the risk family. And at every step, name the question the number is actually answering, because that's what most explainers skip.
Herman
Sensitivity and specificity. Every diagnostic test makes two different mistakes. It can miss someone who has the disease, that's a false negative. Or it can flag someone who's healthy, that's a false positive. Sensitivity is the share of people who truly have the condition that the test correctly catches. Specificity is the share of people who truly don't have it that the test correctly clears.
Corn
So a test with ninety percent sensitivity catches nine out of ten sick people and tells one sick person they're fine. Ninety percent specificity clears nine out of ten healthy people and scares one healthy person into a follow-up.
Herman
Right. And here's the part people miss. Both numbers are calculated within a single disease-status group. Sensitivity only looks at the sick column. Specificity only looks at the healthy column. That's why they're properties of the test itself, and they don't change when the disease gets more or less common in the population. A ninety percent sensitive test is ninety percent sensitive whether you're screening a million people or testing one patient in a specialist clinic.
Corn
That's the setup for the trap that actually matters. Because there's another pair of numbers, positive predictive value and negative predictive value, and those do change with prevalence. They're the ones that answer the question a patient actually asks, which is, I tested positive, what are the odds I really have it.
Herman
And the example that makes this click is brutal. Take a test that's ninety percent sensitive and ninety percent specific. Sounds excellent. In a specialist clinic where half the people being tested actually have the disease, a positive result means you're ninety percent likely to be a true case. Same test, same sensitivity, same specificity, but now you're screening the general population and only one percent of people have the condition. A positive result now means you're eight point three percent likely to be a true case.
Corn
Eight point three. Same test. The only thing that changed is who you pointed it at.
Herman
That's the prevalence paradox. It's why a COVID antigen test behaved differently in a hospital ward than in a random airport screening line. It's why a cancer screen that looks great in a high-risk trial can generate a flood of false alarms in the wild. The test didn't get worse. The crowd got healthier.
Corn
There's a mnemonic I actually like for the two error directions. SnNout. A highly sensitive test, when negative, rules out disease. Because if it's very good at catching true cases, and it says you're clear, you're probably clear. The other one is SpPin. A highly specific test, when positive, rules in disease. Because if it's very good at not flagging healthy people, and it flags you, you're probably not a false alarm.
Herman
And notice neither mnemonic says anything about the other error type. A highly sensitive test can still have terrible specificity. It catches almost all the sick people but also flags half the healthy ones. A highly specific test can have lousy sensitivity. It almost never false-alarms, but it misses a lot of true cases. You cannot tell from one number whether the other one is any good.
Corn
They also trade off against each other. You can move the threshold, the cutoff for what counts as a positive result, and as you catch more true cases you also catch more false alarms. Raising sensitivity lowers specificity. The only way to improve both is a better test, not a different cutoff on the same one.
Herman
And if you want one number that combines both, you look at the area under the ROC curve, the AUC. Point five is no better than flipping a coin. One point zero is perfect. Above point eight is generally considered clinically useful. Below that, limited utility. But AUC hides the threshold question. Two tests can have the same AUC and behave very differently at the cutoff a doctor actually uses.
Corn
So that's the diagnostic pair. The test has two error rates, they're properties of the test, and they don't tell you what a positive result means for you unless you also know how common the disease is in your group.
Herman
Next chunk Daniel asked about, standard deviation, control group size, effect size. These are the how-big-and-how-certain numbers.
Corn
Standard deviation is the one most people half-remember from school and then freeze on. It's just spread. How far individual values typically sit from the group average. A small standard deviation means everyone bunched up near the mean. A large one means the group was all over the place.
Herman
And it matters because an average without a spread is nearly useless. If a drug lowers blood pressure by ten points on average, but the standard deviation is twenty points, that means some patients dropped thirty points and some went up ten. The average is real but it's not the experience of any particular patient. Whereas if the standard deviation is two points, almost everyone got close to that ten point drop.
Corn
Control group size is Daniel's phrase, and it's not a named statistic, but the underlying point is power. How many people do you need in each group to actually detect a real difference if one exists.
Herman
Small samples produce unstable results. Ten people having a bad drug reaction means something very different in a hundred person study than in a hundred thousand person study. Denominators matter. And a study with too few people can't detect a real effect even when it's there. That's a Type II error, concluding there's no effect when one actually exists.
Corn
The convention is you want statistical power of at least eighty percent. That means if there's a real difference, the study has an eighty percent chance of finding it. Below that, a negative result is hard to interpret. It might mean there's no effect, or it might mean the study was too small to see one.
Herman
And here's the thing that connects back to effect size. You cannot calculate how many people you need without first estimating how big the effect is. A huge effect needs a small sample. A tiny effect needs an enormous one. So effect size isn't just a reporting nicety, it's the number that drives study design from the start.
Corn
Effect size itself is the magnitude of a finding, separate from whether it's statistically significant. A result can be statistically significant, meaning unlikely to be chance, and still be so small that no patient would notice or care.
Herman
The most common measure is Cohen's d. It's the difference between two group means expressed in units of the pooled standard deviation. A d of one point zero means the two groups are one standard deviation apart. The famous thresholds are point two small, point five medium, point eight large.
Corn
But those thresholds are a rule of thumb that hardened into dogma. Cohen himself called them arbitrary, a fallback for when a field has no better basis. And different fields have different baselines. Psychology's published average is closer to d of point four. Education uses point two, point four, point six. In clinical medicine, a small effect on Cohen's generic scale can still be a life-saving difference at population scale.
Herman
A d of point five in a medical trial is a big deal. The same d of point five in a personality psychology study sits near the average effect the field routinely publishes. Context beats thresholds every time.
Corn
Which is why the p-value alone is a trap. A paper can report p less than point zero five and you still know nothing about whether the finding is big enough to matter. The p-value tells you the result is unlikely to be noise. The effect size tells you whether it's worth your attention.
Herman
There's also a correction for small samples called Hedges' g. With two groups of eight people each, a d of point five eight drops to about point five five. Once you get above twenty per group, they're nearly identical. It's a small-sample adjustment, and it's one of those details that tells you the authors actually thought about their numbers.
Corn
Now the risk family, which is where the safety and probability language lives. Risk ratio, also called relative risk. It's the proportion with a bad outcome in the treated group divided by the proportion in the control group.
Herman
Above one means the treatment increases risk. Below one means it decreases risk. Exactly one means no change. Worked example, twenty percent of controls develop a bad outcome, twelve percent of treated do. Risk ratio is point six. That's a forty percent relative risk reduction.
Corn
And this is where headlines go to die. Because relative risk reduction sounds enormous and absolute risk reduction is often tiny. A relative risk of point five can mean a drop from fifty percent to twenty-five percent, or from ten percent to five percent, or from one percent to half a percent. Same relative risk. Wildly different real-world magnitudes.
Herman
The ASCOT trial is the classic. Thirty-six percent relative risk reduction in cardiac events. Sounds like a miracle drug. The absolute risk reduction was one point one percent. The drug took a small risk and made it slightly smaller, and the relative framing made it sound like it eliminated a third of all heart attacks.
Corn
Absolute risk reduction is just control risk minus treatment risk. In that example, if the control group had a three percent event rate and the treated group had two percent, the absolute reduction is one percent. Number needed to treat is one hundred divided by that. So NNT of a hundred. You'd treat a hundred people to prevent one event.
Herman
And the FPM refresher Daniel's probably been circling gives the full picture. Relative risk of point eight five, relative risk reduction of fifteen percent, absolute risk reduction of three percent, NNT of thirty-three over four years. For every thirty-three people who take the medicine for four years, one person avoids one of the outcomes.
Corn
Then you look at which outcome. In that example, the composite benefit was driven almost entirely by hospitalization for heart failure. Two point four percent of the three percent absolute reduction was that one outcome. The NNT for that single outcome was forty-two, at twenty-four thousand dollars per person over four years. That's just over a million dollars to prevent one hospitalization.
Herman
Composite outcomes are their own trap. The manufacturer advertises that the drug reduced the likelihood of death, stroke, heart attack, and hospitalization by fifteen percent. It actually did nothing of the sort. It reduced hospitalization, and the other three outcomes were along for the ride.
Corn
Odds ratio is the one that even journal authors get wrong. It's not the same as risk ratio. When the outcome is rare, they're close enough. When the outcome is common, the odds ratio diverges substantially and overstates the effect.
Herman
Multiple papers have documented this as a recognized error. Odds ratios are hard to comprehend directly and are routinely interpreted as equivalent to relative risk. When the outcome is common, that mistake can seriously affect treatment decisions and policy.
Corn
So if you're reading a paper and it reports an odds ratio of two point five, and the outcome happened in forty percent of the control group, you cannot say the treatment more than doubled the risk. The actual risk ratio is smaller. How much smaller depends on the baseline rate.
Herman
The rule of thumb, if the outcome is under about ten percent, odds ratio and risk ratio are close enough that it doesn't matter much. If the outcome is common, demand the actual risk ratio or recalculate from the raw numbers.
Corn
There's also number needed to harm, which is the same logic in reverse. For the Alzheimer's monoclonal antibodies, the number needed to harm for brain swelling was nine. Nine people treated, one case of ARIA edema. For symptomatic brain swelling, it was eighty-six.
Herman
And the benefit side was sobering. Lecanemab produced a one point eight point improvement on a cognitive scale where the minimal clinically important difference is five points on a ninety point scale. So the benefit was real, statistically detectable, and too small for most patients to notice. That's the gap between statistical significance and clinical significance in one drug.
Corn
Minimal clinically important difference is worth naming on its own. Generally a change of at least eight to ten percent in a symptom score before patients actually notice feeling better. Below that, you've got a number that moved but a patient who can't tell.
Herman
Confidence intervals deserve a moment too, because they're the uncertainty band around every one of these numbers. Sampling ten students and finding four men gives a ninety-five percent confidence interval of twelve to seventy-four percent men. Sampling a thousand and finding four hundred gives thirty-seven to forty-three percent. Same point estimate, wildly different precision.
Corn
The confidence interval is the honest version of the finding. A wide interval means the study doesn't know much. A narrow one means it pinned the number down. Any paper that reports a point estimate without a confidence interval is asking you to trust a guess.
Herman
And the reading strategy Daniel's actually going to use. Skip the methods unless you're hunting for a specific detail. Skim the abstract. Read the introduction, the last sentences state the objective. Then results and discussion, first and last paragraphs of the discussion give the highlights. Then limitations.
Corn
Watch the language. Associated with or linked to means correlation, not causation. The ability to derive causation comes from study design and the totality of evidence, not one observational study.
Herman
Check the journal. Predatory journals collect fees without real review. Check the publication date. Check the number of participants. Check who was enrolled, because a trial on forty-year-old men tells you nothing about seventy-year-old women. Check conflicts of interest. If there are no declarations in the paper, assume there could be conflicts among authors.
Corn
Preprints are not peer-reviewed. They're fine for following a field, but they should not be used to support clinical practice or individual health behavior.
Herman
The through-line. Sensitivity and specificity are about the test's two error rates. Standard deviation is spread. Control group size is power. Effect size is magnitude, separate from significance. Risk ratio is relative risk, absolute risk reduction is the real-world number, number needed to treat is how many people you treat to help one. Odds ratio is the impostor that overstates common outcomes.
Corn
The single most counterintuitive fact in the whole bundle is the prevalence paradox. The same ninety-ninety test is trustworthy in a specialist clinic and nearly useless as a population screen. The test didn't change. The crowd did.
Herman
If there's one habit that would fix most misreadings, it's asking which question the number is answering. Is this telling me about the test, the spread, the size of the effect, or the risk? Because each of those is a different question, and the paper will not tell you which one you're holding.
Corn
Daniel's going to be reading these papers with a different set of eyes now. The next time he sees a thirty-six percent reduction, he's going to ask, reduction from what to what.

Hilbert: Point eight.
Herman
What?

Hilbert: The AUC threshold. You said above point eight is clinically useful. The Turkish review says point eight zero or above is generally considered useful. Below that, limited utility. It's point eight zero, not point eight.
Corn
Fair correction. Point eight zero.

Hilbert: I own one of those pulse oximeters. The kind that clips on your finger. Bought it in ninety-eight after my brother-in-law had a scare. It reads oxygen saturation. Sensitivity's fine, specificity's fine, but the display only shows whole numbers. Ninety-five, ninety-six. It never shows ninety-five point five. So you're sitting there watching a number that's rounding itself and you're supposed to feel reassured.
Herman
That's actually a real limitation. The device is averaging and rounding, and the error band on a consumer pulse oximeter is typically plus or minus two percent. So a reading of ninety-five could be ninety-three.

Hilbert: That's what I told him. He said the doctor said it was fine. I said the doctor's reading a better machine. He kept it on the nightstand for six years. Then the battery corroded and he threw it out.
Corn
The rounding point connects to confidence intervals. The machine's giving you a point estimate with no interval. You're trusting a number that's been quietly rounded by firmware.

Hilbert: Firmware. My pulse oximeter doesn't have firmware. It has a battery and a spring. The spring's the part that fails first. After that it clips on but doesn't grip. You'd get a reading of ninety-one and it's just the clip slipping.
Herman
Which is a measurement error, not a clinical change. That's the other thing about these numbers. They assume the measurement itself is reliable. If the device is drifting, the sensitivity and specificity from the validation study don't apply to your unit.

Hilbert: They validated it on a bench in a lab with a calibrated gas mixture. My brother-in-law's was sitting in a drawer with loose batteries for two years. Different instrument.
Corn
The instrument drift point is underrated. A test's published sensitivity is what it did in the validation study under controlled conditions. What it does in your clinic with a tired technician and a machine that hasn't been recalibrated is a different question.

Hilbert: You want the real-world sensitivity, you test it in the real world. Nobody wants to pay for that study. So they quote the lab number and call it a day.
Herman
He's right. That's a known gap between efficacy and effectiveness. Efficacy is what the test does under ideal conditions. Effectiveness is what it does in practice. Most published sensitivity and specificity numbers are efficacy numbers.

Hilbert: My brother-in-law's fine, by the way. It was a panic attack. The oximeter said ninety-one because the clip was loose, and he panicked harder, which made his breathing worse, which made the reading worse. The instrument created the symptom it was measuring.
Corn
That's a feedback loop you don't see in the validation study.

Hilbert: He sold it at a yard sale for four dollars. Told the buyer it worked fine. I didn't say anything.
Herman
The yard sale disclosure is its own study design problem. No blinding, no control group, and one participant with a conflict of interest.
Corn
That's the whole episode in miniature. A number, a missing interval, a measurement error, and someone trusting the point estimate.
Herman
The misconception I'd name is that a statistically significant result is the same as a meaningful one. People see p less than point zero five and think the finding is real, big, and applies to them. The p-value only tells you it's probably not noise. The effect size tells you if it's big enough to notice. The confidence interval tells you how precisely you know it. The absolute risk tells you if it matters in the real world. Four different questions, four different numbers, and the paper only hands you the p-value with a triumphant tone.
Corn
The fix is the question Daniel's now equipped to ask. Significant for whom, measured how, over what baseline, with what interval. If the paper can't answer those, the headline number is doing a lot of unearned work.
Herman
The forward question I'd leave with, there's a gap in the literature Daniel just exposed. Nobody has bundled sensitivity, specificity, effect size, and risk ratio into one layperson's field guide. The material is scattered across diagnostic testing guides, RCT interpretation refreshers, and effect size explainers. Daniel might be the person to write the thing he was looking for.
Corn
That's a project for another day. For now, he can read the results section without flinching, and that was the ask.
Herman
Thanks to Hilbert Flumingtop for producing. This has been My Weird Prompts. Email us at show at my weird prompts dot com, and we'll be back soon.
Corn
Until then, check the denominator.

This episode was generated with AI assistance. Hosts Herman and Corn are AI personalities.