A plain-language guide

A positive test.
Should you panic?

A test that's 99% accurate just said you have a rare disease. Here's the surprising β€” and genuinely reassuring β€” truth about what that number really means, and the one simple habit that keeps you from being fooled by it.

See why not
The surprising truth

A 99%-accurate "yes" can still be wrong β€” usually.

Picture a disease that just 1 in 1,000 people have, and a test that's right 99% of the time. You test positive. Your gut screams "I'm 99% done for." But the real chance you're sick is closer to 9%.

Nothing is broken about the test. The twist comes from something quietly powerful β€” how rare the disease was to begin with. That single fact bends the meaning of your result far more than most people expect. The easiest way to believe it isn't algebra; it's to watch the thing happen to a thousand people at once. So let's do exactly that, and then slowly take the idea apart until it feels obvious.

Try it Β· 1,000 people

Meet 1,000 people. Watch who tests positive.

Every dot is a person. Drag the two dials and watch the crowd sort itself: the few blue dots truly have the disease and got a correct positive, the amber dots are healthy people handed a false alarm. The question that matters isn't "is the test good?" β€” it's "if my dot is coloured, which colour is it likely to be?" The little table underneath keeps a running tally of all four ways the test can land.

1 in 1,000
99%
The same 1,000 people, sorted into four boxes
Actually sickHealthy
Test says + 1true positive 10false alarm
Test says βˆ’ 0missed case 989all clear

Of the 11 people who test positive, only 1 is actually sick.

9%chance you're actually sick after a positive result

So a positive result is overwhelmingly a false alarm. Breathe.

The coloured dots at the top are everyone who tested positive. Notice how the amber false alarms can completely swamp the handful of real cases β€” that gap is the whole idea. Slide "how common" to the right and watch the balance tip back.

Start before the test

The number you already had: the base rate.

Before you took the test β€” before you'd even walked into the clinic β€” there was already an honest answer to "how likely am I to have this?" It's just how common the disease is in people like you. One in a thousand. That starting number has a name: the base rate, or in the language of this idea, the prior. It's your belief before the new evidence arrives.

Most of us quietly throw the base rate away the instant a test result lands. The result feels so vivid, so personal, so now, that the dull background fact β€” "this is rare" β€” stops feeling relevant. That's the mistake. The base rate doesn't disappear when the test beeps. It's the ground the result has to stand on, and when the ground is "almost nobody has this," even a loud, confident beep can't lift you very far.

Here's the intuition pump. Imagine a stadium of 1,000 strangers and a disease that one of them carries. If I point at a random person and say "I bet you don't have it," I'll be right 999 times out of 1,000 β€” with no test at all. That's how strong the prior is. For a test to overturn that starting bet, it has to bring serious evidence to the table. A good test does push your belief upward. But it pushes from the base rate, not from a blank slate β€” and from one-in-a-thousand, even a big push often lands you well short of "definitely."

This is why the very same positive result means different things for different people. A positive on a symptom-free 20-year-old screened "just in case" sits on a tiny base rate. The identical positive on someone with a family history and three matching symptoms sits on a much larger one. Same test, same machine, same printout β€” different priors, different conclusions. The evidence is only ever half the story; the half you bring with you is the other.

What "accurate" really packs in

Two ways to be right, two ways to be wrong.

We've been waving the word "accurate" around as if it were a single dial. In real life it hides two separate skills, because there are two completely different jobs a test has to do β€” and a test can be brilliant at one and lousy at the other.

Sensitivity is how good the test is at catching the sick. Of the people who genuinely have the disease, what fraction does it light up? A test with 99% sensitivity misses just 1 in 100 real cases. When a sick person tests positive, that's a true positive; when the test fails to catch them, that's a false negative β€” a missed case, the scariest kind of error.

Specificity is how good the test is at clearing the healthy. Of the people who don't have the disease, what fraction does it correctly wave through? A test with 99% specificity wrongly alarms 1 in 100 healthy people. A healthy person who's cleared is a true negative; a healthy person who gets flagged anyway is a false positive β€” the false alarm at the heart of this whole story.

Those are the four boxes in the little table you were just dragging: true positive, false positive, false negative, true negative. Every single one of the thousand people lands in exactly one of them. To keep the centerpiece honest and simple, that demo uses one accuracy dial standing in for both skills at once β€” a test that's equally good at catching the sick and clearing the healthy. Real tests usually differ on the two, and the leaflet in the box will quote both numbers separately. But the lesson survives the simplification intact, because the villain of the piece is the false-positive corner, and that corner is fed by the enormous crowd of healthy people. Hold that thought β€” it's the key to the next section.

Why it happens

The disease is rare. That quietly changes everything.

When almost nobody has the disease, the handful of real cases are easy to drown out. A 99% test still gets it wrong for 1 in every 100 healthy people β€” and in a rare disease, there are an enormous number of healthy people. That tiny error rate is small as a percentage, but it's applied to a giant crowd, and a small slice of a giant crowd is still a lot of people.

Out of 1,000 folks where only one is truly sick, that "tiny" 1% error rate produces about ten false alarms. Ten false positives against one real case: even a great test leaves you outnumbered ten to one. So when your dot turns a colour, the odds say it's far more likely to be one of the ten amber ones than the single blue one. The test didn't lie β€” it's just that the pool it's fishing the false alarms out of is vast, while the pool of real cases is a puddle.

This is the lever you felt under your thumb on the prevalence dial. Crank the disease from rare to common and the real cases multiply while the false alarms shrink, until the blue dots finally outnumber the amber ones and a positive becomes genuinely bad news. The exact tipping point depends on the test, but the shape is always the same: rarity inflates false alarms; commonness deflates them. The test's accuracy sets the slope, but the base rate decides which side of the hill you start on.

It's worth saying plainly, because it's the part that surprises people most: you can hold the test perfectly fixed β€” same 99% β€” and the meaning of a positive can swing from "probably nothing" to "almost certainly real" purely by changing how common the disease is. The result on the printout is identical. What it's worth is not.

The rule behind the picture

Meet the formula β€” just this once.

Everything you've watched on the grid is one short rule, written down around 250 years ago by a Presbyterian minister named Thomas Bayes. It looks intimidating in symbols, so we'll show it once, then immediately translate it back into the dots you already understand. The thing it answers is exactly our question: given a positive test, how likely is it that you're actually sick?

That's the whole magic trick, and notice it's nothing more than the readout under the demo, dressed up in symbols. The top of the fraction, P(+ | sick) Γ— P(sick), is the test's catch-rate times the base rate β€” the blue dots, the true positives. The bottom, P(+), is everyone who tests positive β€” the blue dots plus the amber ones. Divide the genuine alarms by all the alarms and you get the only number you actually care about after a scary result: the chance it's real.

People who use this idea give that bottom-of-the-fraction number a job title β€” the posterior: your belief after the evidence. So Bayes' theorem is a recipe with a satisfying shape: take your prior (the base rate), feed in the evidence (the test result and how trustworthy it is), and out comes your posterior (your updated belief). You didn't replace what you knew. You updated it. And updating, not flipping, is the entire spirit of the thing.

Try it Β· stack the evidence

One test nudges. Two tests shove.

A 9% scare isn't a reason to panic β€” but it's absolutely a reason to test again. Here's the beautiful part: today's answer becomes tomorrow's starting point. After one positive, your honest belief is no longer "1 in 1,000"; it's that fresh 9%. So a second positive test doesn't start from scratch β€” it builds on the first. Press the button and watch a couple of independent positives haul the probability across the room.

1 in 1,000
99%

Before any test, about 0.1% of people here actually have the disease. Take a positive test to update that.

The bar fills to your current belief; each notch marks where a positive test left you. The first positive barely moves you off the floor β€” but the second lands on top of the first, and suddenly you've leapt past the halfway line. (Moving a dial resets the chain.)

This is why doctors so rarely diagnose anything off a single screening result, and why a worrying first test usually earns you a second, different one rather than a sentence. Each independent positive multiplies the odds, so the evidence compounds. From a one-in-a-thousand start with 99% tests, one positive leaves you around 9%, a second around 91%, a third past 99% β€” the same machine, the same accuracy, simply asked again. There's a fairness clause, though: the tests have to be genuinely independent. Re-running the exact same flawed test, or one that fails for the same reason every time, just repeats the first answer louder β€” it doesn't add new evidence. Real confidence comes from evidence that could have disagreed and didn't.

The thinking trap

Why your brain falls for it every time.

This blind spot is common enough to have a name: the base-rate fallacy. It's the very human habit of seizing on the vivid new evidence β€” the positive result, the matching description, the alarming headline β€” and forgetting to weigh it against how rare the thing was to begin with. The base rate is boring and abstract; the evidence is specific and emotional. So the boring number gets dropped, and we leap straight to the scary conclusion.

The deepest source of the confusion is that two very different questions sound almost identical, and our brains quietly swap one for the other:

The myth

"It's 99% accurate, so I'm 99% likely to be sick."

This answers a question about the test: when someone is sick, how often does the test catch it? That's a property of the machine β€” and it's genuinely 99%.

The truth

"Given my positive, the chance I'm sick is about 9%."

This answers a question about you: among everyone who tests positive, how many are actually sick? That depends on the base rate β€” and the rare disease drags it way down.

Those two percentages β€” "how often the test is right" and "how likely I am to be sick" β€” feel like the same fact phrased two ways. They are not. One reads down the columns of the table (start with the sick, see what the test does); the other reads along the rows (start with the positives, see who's actually sick). The test's "99%" lives in the columns. Your real worry lives in the rows. Quietly sliding from one to the other is the entire fallacy, and even doctors do it β€” studies repeatedly find that medical professionals, asked this exact kind of question on the spot, often badly overestimate the danger. You're in good, fallible company.

The fix, in one line: whenever an alarm goes off, before you react, ask "how rare was this thing before the alarm?" β€” and let that rarity argue back. The louder the alarm and the rarer the event, the more important that question becomes.

Where this shows up

It isn't just doctors.

The same trap springs anywhere you screen a big crowd for something rare. The rarer the thing, the more of your alarms are false β€” and once you can see the pattern, you start spotting it everywhere.

🩺

Medical screening

Mammograms, genetic panels, mass disease screens β€” a positive on a low-risk person is a reason to look closer, not a diagnosis. It's why follow-up tests exist.

πŸ“§

Spam & fraud filters

Flag the rare bad email or transaction and you face the same balancing act. Tune it too aggressively and the "spam" folder fills with real mail β€” false positives with a cost.

πŸ›‚

Airport & security alerts

Scan millions of harmless bags for the one genuine threat and the overwhelming majority of alarms will, thankfully, be false. The challenge is acting on them without grinding to a halt.

βš–οΈ

DNA & courtroom evidence

"A one-in-a-million match" sounds like certainty β€” but search a database of millions and a few innocent matches become likely. Juries who skip the base rate can convict on a false alarm.

πŸ€–

Machine-learning classifiers

Models that flag rare events β€” disease in a scan, defects on a line, fraud in a feed β€” are judged by precision and recall: data-science names for the very rows and columns of our little table.

πŸ””

Any rare-event alarm

Earthquake warnings, smoke detectors, fault sensors, content moderation β€” when the event is rare, "it went off" is weaker evidence than it feels, and a calm second look pays off.

Notice the recurring tension. Make the alarm more eager and you catch more real cases but drown in false ones; make it more relaxed and you cut the false alarms but start missing the real thing. There's no setting that escapes the trade-off β€” only a choice about which error you can least afford. A smoke detector should err toward false alarms; a spam filter eating your job offer should not. Bayes doesn't make that choice for you, but it tells you, honestly, what each setting will cost.

Run the numbers yourself

The whole thing, on the back of a napkin.

You never need the symbols. The cleanest way to get a Bayes answer right is to do exactly what the demo does: take a round, friendly crowd and just count. Here's our headline case, worked out one plain step at a time β€” and notice there isn't a fraction in sight until the very last line.

Start with a crowd

Take 1,000 people. With the disease at 1 in 1,000, that means about 1 person is genuinely sick and the other 999 are healthy. This split is the base rate, made concrete.

Test the sick

The test catches 99% of real cases, so of our 1 sick person it correctly flags 1 (a true positive) and misses almost nobody. So far: 1 real, correct alarm.

Test the healthy

The test wrongly alarms 1% of healthy people. One percent of 999 is about 10 people β€” ten false alarms, each one a perfectly healthy person handed a scary result.

Count all the alarms

Total positive results: the 1 true alarm plus the 10 false ones makes 11 people staring at a positive. Only one of them is actually sick.

Read off your answer

Your chance of being sick, given a positive, is 1 out of 11 β€” about 9%. Not 99%. The other ten in eleven are false alarms. That's the whole calculation.

Now feel the base rate's grip by changing just one thing. Suppose instead the disease is common β€” say 100 in 1,000, the far end of the demo's dial. Now 100 people are sick, the test correctly flags about 99 of them, and 1% of the 900 healthy people gives only about 9 false alarms. Of the roughly 108 positives, 99 are real: your chance leaps to about 92%. Same test, same 99% accuracy, same arithmetic β€” but a positive now means almost the opposite. Run those exact settings on the first demo and you'll watch the dots agree with the napkin to the dot.

Carry this with you

How to read any alarm, in three moves.

1

Start with the base rate

How common is the thing before any test or alarm? That rarity is your honest starting point β€” don't drop it.

2

Let the result nudge it

A positive pushes your belief up β€” by a lot or a little, depending on how trustworthy the evidence is. Independent results stack.

3

Update, don't replace

You land somewhere between the base rate and certainty β€” and from a rare start, rarely all the way at "definitely."