Measure two things about the same person β and plot them as one dot. Do it for a whole class and a shape appears out of the scatter, whispering how the two are linked.
Start hereWhen you measure two things about the same thing β your height and your shoe size, hours revised and the mark you got β you can plot each pair as a single dot. Stack up enough dots and they stop looking random: a shape appears, and that shape tells you how the two are connected.
That picture is a scatter graph, and the connection it reveals is called correlation. By the end of this page you'll be able to drop dots yourself, watch the relationship form, draw the line that sums it up β and, just as importantly, know exactly what that line is not allowed to claim.
A scatter graph (sometimes called a scatter plot) is a way of showing bivariate data β and that word just means "two variables," two things measured together. The first thing goes along the bottom, on the x-axis; the second thing goes up the side, on the y-axis. Each dot stands for one person, one object, one moment β placed at the spot where its two numbers cross.
Picture your friend Maya: she's 150 cm tall and wears size 4 shoes. On a height-versus-shoe-size graph, Maya isn't a row in a table any more β she's a single dot, sitting at "150 across, 4 up." Plot the whole class the same way and you get a sky full of dots, one per person. Nobody joins them up with a wiggly line, because they're separate people, not steps in a journey. You just let them sit there and you look at the pattern they make together.
That's the quiet magic of a scatter graph: no single dot tells you much, but the crowd of them does. A line graph follows one thing changing over time. A scatter graph asks a different, nosier question β do these two things tend to move together? β and answers it with a shape you can see from across the room.
Reading the graph backwards works too. Spot a dot floating high up the y-axis and far along the x-axis, and you instantly know that this person scored big on both measurements. A dot tucked into the bottom-left corner is small on both. That's why the axes matter so much: deciding what goes across and what goes up is really deciding which two questions you want to ask of every single dot at once.
Tap anywhere on the empty grid to drop a dot. The moment there are two or more dots with some left-to-right spread, a line of best fit appears and a read-out tells you what kind of relationship you've made β and how strong it is. Try to build a rising cloud, then a falling one, then a random mess.
The blue line is the line of best fit β the single straight line that sits as close as it can to all your dots at once. Watch how one stray dot in the corner can tug the whole line toward it.
Notice the read-out is doing two separate jobs. It names the direction β are the dots climbing, falling, or going nowhere? β and it judges the strength β are they hugging the line tightly, or splattered loosely around it? Hold those two questions in your head; the rest of this page is really just the two of them, one at a time.
Whatever cloud of dots you make, its overall lean falls into one of three families. Learning to spot them at a glance is most of the skill.
Here's the careful word in those sentences: tend. Correlation is about a general trend across many dots, never a cast-iron promise about any single one. Some short people have big feet; that's fine. Correlation just means there's an overall pattern in how two things move together β it lives in the crowd, not in any one dot.
Flip back to the demo and try to recreate each of these three pictures with your own dots β then check whether the read-out agrees with your eyes.
Direction is only half the story. Two graphs can both lean uphill, yet one screams its message while the other only mumbles it. That difference is the strength of the correlation, and you read it from how closely the dots hug an imaginary straight line.
When the dots sit almost on a line β a neat, narrow stripe β we call it a strong correlation. Knowing one number lets you guess the other with real confidence. When the dots make the right general lean but spray out in a fat, fuzzy cloud, it's a weak correlation: the trend is real but loose, so any guess you make comes with a big shrug. In between sits moderate β a clear lean with a fair bit of wobble.
Direction asks "which way?" Strength asks "how tightly?" A graph can be strongly positive, weakly negative, or anything in between β the two ideas are completely separate dials.
Statisticians squeeze this whole judgement into a single number called the correlation coefficient, written r. It runs from β1 to +1. A value near +1 means a strong uphill trend, near β1 means a strong downhill one, and near 0 means barely any trend at all. The sign (plus or minus) is the direction; how far it is from zero is the strength. That little r in the play demo's read-out is exactly this number β watch it crawl toward 1 as you line your dots up neatly, and slump toward 0 as you scatter them.
You don't need to calculate r by hand to use it; that's a job computers do happily. What matters at your stage is reading it: see r = 0.9 and picture a tight uphill stripe, see r = β0.3 and picture a loose downhill spray, see r = 0.05 and picture a shapeless blob. The number and the picture are two languages for the very same idea, and the more you flip between them, the faster your eye gets.
Once you can see a trend, it's handy to draw a single straight line that captures it. That's the line of best fit β the line that threads through the middle of the cloud, getting as close as it possibly can to every dot at once. Roughly the same number of dots should sit above it as below it, and it should follow the cloud's lean, not chase any one stray point.
The best part: once that line exists, it becomes a little prediction machine. Got a value for one thing but not the other? Run your finger up from the bottom axis until you hit the line, then across to read off the likely partner value. Below, a real-feeling data set of revision hours versus test marks already has its line of best fit drawn. Slide the hours and watch the line hand you a predicted mark.
The dashed guide does what your finger would: up from the hours, across to the mark. The dots are the real students; the line is our best summary of them.
Reading inside the range of your dots β between the lowest and highest hours people actually revised β is safe and sensible. Stretching the line way beyond your data (say, predicting the mark for 40 hours of revision) is called extrapolation, and it's a gamble: the trend you measured might bend, stop, or break long before you get there. A line of best fit only promises to behave where you have evidence.
Sometimes one dot sits far away from the crowd, breaking ranks with all the others. That loner is an outlier β a value that doesn't fit the general pattern. Maybe it's a genuine surprise (a student who barely revised yet aced the test), or maybe someone simply wrote the number down wrong. Either way, an outlier is worth noticing, because it can yank the line of best fit off course and weaken a perfectly good trend.
Below is a tidy uphill cloud with one stubborn outlier hiding in it. Use the button to pull the outlier out of the calculation and back in again, and watch what it does to the line and to the strength read-out.
See how removing a single misbehaving dot can tighten the whole picture? That's why statisticians always eyeball the scatter before they trust any number.
Outliers aren't villains to be deleted on sight β sometimes they're the most interesting dot on the page. The rule is just: notice them, then think. Ask whether it's a real result or a mistake, and report what you decided. Quietly hiding a dot you don't like is how data goes from honest to dishonest.
Two things rising together does not prove that one is causing the other. A scatter graph can show you that a pattern exists β it can never tell you why.
Here's the classic, friendly trap. Across a year, ice-cream sales and cases of sunburn rise and fall almost in step β plot them and you'd get a lovely strong positive correlation. So does eating ice cream burn your skin? Of course not. A third thing, lurking in the background, is quietly driving both: hot, sunny weather. The sun makes people buy ice cream and the sun burns their skin. The two visible things are linked only because they share a hidden cause.
That hidden third factor has a name β a lurking variable β and hunting for it is a scientist's favourite game. A correlation is an invitation to ask "why might these move together?", and the honest answers include: one really does cause the other, something hidden causes both, or it's pure coincidence. The graph alone can't tell you which.
Scatter graphs aren't a classroom curiosity β they're how people across the world find out whether a hunch is real. Two examples you can picture right now:
Height versus shoe size. Measure everyone in your year group and plot height across, shoe size up. You'll get a clear positive correlation: taller people tend to have bigger feet. It makes sense, too β bigger bodies generally come with bigger feet, so here the connection probably is a genuine one (growing affects both). But notice it's still only a tendency: you'll spot plenty of dots that don't obey, and that's exactly why it's a cloud and not a single neat line.
Revision versus marks. The graph you just played with is one you could make for real. Across a class, more revision hours usually go with higher marks β a positive correlation, often a decent one. Tempting to declare "revision causes good marks"... and revision probably does help. But a careful scientist still pauses: maybe students who revise more also sleep better, or find the subject easier, or were always going to do well. The scatter graph shows the link is there; proving the cause needs a proper, fair experiment, not just a pretty pattern.
That pause β "the pattern is real, but what's behind it?" β is the whole habit of mind a scatter graph is trying to teach you.
This is the single most common mistake people make with data β and now spotting it makes you sharper than most adults reading the news. Two quick checks to lock it in.
Correlation never proves cause. The two things rise together because sunny, hot weather quietly pushes up both ice-cream sales and sunburn. That hidden driver is the lurking variable β and it's why a scatter graph can show a link but never explain it on its own.
Weak doesn't mean nothing, and a trend isn't a guarantee. A weak positive correlation says the lean is real but the dots are scattered, so the link holds on average, not for every person β and even a strong correlation still wouldn't prove that revision caused the marks. Direction, strength, and cause are three different questions.
Measure two things, make each pair one dot, and let the crowd of dots show a shape.
Uphill is positive, downhill is negative, no lean is none β and tight beats fuzzy for strength.
A line of best fit can predict, but a pattern alone never proves one thing causes the other.