Average Stroop effect
Every run taken on this site, drawn as one distribution. No survey, no textbook figure: this is what people actually score here, and you can see how many of them there are before you decide what it is worth.
Every run submitted here
One bar per band of equal width, tallest where most people land. The two conflicts and the two input devices are counted apart, and the chart always says which of the four cuts you are looking at.
Colour and word · Mouse or keyboard
Loading the figures…
- —
- median
- —
- 10th percentile
- —
- 90th percentile
- —
- runs in this cut
Cite this figure
Copy a ready-made line for a post, with the link back to this page.
The scale runs through zero, and the bars to the left of it are real: about one run in 6 comes out negative, meaning the matching trials happened to be slower than the conflicting ones. That is backwards, which is exactly why it is worth drawing — it is the size of the noise, measured on the people who were not trying to produce it.
How this is measured
Seven decisions stand between a run and a bar on that chart. None of them is neutral, so each one is written down — including the two that decide what this page refuses to show.
Why the number is a difference
It is your average time on the conflicting trials minus your average on the matching ones. Subtracting removes how quick you happen to be in general — your screen, your mouse, your morning — and leaves only what the conflict cost you. That is why a quick person and a slow one can score exactly the same interference, and it is the whole reason this, rather than raw speed, is the number worth drawing.
Why your second run lands somewhere else
A difference of two averages carries the noise of both, and a run of 20 trials does not have many of either. We simulated it before this page existed: one person taking the test twice lands about 64 ms apart, while the whole spread between different people is about 40 ms. Read that pairing carefully — the wobble is larger than the difference it is supposed to measure, so roughly three quarters of your number describes this particular run rather than you. Running it again is the honest way to see this for yourself.
Why there is no leaderboard, and no percentile of your own number
Both would be easy to build and both would be fiction. Ranking people by a number that is three quarters run-to-run luck sorts them by luck; printing you a percentile of it prints you your own noise. The usual fix — make the run longer — does not rescue it either: it would take something like 160 trials before the score described the person, which is several times longer than anyone would sit through. A distribution survives all of this, because noise that is in every run equally does not move the shape of the population. So the shape is what we publish, and where this run landed in it, and nothing about where you stand.
Why the two conflicts are two charts
Reading a colour word is a lifetime of practice fighting an answer you have to give deliberately. Judging which of two digits is larger while one of them is physically bigger is the same shape of conflict, but not the same size of one. Putting both on one chart would prove only which task is harder, and the number underneath would describe nobody.
Why a finger and a mouse are counted apart
A touchscreen has its own delay before it reports a touch, and so does a mouse. Subtracting two averages cancels some of that — but not all of it, because the delay does not land equally on trials you answer fast and trials you have to fight. That gap sits between two visitors to this same page, one on a phone and one on a laptop, so splitting by device is the honest cut. Which device a run used is read from the events that produced the answers, never guessed from the screen.
What is thrown away, and what is only flagged
A difference is the one kind of score that improves when you stop trying: answer fast and at random and both averages collapse towards each other, and the gap looks excellent. So every trial is sent, wrong ones included, and a run whose accuracy is too low to have been read at all is kept, marked, and left out of every figure on this page. It is marked rather than refused because a distracted person and a script look identical from a handful of trials, and refusing would only tell whoever wrote the script what to fix. A run with too few correct answers of one kind is refused outright and stored nowhere — there is no average to subtract.
Why a cut sometimes shows nothing at all
Below a fixed number of runs there is no chart and no figure here — not a rounded one, not a hedged one. A median of nine runs is noise wearing the costume of a measurement, and a thin slice is also where one person’s result stops being anonymous. The page says how many it has instead.
Other distributions on this site: Average reaction time
Now put your own number on it
The test runs in your browser and takes about a minute. Submit the run and it joins the chart above — anonymously, with nothing attached to it that could point at you.