In 1935, John Ridley Stroop described a deceptively simple experiment. Participants were asked to name the colour of the ink in which words were printed. When the word itself named a conflicting colour — for example, the word RED printed in blue — responses became substantially slower than when colour naming was performed without that competing information.[1]
That difference became known as the Stroop interference effect. More than 90 years later, variants of the Stroop task remain widely used in experimental psychology, cognitive assessment and neuropsychology.
Its longevity comes partly from the simplicity of the task. But interpreting Stroop performance is less simple. Performance reflects several interacting cognitive processes, there is no single universally administered or scored “Stroop test”, and seemingly minor differences in administration can affect the result.
Those characteristics also make the Stroop task particularly well suited to digital administration: stimulus presentation can be standardised, individual responses timed precisely, and the information generated during each trial preserved for later review.
What produces the Stroop effect?
Reading is a highly practised process in literate adults. When a familiar colour word appears, its meaning is processed rapidly even when reading the word is irrelevant to the task.
Colour naming requires the participant to attend instead to another characteristic of the stimulus: its ink colour. When the word and colour conflict, the irrelevant word information competes with the required response.
The participant therefore needs to maintain the task goal, select the relevant stimulus dimension and resolve interference from the competing response. The additional response time — and often increased error rate — under incongruent conditions is the Stroop effect.[1,2]
For this reason, Stroop performance is frequently used as an index of selective attention, cognitive control and response inhibition.
It should not, however, be considered a pure measure of any one of these processes. Performance can also be influenced by baseline processing speed, reading ability, colour perception, language, age, education and familiarity with the task.[2,3]
What does the Stroop test tell us about the brain?
The Stroop task is often associated with frontal-lobe function, but describing it simply as a “frontal lobe test” overstates its anatomical specificity.
Functional imaging studies demonstrate activity across a distributed network during interference tasks. The anterior cingulate cortex and regions of the prefrontal cortex have received particular attention, with evidence supporting roles in monitoring conflict and implementing cognitive control.[4,5]
Other cortical and subcortical regions also contribute.
This makes the Stroop task useful for probing cognitive control, but poorly suited to anatomical localisation by itself. An abnormal result does not identify a particular brain region or establish a particular diagnosis.
Why is Stroop testing used clinically?
Impaired executive control and slowed information processing occur across numerous neurological and cognitive disorders. Stroop-type tasks have consequently been incorporated into neuropsychological assessment in conditions including traumatic brain injury, stroke, Parkinson’s disease, multiple sclerosis and cognitive impairment.
Its value is not that it identifies any of these conditions specifically. Rather, it places a controlled cognitive demand on the person being assessed.
Someone may converse fluently and appropriately but have considerably greater difficulty when required to suppress a dominant response, select between competing information and repeatedly maintain that selection.
Poor Stroop performance can therefore provide evidence of difficulty managing interference or maintaining an appropriate response set, but the finding needs to be interpreted alongside the history, examination and other cognitive measures.
There isn’t one Stroop score
One complication when discussing “the Stroop test” is that numerous versions exist.
Some use cards containing columns of stimuli and measure the time required to complete each condition. Others present stimuli individually and record response time for every trial. Some compare word reading, colour naming and colour-word interference conditions, while others focus on congruent and incongruent stimuli.
Scoring methods consequently differ.
Depending on the paradigm, relevant outcomes can include completion time, reaction time, accuracy, errors and corrected errors, as well as interference scores derived from differences between conditions.[2,3]
This means a score obtained using one Stroop paradigm should not automatically be treated as equivalent to a score from another.
The protocol, stimulus presentation and scoring method are part of the measurement.
Where traditional administration becomes difficult
One of the strengths of Stroop testing is that it can be remarkably simple to administer. Printed stimuli and a stopwatch may be all that is required.
That simplicity also creates a measurement problem.
With verbal responses, the examiner may need to listen for correctness, identify hesitations or self-corrections, record errors and operate a timer simultaneously. In many card-based versions, timing is available for an entire condition rather than for each individual response.
A participant who pauses for several seconds on two difficult stimuli before responding quickly to the remainder may therefore produce the same overall completion time as someone whose responses are consistently slower.
The final number can hide the pattern that produced it.
Manual administration also makes some responses difficult to classify in real time. Was there a hesitation before the answer? Did the participant begin to give the wrong response and self-correct? Was a response premature? Was an apparent error simply unclear speech?
Capturing those features while simultaneously administering the test is difficult.
Capturing Stroop performance one trial at a time
Stroop Test, part of our Neuro Tools suite, takes a trial-level approach.
Instead of presenting a complete card and timing the condition as a whole, stimuli are presented individually on an iPhone or iPad. This allows the app to capture the response to each stimulus separately.
In Verbal mode, the participant names the ink colour aloud. The app detects voice onset and records a short audio sample for each trial before automatically advancing to the next stimulus.
This means the device does not need to be passed backwards and forwards during testing and the examiner does not need to operate a stopwatch for each response.
Most importantly, the original response is retained.
After the test, individual trials can be reviewed and the recorded audio replayed. The examiner can confirm or change the response classification and flag responses showing:
- hesitation
- freezing
- self-correction
- premature response
- omission
The purpose is not to replace judgement with automatic scoring. It is to give the examiner better information on which to make that judgement.
From individual responses to an interference result
Once the trials have been reviewed, the app calculates the relevant reaction-time, accuracy and interference measures automatically.
That removes another potential source of inconsistency: manually transferring results, calculating averages or deriving an interference score after the test.
It also preserves the information behind the summary result.
A mean reaction time, for example, can be considered alongside the individual responses from which it was calculated. Accuracy can be examined alongside speed. Unusual responses can be revisited rather than relying on what the examiner remembers hearing during the test.
This distinction is important because Stroop performance should rarely be reduced to a single number.
A fast result accompanied by frequent errors represents something different from a similarly fast and accurate performance. Likewise, a slower response pattern may reflect greater interference, but it may also be influenced by general processing speed, language, visual factors, speech production or other aspects of performance.
Digital measurement provides more precise data; it does not remove the need to interpret those data.
Verbal and visual response modes
Verbal colour naming most closely demonstrates the classic interference between the written word and the required colour response, but it is not always the most practical response method.
Stroop Test therefore provides both Verbal and Visual response modes, allowing the task to be administered in different ways while retaining standardised stimulus presentation and automated timing.
Practice, Brief and Standard test options allow the amount of testing to be matched to the purpose and setting.
Whichever mode is used, the same principle applies: the test presents a controlled interference task and records performance consistently rather than relying on manual timing and contemporaneous note-taking.
What digitisation doesn’t solve
Putting the Stroop task on an iPhone or iPad does not make interpretation automatic.
A digital Stroop task should not be assumed to be interchangeable with every established paper-based version. Stimulus characteristics, presentation method, response mode, number of trials and scoring algorithm can all influence performance.
Repeat testing also requires caution. Familiarity and practice effects may affect subsequent performance, particularly when the same task is administered repeatedly.[2]
A change between two assessments therefore cannot automatically be attributed to improvement or deterioration in cognitive function.
What digitisation can improve is how consistently the task is presented and how accurately the resulting performance is captured.
The stimulus can be presented in the same way each time. Response onset can be measured without relying on a handheld stopwatch. Individual responses can be retained. Questionable trials can be reviewed. Summary measures can be calculated consistently from the underlying data.
Those are measurement improvements, not automated clinical conclusions.
Making a simple test easier to administer well
The Stroop task has endured because it places a substantial cognitive demand inside an extremely simple instruction: ignore what the word says and report the colour instead.
What makes the result useful is not simply whether someone can do that. It is the pattern of speed, accuracy, errors and interference produced when they try.
The Stroop Test app is designed to make that pattern easier to capture.
It handles stimulus presentation and timing, preserves individual verbal responses for review, provides structured classification of unusual responses and calculates the resulting performance measures automatically.
That leaves the examiner with the part that should remain theirs: reviewing the responses, considering the pattern of performance and deciding what the findings mean in the context of the broader assessment.
Learn more about Stroop Test →
References
-
Stroop JR. Studies of interference in serial verbal reactions. Journal of Experimental Psychology. 1935;18(6):643–662. doi:10.1037/h0054651
-
MacLeod CM. Half a century of research on the Stroop effect: an integrative review. Psychological Bulletin. 1991;109(2):163–203. doi:10.1037/0033-2909.109.2.163
-
Scarpina F, Tagini S. The Stroop Color and Word Test. Frontiers in Psychology. 2017;8:557. doi:10.3389/fpsyg.2017.00557
-
Botvinick MM, Braver TS, Barch DM, Carter CS, Cohen JD. Conflict monitoring and cognitive control. Psychological Review. 2001;108(3):624–652. doi:10.1037/0033-295X.108.3.624
-
MacDonald AW III, Cohen JD, Stenger VA, Carter CS. Dissociating the role of the dorsolateral prefrontal and anterior cingulate cortex in cognitive control. Science. 2000;288(5472):1835–1838. doi:10.1126/science.288.5472.1835
This article is general clinical education and isn't a substitute for formal training or guidelines. Always interpret bedside tests in the context of the full examination.