The Stroop test is a family of timed color-word interference measures, not one instrument: Golden, Victoria, Comalli and Kaplan, Trenerry SNST, and the D-KEFS Color-Word Interference Test each use different scores and norms. Neuropsychologists and psychologists use them to assess performance under response conflict. A score is uninterpretable without its version and norm source. This page covers how to document Stroop results, with a fictional sample.
Psychologists and neuropsychologists select and interpret it; trained psychometrists may administer and score under supervision; purchaser qualification level B or C depending on version and seller
Neuropsychologists and psychologists, referring physicians, ADHD and dementia evaluators, schools and disability reviewers, courts and attorneys, payers reviewing testing claims
4 to 10 chart lines or a short report paragraph · administration about 4 to 10 minutes per version, plus scoring and integration
Timed color-word interference test family (fixed-time counts or fixed-card completion times, depending on version)
Neuropsychological and psychoeducational batteries, ADHD and executive function evaluations, dementia workups, medicolegal, driving fitness, and fitness-for-duty assessments
Version-specific products from Stoelting, PAR, and Pearson plus academic versions; no authority mandates any version by name; described here for documentation purposes, no test content reproduced
The Stroop test is any of several standardized clinical instruments built on the interference effect J. R. Stroop published in 1935: reading a color word is more automatic than naming the ink color it is printed in, so naming the ink color of an incongruent color word costs time and errors. The 1935 article appears to be in the US public domain, but every modern clinical packaging is proprietary. The versions in current clinical use differ in structure, raw metric, and rights holder: the Golden Stroop Color and Word Test, whose current edition is the Normative Update (Golden, Freshwater, Syzdek, and Ailes; Stoelting lists a 2022 publish date, adult and children's versions, with PAR and WPS as distributors); the Victoria modification (Regard, 1981), a brief three-card version sold through the University of Victoria Psychology Clinic to verified professionals; the Comalli and Kaplan lineage (Comalli, Wapner, and Werner, 1962), which has no current commercial kit; the Stroop Neuropsychological Screening Test (Trenerry and colleagues, 1989; PAR); and the D-KEFS Color-Word Interference Test (Pearson, 2001), joined in July 2025 by the separately normed, digital-only D-KEFS Advanced. As of August 2026 no successor edition has been announced for the Golden, Victoria, or SNST products.
The load-bearing fact for documentation is that there is no such thing as a generic Stroop score. Golden and the SNST count items completed in a fixed time; Victoria, Comalli and Kaplan, and the D-KEFS time completion of a fixed display; derived interference is a predicted-score residual in Golden, a ratio or difference in Victoria and legacy versions, and an age-scaled contrast score in the D-KEFS. A chart entry of "Stroop = 42, impaired" could be a count, a time in seconds, a T score, a scaled score, or a derived index, and nothing in the number says which. That is why a defensible note names the version, edition, publisher, norm source, and unit for every score it reports. The D-KEFS report page covers the full nine-test battery this family's four-condition member belongs to, and the Trail Making Test page covers the set-shifting task it is most often paired with.
Neuropsychologists and clinical psychologists administer a Stroop version inside most adult executive function batteries: ADHD and executive complaints, dementia and mild cognitive impairment workups, traumatic brain injury, epilepsy and movement disorder evaluations, and medicolegal or fitness-for-duty questions where response control matters. School and pediatric psychologists reach for the child-normed options (the Golden children's version from age 5, the D-KEFS from age 8, or the related NEPSY-II inhibition subtest covered on the NEPSY-II page). Screening-oriented settings sometimes use the SNST for a brief adult check. The documentation reader is rarely just the author: the next evaluator needs the version and norms to test for change, payers reviewing 96130 to 96139 claims need the medical-necessity chain, and in forensic work the version and norm choice is routinely cross-examined. Results land in a neuropsychological report or feed an ADHD evaluation report; the RBANS page covers the screening battery it sometimes accompanies.
No US, Canadian, or Australian authority prescribes a Stroop documentation format, and none requires a specific version. What survives review is a record that pins the result to one named instrument, one named norm system, and explicit units, then keeps the interpretation inside what a timed interference task can support. Each element below carries the pitfall that most often undermines it.
Version, edition, and publisher. Name the exact product: for example "Stroop Color and Word Test Normative Update, adult version (Stoelting)" or "D-KEFS Color-Word Interference Test (Pearson, 2001 norms)", never "the Stroop." For D-KEFS work, state whether scores come from the 2001 battery or the separately normed D-KEFS Advanced (2025). For a Comalli or Kaplan administration, cite the actual protocol and publication followed, because no current commercial product defines one. Pitfall: "Stroop" with no version. Fixed-time counts and fixed-card completion times are different scales, so an unnamed score cannot be interpreted, compared, or defended later.
Norm source and corrections applied. Name the normative set actually used (publisher norms; Norman and colleagues' 2011 HNRC norms; the Revised Comprehensive Norms; Mayo's Older Americans studies; Troyer's 2006 Victoria data; a published regression source for the SNST) and which demographic corrections it applies: age only, age and education, or fuller demographic models. Say "Norman et al. 2011" rather than the ambiguous "Heaton norms." Pitfall: A T score or scaled score reported as if demographically neutral. What was corrected depends entirely on the norm system, and mismatched norms are a standard cross-examination target.
Raw scores with their units, by condition. Report each administered condition with its unit: items completed in the interval (Golden, SNST) or seconds to completion (Victoria, Comalli and Kaplan, D-KEFS), plus the standardized metric (T score, scaled score, percentile) for each. Keep baseline word-reading and color-naming results in the record; they determine what the interference condition can mean. Pitfall: A bare number. "Stroop 42" could be a count, a time, or a standardized score, and a low color-word result means something different over intact versus slowed baselines.
Derived interference score, named. State which derived index was computed: Golden's predicted-score residual, a Victoria-style ratio or difference, or a D-KEFS contrast score, and report it beside the condition-level scores rather than instead of them. Keep the computation in the authorized manual or scoring software; the note needs the name and the standardized result, not the formula. Pitfall: Letting a derived score do decisional work alone. Difference and contrast scores compound the error of both components; in the D-KEFS reliability analysis none of the 51 contrast coefficients reached 0.70.
Errors, self-corrections, and observations. Record uncorrected errors and self-corrections separately where the version scores them, and describe the performance: impulsive starts, loss of set, slowed articulation, repeated checking, fatigue, frustration, strategy statements. A slow-and-accurate record and a fast-and-error-prone record can produce similar times with different meanings. Pitfall: Time or count charted as the whole result. Victoria and D-KEFS interpretation can turn on the error pattern, and the observations are what make a low score explainable.
Validity conditions, screened and stated. Document the color-vision check before timing (congenital deficiency affects up to about 8 percent of males), corrected acuity, reading fluency and history (the paradigm presumes automatized reading), language of administration, dominance, and interpreter status, speech-motor factors, effort indicators, and any deviation from standard administration. Close with a validity conclusion: interpretable, interpretable with caution, or not interpretable. Pitfall: A color-vision, reading, language, or motor-speech barrier scored as executive impairment. Document the barrier and use a measure that does not depend on the impaired function; do not invent an adjusted score.
Interpretation inside the battery. Describe the finding as performance under color-word conflict relative to the person's own baselines, state what it does not establish (no version diagnoses ADHD or dementia by itself), and integrate it with history, other executive and attention measures, and validity evidence. For serial testing, compare only within the same version and norm system, with prior exposure and interval documented. Pitfall: "Stroop confirms ADHD," or decline calculated between a Golden count and a D-KEFS time. Group-level effects do not support individual diagnosis, and there is no valid cross-version conversion.
STROOP TEST DOCUMENTATION BLOCK Date: [ ] Setting: [ ] Administered by: [ ] Interpreted by: [ ] Referral question / battery context: [ ] Version + edition + publisher: [Golden SCWT Normative Update (Stoelting) / Victoria / Comalli or Kaplan protocol (cite source) / SNST (PAR) / D-KEFS Color-Word Interference (Pearson; 2001 battery or Advanced 2025)] Norm source + corrections: [publisher / Norman 2011 HNRC / RCN / MOANS / Troyer 2006 / other; age / age+education / demographic model] Language + administration: [language, dominance, interpreter, paper or digital, deviations from standard procedure] Results (each condition, with units): Word reading: [raw + unit] [std score + metric] Color naming: [raw + unit] [std score + metric] Color-word / inhibition: [raw + unit] [std score + metric] Switching condition (D-KEFS): [raw + unit] [std score + metric] Derived interference index: [predicted residual / ratio / difference / contrast] Result: [ ] Errors: [uncorrected] Self-corrections: [ ] Observations: [ ] Validity: [color-vision check + method, acuity, reading fluency, language, speech-motor, effort] Conclusion: [interpretable / with caution / not interpretable: reason] Interpretation: [performance under conflict vs own baselines; what it does not establish; convergence with other measures] Prior testing: [version, date, norms; same-version change only] Plan / integration: [ ] Clinician signature / credentials: Date:
Free to use and share, no signup. The PDF includes a one-page cheat sheet with element-by-element pitfalls and a pre-sign checklist; the DOCX is the blank documentation block, ready to adapt. Neither reproduces stimuli, card layouts, scoring formulas, or norm values.
Scenario: a Golden Stroop administration inside an adult ADHD evaluation, with intact reading and a normal color-vision check, documented at the level a records reviewer or later evaluator needs. All details are fictional.
Patient: J.T., 28 · Setting: Outpatient neuropsychology clinic, ADHD evaluation · Clinician: R. Okafor, PhD · Note date: 08/14/2026
Measure and administration: Stroop Color and Word Test Normative Update, adult version (Stoelting), administered in English, paper format, by this examiner as part of a four-hour attention and executive function battery. English is the patient's first and dominant language; he reads for work daily and reported no reading history. Corrected vision worn; he named all test colors accurately on the pre-timing check, with no reported color-vision history. Standard administration, no deviations. Scoring used the publisher's adult age-and-education norms; T scores, mean 50, SD 10.
Results: Word reading raw 96 items in the interval, T 49. Color naming raw 71, T 46. Color-word raw 34, T 37. Interference score (predicted-score residual method) T 41: observed color-word performance modestly below the level predicted from his own reading and naming rates. Two self-corrections, no uncorrected errors. He started quickly on each card, slowed visibly midway through the color-word condition, and twice paused with a whispered restart; effort was engaged throughout and embedded validity indicators across the battery were unremarkable.
Validity: Standard administration in the dominant language with intact baselines, an accurate pre-timing color check, and adequate corrected acuity: result interpretable. Baseline reading and naming within expected limits, so the interference result reflects performance under response conflict rather than a reading or naming bottleneck.
Interpretation: Mildly reduced efficiency when a dominant reading response must be suppressed, relative to his own intact baselines. This is a nonspecific finding: it is consistent with the referral concern but does not establish ADHD or any diagnosis by itself, and no driving, capacity, or fitness conclusion follows from it. It converges with reduced performance on the other conflict-loaded task in the battery and with collateral ratings of distractibility, while sustained attention and processing speed measures were within expected limits.
Integration and comparison note: Findings integrated in the evaluation summary with developmental history, symptom ratings, collateral report, and the remaining battery. A 2021 outside report listed "Stroop within normal limits" without version, norms, or scores; that result is treated as version and norms unspecified, not interpretable for change, and no decline or improvement is calculated against it. Any retest should use this same version and norm system with the interval and prior exposure documented.
This sample is fictional and for educational purposes. It does not describe a real patient or record; the scores, dates, and details are invented to show documentation structure and are not clinical guidance. No test stimuli, formulas, or norm-table values are reproduced.
Writing these after every session? BastionGPT drafts complete notes from bullets, dictation, or a transcript.
Generate a note from bulletsUnited States: the requirement is a defensible rationale, not a brand. Psychological and neuropsychological testing is billed under the 96130 to 96139 code family (with 96136 and 96137 for administration and scoring, and 96146 for a single automated instrument), and Medicare contractor policy (for example LCD L34646 and its billing articles) requires documented medical necessity: the diagnostic question, why testing was needed, why these tests were selected given education, culture, sensory and physical status, fatigue, and premorbid function, the tests given, the time, and integrated findings (PAYER POLICY). No CMS policy names a Stroop version, so the chart must show why this version and norm system fit this examinee, which is exactly what the version-and-norms line provides. For ADHD referrals, the American Academy of Pediatrics' 2019 guideline states that neuropsychological testing generally does not improve diagnostic accuracy for ADHD, useful as that testing is for profiling strengths and weaknesses, so a Stroop result is adjunctive evidence, never the diagnosis (CONVENTION). In forensic and disability work, APA forensic guidelines and test-security standards make the version, edition currency, norm fit, confound handling, and derived-score restraint foreseeable cross-examination targets (CONVENTION). Driving and return-to-duty evaluations use interference tasks as supporting evidence only; no US authority sets a Stroop cutoff for either decision.
Canada and Australia change the funding frame, not the documentation standard. Canadian provincial plans generally do not pay for outpatient psychological testing: the Ontario Psychological Association notes OHIP does not cover private psychological services, and access runs through hospital programs, auto insurers, workers' compensation, or private benefits, none of which names a Stroop version (PAYER POLICY, provincial). British Columbia's driver-fitness guidance treats office cognitive screening as informative while recognizing that functional driving assessment may better determine capability (REGULATORY GUIDANCE). In Australia, psychometric and neuropsychological testing sits outside the Better Access rebate structure, so Stroop-inclusive assessments are typically private, NDIS, or privately insured, and no MBS item names an instrument (PAYER POLICY); Austroads' Assessing Fitness to Drive standard is condition- and function-based and mandates no cognitive test by name (REGULATORY STANDARD). The evidence base argues for the same restraint everywhere: the apparent ADHD interference deficit moves from about 0.24 to about 1.11 standardized units depending on how interference is computed, meta-analytic work on the paradigm rejects any single-mechanism reading of the effect, and reduced interference performance should be written as consistent-with, never diagnostic-of (CONVENTION).
Rights and test security reward the same precision as psychometrics. The 1935 Stroop article is a public-domain scientific finding, but the clinical packagings are not: Stoelting publishes the Golden Normative Update (PAR and WPS distribute it, and seller qualification levels differ, Level B at Stoelting versus Level C listings elsewhere, so state the seller when qualification matters); the University of Victoria Psychology Clinic sells the Victoria materials, by arrangement with the copyright holders, only to professionals with verified qualifications; PAR publishes the SNST; Pearson publishes the D-KEFS and D-KEFS Advanced. Under APA Ethics Standards 9.04 and 9.11, the patient's responses, scores, and your interpretation are test data and belong in the record; stimulus cards, record-form layouts, administration scripts, scoring constants, and conversion tables are test materials and stay out of the chart, the patient portal, and any public page, with no blanket web or EHR carve-out from any of these publishers. A practice that wants a web-based or in-EHR administration needs a written, product-specific arrangement, not an assumption. The Stroop Color and Word Test, the Stroop Neuropsychological Screening Test, and the D-KEFS are products and trademarks of their respective publishers (Stoelting Co., PAR, Inc., and NCS Pearson, Inc.). BastionGPT is not affiliated with, or endorsed by, any of these publishers. This page reproduces no test items, stimuli, norms, or scoring materials.
The numbers behind these errors are specific. In the 2008 reliability analysis of 51 D-KEFS contrast coefficients, none reached 0.70 (mean 0.27), and in clinical patients the inhibition/switching condition was not reliably harder than inhibition (57.1 percent showed an atypical pattern); the 2007 ADHD meta-analysis found the interference effect ranged from about 0.24 with difference scores to about 1.11 with time-per-item methods; and congenital color-vision deficiency affects up to about 8 percent of males. The BastionGPT Clinical Advisory Board sees the same errors most often in Stroop documentation reviews:
BastionGPT is specifically trained, tuned, and clinically tested on psychological and neuropsychological evaluation reports.
See how clinicians use it day to day on the AI therapy notes page.
Many BastionGPT users report saving more than 90 minutes per day on documentation.
HIPAA-compliant with a signed BAA on every plan. Your data is never used to train models. BastionGPT drafts, you review and sign.
It depends entirely on the version. The Golden test and the SNST count how many items were completed in a fixed time; Victoria, Comalli and Kaplan, and the D-KEFS Color-Word Interference Test record seconds to complete a fixed display, so a raw "42" is a count on one instrument and a time on another. Raw scores convert to the version's standardized metric (T scores for Golden, scaled scores for the D-KEFS), and a separate derived interference index estimates the cost of naming ink colors against an incongruent word: Golden predicts expected color-word performance from the person's own reading and naming speeds and scores the difference, Victoria-style approaches use ratios or differences, and the D-KEFS uses contrast scores. A defensible note reports the condition scores with units, names the derived index, and names the norm system, because none of these numbers is interpretable without that frame.
No universal cutoff exists, and any page or report that states one without naming a version and norm set is wrong on its face. Impairment ranges depend on the version, the raw metric, the normative sample, and the corrections that sample applies: publisher norms for the Golden Normative Update incorporate age and education, Norman and colleagues' 2011 HNRC norms apply fuller demographic models, Mayo's older-adult studies adjust for age and estimated ability, Victoria and SNST research norms have their own samples and variables, and standard D-KEFS scores correct for age only. The same raw performance can sit at different standardized levels under different systems, which is why the note names the system rather than asserting a free-floating threshold, and why consumer websites' millisecond bands for browser Stroop tasks have no clinical standing.
Choose by referral mix, age range, and norm fit, then document whichever you chose by name. The Golden Normative Update (Stoelting; adult and children's versions from age 5) is a brief, inexpensive standalone with current age-and-education norms; note that seller qualification levels differ, Level B at Stoelting versus Level C on some distributor listings. The Victoria modification is the shortest fixed-card option, sold through the University of Victoria to verified professionals, with norms in the research literature rather than a publisher manual. The SNST (PAR) is a brief adult screen with two broad publisher age bands and newer independent regression norms. The D-KEFS Color-Word Interference Test (Pearson) adds an inhibition/switching condition and formal error scoring inside a full executive battery, with the digital-only, separately normed D-KEFS Advanced available since 2025. A Comalli or Kaplan administration has no current commercial kit, so cite the protocol and norms you actually followed.
Not quantitatively. Golden scores are counts in a fixed interval with a predicted-residual interference index on T-score norms; D-KEFS scores are completion times on age-scaled norms with contrast-score architecture, and the D-KEFS Advanced adds a third, separate norm system. A records reviewer can describe a shared qualitative pattern, for example that both evaluations showed reduced efficiency under color-word conflict, but calculating decline, converting scores, or comparing percentile movement across versions has no psychometric basis. The defensible language for a prior result reported without version or norms is "version and norms unspecified, not interpretable for change." For what the full nine-test battery's scores mean, see the D-KEFS report write-up guide.
No and no. US neuropsychological testing coverage (the 96130 to 96139 family) turns on documented medical necessity and clinician-justified test selection; no CMS policy names a Stroop version (PAYER POLICY). Canadian provincial plans generally do not fund outpatient psychological testing, and Australian Medicare's Better Access structure does not rebate psychometric testing, so access runs through hospitals, insurers, NDIS, or private pay, again with no named instrument (PAYER POLICY). On diagnosis, the American Academy of Pediatrics' ADHD guideline states that neuropsychological testing generally does not improve diagnostic accuracy, and the meta-analytic ADHD interference effect swings from about 0.24 to about 1.11 depending on the computation method, so a low interference score is written as consistent with executive inefficiency inside a fuller formulation, never as the diagnosis (CONVENTION).
Screen first, then let the screen decide what the scores can mean. Check color discrimination on the actual test colors before timing (formal plates where history or behavior raises doubt) and record the method and result; congenital red-green deficiency affects up to about 8 percent of males, and an examinee who cannot reliably discriminate the hues has not shown executive impairment by failing a color condition. Reading must be automatic for the paradigm to work, so document reading history and baseline word-reading performance, and treat reduced interference over effortful reading as a reading finding, not intact inhibition. Record the language of administration, dominance, proficiency, and interpreter status; an improvised translation is a nonstandard test without norms. In each case, chart the barrier, mark the attempt interpretable with caution or not interpretable, and assess the construct with a measure that does not depend on the impaired function.
Not with any published clinical version's materials. Under APA Ethics Standards 9.04 and 9.11, the patient's responses, scores, and your interpretation are test data and belong in the record, while stimulus cards, record-form layouts, administration scripts, scoring constants, and conversion tables are test materials: publishers treat them as protected, security-sensitive content, and none of Stoelting, PAR, or Pearson publishes a free-web or EHR reproduction carve-out. PAR states that individual Golden items are not available for licensing, and digital administration runs through written, product-specific arrangements such as Pearson's platforms. What a practice may build freely is an original Stroop-like task with independently created stimuli and its own validation, described as such, never presented or scored as the Golden, Victoria, SNST, or D-KEFS. The same boundary governs this page and its templates: structure in words, no stimuli, no norms.
Yes. Give it the facts (version, edition, publisher, norm source, condition scores with units, derived index, errors and self-corrections, validity screens, observations, and prior testing) and it drafts the documentation block or report paragraph: version and norms named, units stated, validity conclusion drawn, and interpretation kept to what an interference task supports, ready for your review. It can also cross-check a finished note for the errors reviewers flag: a score with no version, an unnamed norm source, a cross-version comparison, or missing confound documentation. BastionGPT is HIPAA-compliant with a signed BAA on every plan, and your data is never used to train models.
The instrument facts and compliance claims on this page trace to these sources, last verified August 2026:
Educational content, not legal or billing advice. Sample notes are fictional. Follow your organization's policies and your board, payer, and jurisdiction requirements.