Trail Making Test Documentation: Score Interpretation & Sample Note

The Trail Making Test is a timed paper-and-pencil measure of visual search, sequencing, and set shifting: Part A connects numbered circles, Part B alternates numbers and letters. Neuropsychologists, psychologists, and occupational therapists use it in cognitive batteries and driving screens. Raw seconds mean nothing without a named norm system. This page covers how to document and interpret Trail Making results, with a fictional sample note.

Free to use and share. No signup required.
Already have session bullets or a transcript? Generate a structured draft with BastionGPT — you review and sign it.
Who writes it

Neuropsychologists and psychologists select and interpret it; trained psychometrists may administer and score under supervision; occupational therapists use it in driving and functional screens

Audience

Neuropsychologists and psychologists, referring physicians, occupational therapists and driving evaluators, licensing authorities, courts and attorneys, payers reviewing testing claims

Typical length

4 to 10 chart lines or a short report paragraph · administration about 5 to 10 minutes, plus scoring and integration

Format family

Timed visual search and set-shifting test (two paper conditions; completion time in seconds is the raw metric)

When it's used

Neuropsychological batteries for dementia, TBI, ADHD, and executive complaints; older-driver and fitness-to-drive screens; serial monitoring with alternate forms

Standards context

Classic form from the 1944 Army battery with contested rights messaging; commercial relatives from Pearson, PRO-ED, and PAR; no authority mandates it or sets a universal cutoff; no test content reproduced

What is the Trail Making Test?

The Trail Making Test is the most widely administered paper-and-pencil measure of visual search, sequencing speed, and set shifting. Part A asks the examinee to connect 25 numbered circles in order; Part B alternates between numbers and letters (1, A, 2, B), adding the demand of holding and switching between two sequences. The task began as John Partington's late-1930s distributed-attention test, entered the record as a component of the US Army Individual Test Battery published by the War Department in 1944, and was adopted and clinically validated by Ralph Reitan in the 1950s within the Halstead-Reitan tradition, whose laboratory later published the standard administration and scoring manual. The primary raw score for each part is elapsed completion time in seconds; there is no standardized score until a named normative system converts it.

The load-bearing fact is that the classic two-part test is only one member of a family whose scores do not interchange. The D-KEFS Trail Making Test (Pearson, 2001) splits the task into five conditions with its own scaled and contrast scores, and its number-letter switching condition is not classic Part B; the digital D-KEFS Advanced version (2025) is normed separately again. The Comprehensive Trail-Making Test, Second Edition (PRO-ED, 2020) uses five trails and proprietary T scores, and the Color Trails Test and Oral Trail Making Test were built to reduce alphabet and motor demands, at the price of separate norms and no score conversion. A chart entry that says only "Trails B = 130, impaired" is therefore close to meaningless: the reviewer needs the exact form, the raw seconds, the error count, and the norm system, because the same 130 seconds is unremarkable for an 80-year-old with eight years of education and clearly abnormal for a 30-year-old graduate. The Stroop test page covers the interference measure this test is most often paired with in executive batteries.

Who uses Trail Making Test documentation and when

Neuropsychologists and clinical psychologists administer it inside nearly every adult battery: dementia and mild cognitive impairment workups, traumatic brain injury, ADHD and executive complaints, epilepsy, and medicolegal evaluations, with results integrated into a neuropsychological report. Occupational therapists and physicians use Part B in older-driver screens: the AGS and NHTSA clinician's guide keeps it in its office battery (renamed CADReS in the 2019 edition), and the Canadian Medical Association's driver's guide discusses a Part B screening trigger, in both cases as one input to a functional decision, never a pass or fail. Pediatric evaluators use age-appropriate systems instead (the D-KEFS from age 8, the CTMT-2, or the Children's Color Trails Test; the NEPSY-II page covers the pediatric battery context). The documentation reader is usually downstream: the next evaluator re-deriving a standardized score, a payer reviewing a 96130-series claim, a licensing authority weighing a driving referral, or an attorney testing whether the norms fit the examinee, which is why the raw data, error counts, and norm identification carry the note. The RBANS page covers the repeatable screening battery it often runs alongside.

How to document Trail Making Test results in the chart

No US, Canadian, or Australian authority prescribes a Trail Making documentation format, mandates the test, or sets a universal cutoff. What survives review is a record from which an experienced reader could re-derive the standardized result: the exact form, raw seconds, errors under the standard correction convention, completion status, the named norm system with its demographic inputs, and the confounds that bound interpretation. Each element below carries the pitfall that most often undermines it.

Exact test identity and form. Name the instrument precisely: classic paper TMT (with vendor or source of the form), D-KEFS Trail Making Test (2001 battery or the digital Advanced version), CTMT-2, Color Trails, Oral TMT, or a named digital task, plus the language, alternate form if any, modality (paper, tablet, oral), and hand used. Digital and paper performances are not automatically equivalent, especially for Part B. Pitfall: "Trails" without a version, or a D-KEFS switching score reported as "TMT-B." The relatives use different conditions, scores, and norms, and none converts to a classic result.

Raw times in seconds, both parts. Report Part A and Part B as elapsed seconds, each with completion status. Seconds are the raw datum a later reviewer needs to re-derive any standardized score, so they stay in the record even after conversion, together with the standardized metric and its direction. Pitfall: A standardized score with no raw time, or a raw time with no norm system. Either half alone cannot be checked, compared, or defended.

Errors and the correction convention. Under the standard administration the examiner points out each error immediately, the examinee corrects it, and the clock keeps running, so the cost is absorbed into the time with no separate deduction. State that convention and still record the number and type of errors per part: sequencing slips versus losses of the alternating set carry different information. Pitfall: Errors left uncounted because "only time scores." Part B errors carry independent clinical information; in one lesion study the errors, not the completion time, tracked the frontal finding.

Discontinuation and the ceiling, sourced. If a part is stopped, record the elapsed time at discontinuation, the prespecified ceiling and where it comes from (protocols differ: 150 or 100 seconds for Part A and 300 for Part B in common research protocols, a five-minute clinic convention, or a norm system's own rule), the errors to that point, and the reason: time, set loss, fatigue, distress, or a motor or visual barrier. Pitfall: Silently charting 300 or 301 seconds as an observed completion time. Score at ceiling only where the named norm system directs it, and say the ceiling creates a floor effect that limits severity grading.

Norm system and demographic inputs. Name the normative source and its corrections, and record the inputs it used: exact age at testing and completed education at minimum. Options differ materially: Tombaugh's 2004 age-and-education strata (ages 18 to 89), the Heaton Revised Comprehensive Norms (PAR, demographic T scores), Mayo's older-adult systems and their 2024 update, population-specific sets, and national norms abroad. Pitfall: "Age norms" or "Heaton norms" without the edition and inputs, or strata from one system combined with corrections from another. The same raw seconds can shift by roughly half a standard deviation between accepted systems.

Derived indexes, named and supported. If you report B minus A or B divided by A, name the formula and the normative or regression source that supports interpreting it. The difference score still tracks baseline speed and intelligence, the ratio destabilizes when Part A is unusually fast or slow, and large clinical studies found neither adds classification value over the two direct times. Pitfall: "Trails ratio 3.2, impaired." No universal B/A boundary exists, and a derived score reported without its components and source is uninterpretable.

Validity qualifiers and interpretation. Document tremor, weakness, arthritis, or other motor limits with the hand used; visual acuity, field loss, or neglect and any asymmetric search behavior; literacy and alphabet familiarity for Part B (years of schooling is an imperfect proxy, and non-Latin-alphabet backgrounds change the task); medications, sleep, and engagement. Then interpret inside the battery: Part B adds set alternation to a task already loaded on search and speed, so write pattern-level conclusions, and frame any driving concern as a referral trigger, not a determination. Pitfall: All slowing read as executive dysfunction, or a driving conclusion from one office score. Guidelines in all three countries treat the test as one input that can trigger functional or on-road assessment.

Blank template (copy and adapt)

TRAIL MAKING TEST DOCUMENTATION BLOCK
Date: [ ]   Setting: [ ]   Administered by: [ ]   Interpreted by: [ ]
Referral question / battery context: [ ]
Test identity: [classic paper TMT (form source/vendor) / D-KEFS TMT (2001 or
   Advanced 2025) / CTMT-2 / Color Trails / Oral TMT / named digital task]
Form + modality: [standard or alternate form, language, paper or digital,
   hand used, deviations from standard administration]
Part A: [ ] seconds   Errors: [ ]   Completed: [Y/N]
Part B: [ ] seconds   Errors: [ ] ([sequencing / set-loss])   Completed: [Y/N]
Error convention: [errors corrected in real time; timing continued;
   no deduction]
Discontinuation (if any): [elapsed time, prespecified ceiling + its source,
   errors to that point, reason; no invented completion time]
Norm system: [Tombaugh 2004 / Heaton RCN 2004 (PAR) / MOANS / MOAANS /
   updated Mayo 2024 / other]  Inputs used: [age, education, others]
Standardized scores: [metric + direction; Part A, Part B]
Derived index (optional): [B-A / B/A, formula named, supporting source,
   result]
Observations: [search pattern, pencil lifts, set losses, self-corrections,
   pace, fatigue, frustration]
Validity: [motor, visual, literacy and alphabet familiarity, language,
   medications, engagement]  Conclusion: [interpretable / with caution /
   not interpretable: reason]
Interpretation: [pattern within the battery; what it does not establish;
   driving framed as referral trigger only, if relevant]
Prior testing: [date, form, modality, norms; practice effects and
   reliable-change framing for serial results]
Plan / integration: [ ]
Clinician signature / credentials:            Date:

Free to use and share, no signup. The PDF includes a one-page cheat sheet with element-by-element pitfalls and a pre-sign checklist; the DOCX is the blank documentation block, ready to adapt. Neither reproduces the test form, circle layouts, norms, or scoring tables.

Sample Trail Making Test documentation (fictional)

Scenario: a dementia workup in which Part B is discontinued at the clinic's prespecified ceiling, documented without inventing a completion time, with the driving question framed as a referral. All details are fictional.

Patient: H.L., 79  ·  Setting: Outpatient neuropsychology clinic, cognitive decline evaluation  ·  Clinician: M. Beaudry, PhD  ·  Note date: 08/18/2026

Measure and administration: Trail Making Test, classic paper form, standard administration in English with the dominant right hand, administered by this examiner within a dementia evaluation battery. Corrected vision worn; no field cut on confrontation screening earlier in the battery. Sample items completed correctly on both parts after one repetition of the Part B instructions. Errors were pointed out and corrected in real time with timing continuing throughout, per the standard convention; no deduction applied.

Results: Part A: 74 seconds, completed, one examiner-corrected sequencing error. Part B: discontinued at the clinic's prespecified 300-second ceiling, not completed, four examiner-corrected errors (one sequencing, three losses of the alternating set, each a return to the number sequence). Part A was scored against Tombaugh (2004) age-and-education norms (age 79, 10 years of education, documented in the scoring record), falling in the low-average range for her stratum. Because Part B was not completed, no completion-time standardized score was assigned; the discontinuation, elapsed ceiling, and error pattern are reported in place of a number, and the ceiling's floor effect is noted as limiting severity grading. No derived index was computed on an incomplete part.

Observations and validity: Visual search on Part A was slow but systematic, with no lateralized omissions. On Part B she articulated the rule correctly, managed the first three transitions, then repeatedly reverted to consecutive numbers despite prompts, appearing effortful and increasingly frustrated; no tremor or motor slowing interfered with drawing. Attention was adequate and effort was engaged across the battery, with validity indicators unremarkable: the pattern is read as cognitive, not motor, visual, or effort-related.

Interpretation: Marked difficulty maintaining and alternating between two sequences, beyond the generalized slowing evident on Part A, convergent with the memory and category fluency findings elsewhere in the battery and with the informant's report of declining instrumental function. The result is one component of the evaluation, not a stand-alone diagnostic finding, and no severity grade is assigned from a ceiling score.

Driving and plan: Because she drives locally, the performance exceeds the screening triggers discussed in the older-driver guidance (Part B beyond three minutes with three or more errors) and supports referral, not a determination: an occupational therapy driving evaluation with on-road assessment was recommended and discussed with the patient and her son, consistent with guidance that no single office test decides fitness to drive. Full integration, diagnosis, and recommendations appear in the evaluation summary; any retest will use an alternate form with the interval and practice effects addressed by a reliable-change method.

This sample is fictional and for educational purposes. It does not describe a real patient or record; the times, errors, dates, and details are invented to show documentation structure and are not clinical guidance. No test form, layout, or norm-table values are reproduced.

↑ Back to the template and downloads

Why this sample works

  • The raw seconds, error counts, completion status, norm system, and demographic inputs are all present, so a later reviewer can re-derive the Part A score and see exactly why Part B has no number.
  • The discontinuation is handled the defensible way: elapsed time at a named, prespecified ceiling, the error pattern in place of an invented completion time, and the floor effect stated.
  • Errors are counted and typed (sequencing versus set loss) even though the correction convention absorbs them into the time, which is what makes the set-shifting interpretation checkable.
  • Motor, visual, and effort explanations are explicitly examined and set aside before the slowing is read as cognitive, and the interpretation stays at the pattern level inside the battery.
  • The driving concern becomes a referral with its guidance frame named, not a pass-fail call, and the serial-testing rule (alternate form, reliable change) is set before any retest.

Writing these after every session? BastionGPT drafts complete notes from bullets, dictation, or a transcript.

Generate a note from bullets

Documentation and compliance considerations

United States: the external requirements are documentation requirements, not test requirements. Psychological and neuropsychological testing is billed under the 96130 to 96139 family (96136 and 96137 for administration and scoring), and the governing CMS billing articles (A57780 and A57481, current revisions effective October 2025) expect the record to show the referral question, medical necessity, the tests administered with time, scoring, interpretation, and recommendations, with test selection accounting for age, education, ethnicity, and sensory or physical limitations (PAYER POLICY); the Trail Making Test has no standalone code and no CMS policy names it. Driving is state law: licensing and medical review boards decide fitness, clinician reporting duties vary by state (LAW), and the AGS and NHTSA clinician's guide keeps Part B in its CADReS office battery as screening guidance only (CONVENTION); the American Academy of Neurology's dementia-driving guideline found the evidence insufficient for any neuropsychological test to determine driving risk by itself, and FMCSA sets no Trail Making requirement for commercial drivers. In aviation work, the FAA's current guide requires complete score reporting with the normative comparison group identified, using pilot norms when available, but its core-test list has been moved behind a secure site, so a report should not assert that the FAA requires the TMT (CONVENTION). Concussion frameworks (the Amsterdam consensus, CDC guidance, the military's progressive-return protocols) individualize return decisions and mandate no Trail Making test (CONVENTION).

Canada and Australia put the same test inside different frames. Canadian driver fitness runs on provincial law with physician reporting mandatory in most provinces (LAW); the Canadian Medical Association's driver's guide states that no single office test determines fitness and discusses a Part B screening trigger of three minutes or three or more errors, drawn from a systematic review of 47 studies that itself found the underlying cutoff literature methodologically weak, with candidate cutoffs scattered from 90 to 180 seconds (CONVENTION); outpatient neuropsychological testing is generally privately funded, through auto insurers, workers' compensation, or extended benefits rather than provincial plans (PAYER POLICY). In Australia, Austroads' Assessing Fitness to Drive (sixth edition, 2022, applied by licensing authorities and under a national review announced for 2026 to 2028) does not name the Trail Making Test at all: it directs medical, functional, occupational therapy, and on-road evidence (REGULATORY STANDARD), and Australian discrepancy data show that the expected difference between Parts B and A depends on the level of Part A performance, one more reason a raw difference is not a fixed rule. Across all three countries the widely circulated absolute thresholds (a 273-second Part B "impairment" line, a universal B/A ratio of 3, a single NHTSA seconds figure) trace to secondary handouts, single protocols, or conflicting program documents rather than to any demographically corrected norm system, and should be described that way when a record must address them (CONVENTION).

Rights and versions demand unusual care here, because even federal sources disagree. The classic form descends from the 1944 War Department publication, which has a strong basis for US public-domain treatment as a government work, yet one NIH institute's copyright statement describes the test as copyrighted and directs licensing through the current Halstead-Reitan vendor while another NIH page calls it public domain (and misdates the Army battery), and the vendor sells current forms, manuals, and Canadian and Australian English translations under stated use restrictions. The defensible position for a clinic is to separate three questions: the historical task structure can be described freely; a specific modern form, manual text, or translation should be treated as the vendor's product unless counsel has verified the exact edition's provenance; and the norm systems that give raw seconds meaning are unambiguously proprietary (the Heaton Revised Comprehensive Norms are a commercial PAR product, Tombaugh's tables sit in a copyrighted journal article, and D-KEFS, CTMT-2, and Color Trails norms belong to their publishers), so scores and interpretations belong in the record while reproduced forms, tables, and scoring software output stay out of public pages and portals, and a free web task may not present itself as producing classic TMT or any publisher's scores without validation and permission. One scheduled change is worth noting in serial-testing plans: Pearson has announced that support for the original D-KEFS desktop scoring software ends from 2027, while the 2001 battery and the separate D-KEFS Advanced continue. The D-KEFS is a product and trademark of NCS Pearson, Inc.; the CTMT-2 of PRO-ED, Inc.; the Color Trails Test of PAR, Inc.; Reitan-branded materials are sold through the Reitan Neuropsychology Laboratory and the Neuropsychology Center. BastionGPT is not affiliated with, or endorsed by, any of these publishers or vendors. This page reproduces no test items, stimuli, norms, or scoring materials.

↑ Back to the template and downloads

Common Trail Making Test documentation errors reviewers flag

The numbers behind these errors are specific. A 2021 NACC analysis found accepted norm systems moved the same raw times by up to about half a standard deviation; the driving-cutoff systematic review of 47 studies found none justified its sample size and candidate cutoffs scattered from 90 to 180 seconds; in a 2015 lesion study Part B errors, not completion time, tracked the right frontal finding; and in 571 people with TBI derived indexes added nothing over the two direct times. The BastionGPT Clinical Advisory Board sees the same errors most often in Trail Making Test documentation reviews:

  • A time with no norms, or norms with no inputs. "Trails B = 130 seconds, impaired." The same time is unremarkable at 80 with eight years of education and abnormal at 30 with a degree. Chart raw seconds plus the named norm system, its edition, and the exact age and education inputs, so the standardized score can be re-derived and checked.
  • Internet cutoffs charted as authority. The circulating 78-second and 273-second "deficient" lines, a universal ratio of 3, and single NHTSA seconds figures come from secondary handouts and conflicting program documents, not from any demographically corrected norm system. If a referral source cites one, address it as an unsourced heuristic and report the normed result instead.
  • An invented completion time at the ceiling. A discontinued Part B charted as "300 seconds" as if observed, or a percentile assigned to an unfinished part. Record discontinued status, the elapsed time, the ceiling and its source (protocols disagree: 150 or 100 seconds for Part A, 300 for Part B, five-minute conventions), the errors, and the floor effect; apply an above-ceiling code only where the named norm system directs it.
  • Derived indexes doing decisional work. B minus A called a pure executive score, or a B/A ratio deciding impairment. The difference still tracks baseline speed and ability, the ratio destabilizes at fast or slow Part A times, large TBI studies found no added classification value, and the expected discrepancy itself varies with Part A. Report the two direct times first and any index with its formula and source.
  • The wrong instrument's score in the classic slot. A D-KEFS switching scaled score reported as "TMT-B," a CTMT-2 composite abbreviated to "Trails," or an Oral TMT, Color Trails, or unvalidated digital task scored with paper norms. Each relative has its own conditions, norms, and scores, and none converts to a classic result; name what was actually given.
  • Confounded slowing read as executive impairment. Tremor, hemiparesis, field loss or neglect, low literacy, or unfamiliarity with the Latin alphabet scored as set-shifting failure, or a driving determination made from one office score. Chart the confound, qualify or withhold interpretation, consider the alternatives built for these cases, and frame driving as a referral for functional assessment.
How BastionGPT helps

BastionGPT is specifically trained, tuned, and clinically tested on psychological and neuropsychological evaluation reports.

  • Give it the facts (form, modality, raw seconds, errors and their types, completion status, ceiling, norm system and inputs, observations, confounds) and it drafts the documentation block or report paragraph: units stated, discontinuation handled without invented times, norms named, interpretation kept to the pattern, ready for your review.
  • Cross-check a finished note or results section for the gaps reviewers flag: a time without norms, an absolute cutoff doing decisional work, a percentile on an unfinished part, a derived index without its components, or a D-KEFS score sitting in a classic TMT slot.
  • Draft the serial-testing or driving-referral paragraph: prior form and interval, practice-effect and reliable-change framing, or the screening-trigger language that supports an on-road referral without overstepping into a fitness determination.

See how clinicians use it day to day on the AI therapy notes page.

Many BastionGPT users report saving more than 90 minutes per day on documentation.

HIPAA-compliant with a signed BAA on every plan. Your data is never used to train models. BastionGPT drafts, you review and sign.

Frequently asked questions

There is no normal range without a norm system. The raw score is seconds to complete each part, and its meaning depends on age and education at minimum: published systems include Tombaugh's 2004 age-and-education strata from 911 adults aged 18 to 89, the Heaton Revised Comprehensive Norms with demographic T scores (a commercial PAR product), Mayo's older-adult systems and their 2024 update, and population-specific sets. A 2021 analysis found that accepted systems standardize the same raw times differently, with average differences reaching about half a standard deviation, so the note names the system and the inputs rather than citing a free-floating range. The widely copied absolute lines (about 29 seconds average and 78 seconds "deficient" for Part A, 75 and 273 for Part B) are unsourced heuristics that over-pathologize older and less-educated examinees.

Not as a deduction, and yes as information. Under the standard administration the examiner points out each error the moment it happens, the examinee corrects it, and the clock keeps running, so the error's cost is already inside the completion time; no points are subtracted afterward. Record the count and type anyway, per part: sequencing slips differ from losses of the alternating set, and the research argues errors carry independent signal, from geriatric samples where time and errors each discriminated impairment groups to a lesion study where Part B errors, not time, tracked the right frontal finding. Two patients with identical times, one clean and one with repeated set losses, are different clinical pictures, and only the error record shows it.

Chart it as discontinued, then give the data that replaces the number: elapsed time at discontinuation, the prespecified ceiling and its source (protocols genuinely differ: a common clinic convention stops Part B at five minutes, major research protocols use 300 seconds with 150 or 100 for Part A, and some norm systems have their own above-ceiling coding), the error count and pattern to that point, whether the rule was understood, and the reason: time, set loss, fatigue, distress, or a motor or visual barrier. Do not chart the ceiling as an observed completion time or assign a completion-time percentile to an unfinished part unless the named norm system explicitly directs that coding, and say that scoring at ceiling creates a floor effect that limits severity grading among the most impaired examinees.

Use a system whose sample, age range, and corrections actually fit, and say why. For an 85-year-old with eight years of education: Tombaugh (2004) stratifies by age and education through 89; the Heaton system applies simultaneous demographic corrections but tops out around 85, putting this patient at its boundary; Mayo's older-adult systems reach into the 90s with an ability-adjusted model, with a 2024 update modeling age, sex, and education; and population-specific sets exist where they match the examinee. Never combine one system's strata with another's corrections, and remember that years of schooling is an imperfect proxy: reading ability explains performance beyond education in some groups, and Part B assumes automatic Latin-alphabet knowledge, which is why alphabet-reduced alternatives like the Color Trails Test exist for the cases where that assumption fails.

No, in all three countries, and the note should say what role it did play. In the US, licensing and medical review are state law, the AGS and NHTSA clinician's guide keeps Part B in its CADReS office screen as guidance, and the neurology dementia-driving guideline found no neuropsychological test sufficient to determine risk alone. In Canada, the CMA driver's guide states no single office test determines fitness and discusses a screening trigger of three minutes or three or more errors on Part B, but the systematic review behind such cutoffs found the literature methodologically weak, with candidates scattered from 90 to 180 seconds. Austroads in Australia does not name the test at all, directing medical, functional, and on-road evidence. Defensible language: performance exceeded the screening trigger and supports referral for occupational therapy and on-road assessment; the licensing authority makes the determination.

The honest answer is that the rights messaging conflicts, so separate the three questions. The task descends from a 1944 US War Department publication, which gives the exact federal edition a strong basis for US public-domain treatment, and yet one NIH institute's copyright statement calls the test copyrighted and directs licensing through the current Halstead-Reitan vendor while another NIH page labels it public domain; the vendor meanwhile sells current forms, manuals, and restricted Canadian and Australian translations. Practically: the task structure can be described freely; treat any specific modern form, manual, or translation as the vendor's product unless the exact edition's provenance has been verified; and the norm tables that make seconds meaningful are unambiguously proprietary (Heaton through PAR, Tombaugh in a copyrighted journal, D-KEFS and CTMT-2 and Color Trails through their publishers). Scores, times, errors, and interpretations belong in the EHR; reproduced stimulus sheets and norm tables do not belong in portals or public pages, and a web task may not present itself as producing classic or publisher scores without validation and permission.

No. The D-KEFS version restructures the task into five conditions (visual scanning, number sequencing, letter sequencing, number-letter switching, motor speed) on its own page layout, converts times to age-scaled scores, records error types formally, and evaluates switching through contrast scores; its number-letter switching condition is analogous to Part B but normed and scored differently, and the digital D-KEFS Advanced (2025) is normed separately again, so no D-KEFS result converts to a classic TMT score in either direction. The same boundary holds for the CTMT-2's five trails and composite T scores, the Color Trails Test, and the Oral TMT. In serial testing, say plainly that the instrument changed and describe qualitative convergence instead of calculating change. The D-KEFS report write-up guide covers the full battery's structure and scores.

Yes. Give it the facts (form and modality, raw seconds, errors and types, completion status, any ceiling and its source, norm system and demographic inputs, observations, confounds, and prior testing) and it drafts the documentation block or report paragraph: units stated, the correction convention named, discontinuation handled without invented times, norms identified, and interpretation kept to the pattern, ready for your review. It can also cross-check a finished note for a time without norms, an absolute cutoff doing decisional work, a percentile on an unfinished part, or a D-KEFS score in a classic slot. BastionGPT is HIPAA-compliant with a signed BAA on every plan, and your data is never used to train models.

Primary sources

The instrument facts and compliance claims on this page trace to these sources, last verified August 2026:

  1. Lineage and administration: Smithsonian Institution, Army Individual Test Battery catalog record (War Department, 1944; National Research Council committee development); Reitan RM, 1955 and 1958 validity studies; Bowie CR, Harvey PD, 2006, Nature Protocols, administration and interpretation of the Trail Making Test (standard error-correction convention).
  2. Rights record: NIH/NINDS, TMT copyright statement (describes the test as copyrighted; licensing via the vendor); NIDA Data Share, TMT entry (labels it public domain); US Copyright Office, 1909 Copyright Act; Neuropsychology Center, current Halstead-Reitan materials including restricted translated forms.
  3. Norm systems: Tombaugh TN, 2004, Archives of Clinical Neuropsychology, age-and-education norms, 911 adults; PAR, Inc., Heaton Revised Comprehensive Norms; Steinberg BA and colleagues, 2005, Mayo's Older Americans Normative Studies; Lucas JA and colleagues, 2005, MOAANS; Karstens AJ and colleagues, 2024, updated Mayo Normative Studies; Patel BM and colleagues, 2021, norm-system comparison in the NACC cohort (about half a standard deviation between systems).
  4. Errors, derived indexes, and construct: Ashendorf L and colleagues, 2008, TMT errors in normal aging, MCI, and dementia; Kopp B and colleagues, 2015, errors, not time, and right frontal lesions; Sanchez-Cubillo I and colleagues, 2009, construct validity; Martin TA and colleagues, 2003, ratio-score utility in TBI; Lange RT and colleagues, 2005, derived indexes in 571 TBI cases; Senior G and colleagues, 2018, Australian B-minus-A discrepancy data.
  5. Ceilings and protocols: NHTSA and AGS older-driver materials, TMT protocol with 150- and 300-second ceilings; CENTER-TBI, 100- and 300-second protocol; Ahern L and colleagues, 2019, efficiency scoring for noncompleters (research approach).
  6. Serial testing and validity: Dikmen SS and colleagues, 1999, test-retest reliability and practice effects; Calamia M and colleagues, 2012, practice-effect meta-analysis; LoSasso GL and colleagues, 1998, alternate-form difficulty differences; AACN, 2021, performance-validity consensus; Iverson GL and colleagues, 2002, extreme TMT scores as low-sensitivity red flags.
  7. Confounds and adaptations: Schneider BC, Lichtenberg PA, 2011, reading ability beyond years of education; Waggestad TH and colleagues, 2023, Frontiers in Psychology, alphabet knowledge and misclassification; Abeare CA and colleagues, 2019, demographic bias in raw cutoffs; PAR, Inc., Color Trails Test; digital-paper equivalence, computerized trail-making comparison.
  8. Driving and coverage: Roy M, Molnar F, 2013, Canadian Geriatrics Journal, systematic review of TMT driving cutoffs (47 studies); Canadian Medical Association, CMA Driver's Guide, 10th edition; American Academy of Neurology, 2010, dementia and driving guideline; Austroads, Assessing Fitness to Drive; FAA, neurocognitive evaluation protocol; CMS billing articles A57780 and A57481; Pearson product pages for the D-KEFS and PRO-ED for the CTMT-2.

Educational content, not legal or billing advice. Sample notes are fictional. Follow your organization's policies and your board, payer, and jurisdiction requirements.