The D-KEFS (Delis-Kaplan Executive Function System, 2001) is a battery of nine co-normed stand-alone tests of executive function for ages 8 to 89, with no overall composite by design. Neuropsychologists, clinical psychologists, and school psychologists select tests from it in ADHD, TBI, epilepsy, and dementia evaluations. This page covers how to write up D-KEFS results, with a fictional sample and free templates.
Neuropsychologists, clinical psychologists, and school psychologists; publisher qualification level C
Referring physicians and neurologists, IEP and Section 504 teams, rehabilitation teams, accommodation reviewers, courts and payers
300 to 700 words for the D-KEFS section · administration varies by tests selected, about 90 minutes for the full battery
Norm-referenced executive function battery
Neuropsychological and psychoeducational referrals where executive function is in question: ADHD, TBI, epilepsy, dementia workups, rehabilitation planning
Published by NCS Pearson (2001; all-digital D-KEFS Advanced released 2025); described here for write-up purposes, no test content reproduced
The Delis-Kaplan Executive Function System (D-KEFS) is a battery of nine stand-alone tests of executive function published in 2001 by The Psychological Corporation, now NCS Pearson, and authored by Dean Delis, Edith Kaplan, and Joel Kramer. It was the first executive battery co-normed on a large national standardization sample, it spans ages 8 through 89, and it carries the authors' cognitive-process approach: each test samples a higher-level skill through layered conditions, so the examiner can watch how performance changes as executive demand is added to an intact lower-level skill. The nine tests are Trail Making, Verbal Fluency, Design Fluency, Color-Word Interference, Sorting, Twenty Questions, Word Context, Tower, and Proverb. Primary achievement measures are age-referenced scaled scores with a mean of 10 and a standard deviation of 3; contrast scores put the difference between two related conditions on that same metric.
Since 2025 the name covers two current products, and sorting that out is now the first job of any write-up. The original 2001 battery remains in print, sold and supported, with no announced retirement. The D-KEFS Advanced, published in July 2025 and announced by Dean Delis that September, is administered and scored entirely on iPad through Q-interactive, carries six tests (updated Trail Making, Verbal Fluency, Color-Word Interference, and Tower, plus two new ones, the Social Sorting Test and the Risk-Reward Decision Test), spans ages 8 to 90, and is normed on a national sample of over 1,280 people matched to 2020 US Census figures. Pearson states that "The D-KEFS Advanced complements and does not replace the original D-KEFS", and Delis explained the naming directly: "That's also why we named the new edition the D-KEFS Advanced rather than the D-KEFS-2." A defensible report names the edition it used, because the two products share a name but not tests, format, or norms.
The design choice that shapes every D-KEFS write-up is the absence of an overall composite score. The battery is built for profiling, not for a single executive number, so results are organized by construct rather than summed. It is also strictly a performance measure: it samples what a person can produce in a structured one-to-one session, which research shows is a different construct from the everyday executive behavior captured by rating scales such as the BRIEF-2. In practice, a D-KEFS section lives inside a fuller neuropsychological report or psychological evaluation report rather than standing alone.
Neuropsychologists, clinical psychologists, and school psychologists administer the D-KEFS (Pearson qualification level C) whenever a referral turns on executive function: ADHD and learning evaluations, traumatic brain injury and rehabilitation planning, epilepsy workups, dementia and other adult neurocognitive questions, and psychiatric differentials where planning, inhibition, or flexibility is in doubt. Almost nobody gives all nine tests: Pearson's own product page invites examiners to administer any combination, and published clinical studies routinely use selected subtests, most often Trail Making, Verbal Fluency, Color-Word Interference, and Tower. Co-norming is the reason to select within the D-KEFS rather than assembling standalone instruments with different normative baselines. In a typical battery it runs alongside an intelligence measure, the WAIS for adults or the WISC-V for school-age clients, with a rating scale such as the BRIEF-2 covering the everyday-behavior dimension that performance testing does not reach; for young children who need a broader developmental battery, the NEPSY-II covers overlapping ground for ages 3 to 16.
No statute, payer, or publisher mandates a D-KEFS results-section format. The structure below is the convention experienced neuropsychologists converge on because it survives review: it names the edition and the tests given, organizes findings by executive construct instead of test by test, keeps primary scaled scores as the interpretive backbone, and quarantines process-level detail behind an explicit caution. Each section carries the pitfall that most often undermines it.
Measures and edition statement. Name the battery and edition (original D-KEFS, 2001, or D-KEFS Advanced, 2025), list the tests administered, and record the format: paper, Q-interactive, or telepractice. Selective administration is standard practice, so the list is information, not a confession; add a clause on why these tests fit the referral question. Pitfall: writing "the D-KEFS" with no edition. Two current products share the name with different tests, formats, and norms, so an unnamed edition leaves the reader unable to place a single score.
Behavioral observations during testing. Two or three sentences on engagement, pace, frustration tolerance, and strategy use, written to inform interpretation: whether instructions landed on first presentation, how the examinee handled timed switching demands, whether errors were noticed and self-corrected. Pitfall: observations that never connect to a score. If frustration on switching tasks was worth recording, it is worth referencing when the switching scores are interpreted.
Results by executive construct. Group findings by what they measure across tests: inhibition (the Color-Word Interference inhibition condition), flexibility and switching (Trail Making number-letter switching, inhibition/switching, category switching), fluency and generativity (letter, category, and design fluency), planning and problem solving (Tower, Twenty Questions), and concept formation and abstraction (Sorting, Word Context, Proverb). Report primary achievement scaled scores (mean 10, SD 3) as the backbone, and use the intact lower-level baseline conditions to localize a weakness before calling it executive. Factor-analytic work supports construct-level grouping over test-level lists. Pitfall: the test-by-test laundry list. It buries the profile, invites over-reading of isolated scores, and reads like a score printout instead of an interpretation.
Process and contrast observations. The battery's signature comparisons, such as switching versus baseline or category versus letter fluency, belong here as hypothesis-level detail: describe the pattern, say what it suggests, and state the caution in the section itself. Published reliability estimates for contrast measures are low, so no diagnostic conclusion should rest on one. Pitfall: a contrast score doing decisional work. Reviewers who know the psychometrics read that as a validity error, not a style choice.
Integration with ratings, records, and history. Put the D-KEFS findings next to rating-scale results, school or work records, and interview history, and explain convergence and divergence rather than averaging them. Performance measures and rating scales sample different conditions and correlate weakly, so disagreement is an expected finding to interpret, not a problem to hide. Pitfall: declaring whichever source is inconvenient "invalid" instead of explaining what each one samples.
Interpretive summary. A short paragraph stating the construct profile in plain language, tied to the referral question, with the limits of the data stated where a reviewer looks for them: the norms year, the selective battery, and the single-session sample. The D-KEFS has no overall composite, so the summary is a profile statement, never a summary score. Pitfall: inventing a composite by averaging scaled scores. The battery withholds that number on purpose.
Recommendations linked to findings. Every recommendation traces to a documented strength or weakness: supports for switching demands if switching was weak, strategy instruction where planning broke down, re-evaluation timing that accounts for the battery's limited alternate forms. Pitfall: a boilerplate list that could sit under any profile.
D-KEFS RESULTS SECTION SKELETON (adapt; delete guidance in parentheses before signing) MEASURES AND EDITION Battery and edition: original D-KEFS (2001) / D-KEFS Advanced (2025) Tests administered and why they fit the referral question: ____ Format (paper / Q-interactive / telepractice) and norms year stated: ____ BEHAVIORAL OBSERVATIONS DURING TESTING Engagement, pace, frustration tolerance, strategy use: ____ Observations that qualify specific scores: ____ RESULTS BY CONSTRUCT (primary scaled scores: mean 10, SD 3) Inhibition: ____ Flexibility and switching: ____ Fluency and generativity: ____ Planning and problem solving: ____ Concept formation and abstraction: ____ (Localize first: intact baseline conditions before any executive claim) PROCESS AND CONTRAST OBSERVATIONS (hypothesis level only) Pattern observed, what it suggests, and the reliability caution: ____ INTEGRATION WITH RATINGS, RECORDS, AND HISTORY Convergence and divergence, explained rather than averaged: ____ INTERPRETIVE SUMMARY (construct profile; the D-KEFS has no composite score) Profile statement tied to the referral question: ____ Limits of the data (norms year, selective battery, single session): ____ RECOMMENDATIONS LINKED TO FINDINGS 1. ____ 2. ____ 3. ____ Clinician signature, credentials, date: ____
Free to use and share, no signup. The PDF includes a one-page cheat sheet with section-by-section pitfalls and a pre-sign checklist; the DOCX is the blank results-section skeleton, ready to adapt.
Scenario: a 15-year-old tenth grader referred by his pediatrician for an ADHD evaluation after two years of declining grades and missed assignments. Four D-KEFS tests were administered inside a fuller battery that included the WISC-V and parent and teacher executive function ratings; this is the D-KEFS section of the larger report only, condensed but structurally complete. All details are fictional.
Client: J.T., 15 · Referral: ADHD evaluation, referred by pediatrician · Evaluator: M. Okafor, PsyD, Licensed Psychologist · Testing date: 07/08/2026 · Report date: 07/15/2026
Measures and edition: Selected tests from the Delis-Kaplan Executive Function System (D-KEFS, 2001 print edition), administered on paper in one session: Trail Making, Verbal Fluency, Color-Word Interference, and Tower. Scores are age-referenced scaled scores (mean 10, SD 3) from the 2001 national norms; the age of those norms is noted under limitations below. Tests were selected to examine the inhibition, switching, and planning questions raised by the referral.
Behavioral observations: J.T. engaged readily, understood each instruction on first presentation, and worked quickly and accurately on baseline conditions. On the switching conditions he slowed markedly, showed visible frustration, and caught most of his own errors, saying he could feel himself losing the rule. These observations inform the interpretation of the switching scores below.
| Test and condition | Scaled score |
|---|---|
| Trail Making: Visual Scanning | 10 |
| Trail Making: Number Sequencing | 11 |
| Trail Making: Number-Letter Switching | 6 |
| Verbal Fluency: Letter Fluency | 11 |
| Verbal Fluency: Category Fluency | 12 |
| Verbal Fluency: Category Switching | 9 |
| Color-Word Interference: Color Naming | 10 |
| Color-Word Interference: Word Reading | 11 |
| Color-Word Interference: Inhibition | 6 |
| Color-Word Interference: Inhibition/Switching | 5 |
| Tower: Total Achievement | 9 |
Results by construct: Lower-level component skills are intact: visual scanning, number sequencing, color naming, and word reading all fell in the average range (scaled scores 10 to 11), which localizes the weaknesses that follow to executive demands rather than to speed, sequencing, or reading. Inhibition was a clear weakness: the Color-Word Interference inhibition condition (scaled score 6) fell at approximately the 9th percentile for age, and the combined inhibition/switching condition (5) near the 5th. Cognitive flexibility showed the same pattern on a second, independent test, with Trail Making number-letter switching (6) well below that test's intact baseline conditions. Fluency was average to high average (letter fluency 11, category fluency 12), with category switching (9) at the lower edge of the average range. Planning and rule-governed problem solving on the Tower test (9) was average.
Process observations: Errors concentrated on the switching conditions and were mostly self-corrected, consistent with the observed loss of set under dual demands. Contrast comparisons pointed in the same direction and are noted only as supporting observation, given the published reliability limits of D-KEFS contrast measures; no conclusion in this report rests on a contrast score.
Integration: Parent and teacher ratings on the BRIEF-2, reported earlier in this evaluation, were elevated for everyday inhibition, shifting, and task completion. The performance findings converge on inhibition and flexibility while the ratings describe the broader everyday picture, an expected relationship because structured one-to-one testing and daily classroom demands sample different conditions; the two sources are integrated here rather than averaged. School records and interview history, also reported earlier, describe longstanding inattentive and executive difficulties.
Summary and recommendations: Across two independent tests, J.T. shows a specific weakness in inhibition and set-shifting against intact lower-level skills, average planning, and average to strong fluency. The D-KEFS yields no overall composite by design, so no summary score is reported; this construct profile, together with the rating-scale, record, and interview evidence elsewhere in this evaluation, supports the diagnostic formulation in that section. Limitations: a selective four-test battery in a single session, scored against the 2001 norms. Recommendations tied to these findings: classroom supports through his Section 504 plan that reduce rapid task-switching demands and chunk multi-step instructions, explicit strategy instruction in self-monitoring under time pressure, and re-evaluation planning that accounts for the D-KEFS having alternate forms for only three of its nine tests.
This sample is fictional and for educational purposes. It does not describe a real client or record; the scores are invented for illustration and correspond to no real child or record.
Writing these after every session? BastionGPT drafts complete notes from bullets, dictation, or a transcript.
Generate a note from bulletsNo law, payer policy, or education regulation names the D-KEFS; every requirement attaches to the evaluation around it, so the write-up's job is to satisfy the framework the report will be used in. Under IDEA, which is law, special-education eligibility rests on multiple sources: 34 CFR 300.304(b)(2) bars any single measure from serving as the sole criterion, and executive function evidence conventionally enters under categories such as traumatic brain injury or other health impairment as one contributor among several. Testing-accommodation documentation runs on agency policy layered over the ADA and Section 504: the ETS documentation guidelines (2026 edition), for example, expect a current comprehensive evaluation with each requested accommodation "explicitly linked to the individual's disability-related functional limitations". US payer coverage is policy about services, not instruments: neuropsychological testing bills under evaluation services 96132 and 96133 with test administration and scoring under 96136 and 96137 (96138 and 96139 by technician), Medicare billing articles A57481 and A57780 expect the record to show medical necessity, the tests administered, scoring and interpretation, and time, and psychological (96130, 96131) and neuropsychological evaluation services are not billed in the same episode; the neuropsychological report page covers that billing chain in depth. In Australia, a search of the MBS for neuropsychology returns no general item as of July 2026 (one narrow item exists for complex neurodevelopmental assessment under age 25), so funding typically runs through private, insurer, NDIS, or state-clinic routes; Canadian funding is a provincial and private mix.
The defensibility question specific to the D-KEFS is edition and norms currency. Name the edition in every report. The original battery's norms date to 2001, a quarter century old as of July 2026, and that age is a disclosable limitation rather than a disqualification: the battery remains published, sold, and supported, with no announced retirement. The 2025 D-KEFS Advanced sharpens the question simply by existing, with norms collected against 2020 Census figures, an all-digital format, and a partly different test list, so edition choice is now a real clinical decision: weigh the constructs the referral needs (only the Advanced carries the new social-cognition and risk-decision tests; only the original carries Design Fluency, Sorting, Twenty Questions, Word Context, and Proverb), the format constraints of your setting, and the documentation-currency expectations of high-stakes reviewers. For serial assessment, alternate forms exist for only three tests of the original battery (Sorting, Verbal Fluency, and Twenty Questions), so practice effects belong in the interpretation of everything else. Record the administration format: Pearson has published digital-to-paper equivalence evidence for four subtests (Design Fluency, Verbal Fluency, Trail Making, and Color-Word Interference, in Q-interactive Technical Report 3) and maintains telepractice guidance for the battery. Purchase requires Pearson qualification level C; qualification to buy is not competence for every interpretive use, which remains a licensing and professional-standards matter.
D-KEFS and D-KEFS Advanced are trademarks, in the US and other countries, of Pearson plc; both editions are published by NCS Pearson, Inc. BastionGPT is not affiliated with, or endorsed by, the publisher. This page reproduces no test items, stimuli, norms, or scoring materials.
Two published findings explain most of the failures. Crawford, Sutherland, and Garthwaite (2008, Journal of the International Neuropsychological Society) estimated reliability for all 51 D-KEFS contrast measures: none exceeded 0.7, the mean was 0.27 and the median 0.30, standard errors of measurement often approached the scores' own standard deviations, and the authors concluded that contrast measures "should not be used in neuropsychological decision making". Toplak, West, and Stanovich (2013, Journal of Child Psychology and Psychiatry) reviewed 20 studies comparing performance-based executive measures with rating scales: 68 of 286 correlations (24 percent) reached significance, and the median correlation was .19. A write-up that ignores either finding reads as naive to the reviewers who know them. The BastionGPT Clinical Advisory Board sees the same errors most often in D-KEFS reviews:
BastionGPT is specifically trained, tuned, and clinically tested on psychological and neuropsychological evaluation reports.
See how clinicians use it day to day on the AI therapy notes page.
Many BastionGPT users report saving more than 90 minutes per day on documentation.
HIPAA-compliant with a signed BAA on every plan. Your data is never used to train models. BastionGPT drafts, you review and sign.
Primary D-KEFS measures are age-referenced scaled scores with a mean of 10 and a standard deviation of 3, so each 3 points is one standard deviation: a score of 10 sits at the middle of the distribution for the person's age, a 7 falls about one standard deviation below it (roughly the 16th percentile), and a 4 about two below (near the 2nd percentile). Reports translate the numbers into qualitative descriptor words; whichever descriptor convention you use, apply it consistently and keep the number next to the word. Error scores and several optional process measures use other metrics, including cumulative percentile ranges, so a careful write-up says which score type each number is. There is no overall battery score to interpret: the D-KEFS reports no composite by design.
They are two current products from the same author team and publisher. The original D-KEFS (2001) is the nine-test battery, print-based at its core, normed for ages 8 through 89. The D-KEFS Advanced (July 2025) is administered and scored entirely on iPad through Q-interactive, requires a Q-interactive Standard license or a Digital Assessment Library for Schools account, spans ages 8 to 90, was normed on over 1,280 people matched to 2020 US Census figures, and carries six tests: updated Trail Making, Verbal Fluency, Color-Word Interference, and Tower, plus the new Social Sorting Test and Risk-Reward Decision Test. Pearson frames the Advanced as a complement rather than a replacement, both stay on the market, and the deliberate naming (Advanced, not a second edition) signals exactly that. Name the edition in your report, and where currency matters, say which norms stand behind the scores.
No. No publisher, payer, or standard requires the full battery. Pearson's product page says "Administer any or all of the D-KEFS tests to customize your assessments", and published clinical research routinely uses selected subtests, most often Trail Making, Verbal Fluency, Color-Word Interference, and Tower. The reason to select within the D-KEFS rather than mixing standalone instruments is co-norming: every test is referenced to the same standardization sample. Document which tests you gave and why they fit the referral question, treat the constructs you did not test as untested rather than assumed intact, and remember when planning re-evaluation that alternate forms exist for only three tests: Sorting, Verbal Fluency, and Twenty Questions.
No authority names it. In schools, IDEA's rule is instrument-neutral: evaluators may "not use any single measure or assessment as the sole criterion" (34 CFR 300.304(b)(2)), and executive function evidence enters eligibility categories such as traumatic brain injury or other health impairment by convention, not by mandate. Accommodation reviewers apply agency policy: the ETS documentation guidelines (2026 edition) ask for a current, comprehensive evaluation that ties each requested accommodation to documented functional limitations. Medicare pays for the testing service, not the instrument: the 96132 to 96139 code family under billing articles A57481 and A57780, none of which names the D-KEFS. In Australia, an MBS search for neuropsychology returns no general item as of July 2026. What every framework does require is documented reasoning: why this battery, what it showed, and how the findings connect to the decision; the neuropsychological report page covers those frameworks in depth.
By design. The authors built the battery on a process approach in which executive function is multidimensional and the diagnostic meaning lives in the profile: comparing conditions within a test and constructs across tests. A single executive quotient would flatten exactly the variability the battery exists to reveal, so no composite is provided and none can be legitimately derived by averaging. Independent test reviews describe the same architecture: nine co-normed stand-alone tests interpreted through comparison (Homack, Lee, and Riccio, 2005, Journal of Clinical and Experimental Neuropsychology). In the write-up, the summary paragraph does the work a composite cannot: a construct profile in plain language, tied to the referral question. Pages offering to help you interpret your "D-KEFS composite score" are describing an instrument that does not exist.
Report them as supporting observation, never as the basis of a decision. Crawford, Sutherland, and Garthwaite (2008) estimated reliability for all 51 contrast measures: none exceeded 0.7, the mean was 0.27, the median 0.30, and standard errors of measurement were large enough that the authors concluded contrast measures should stay out of neuropsychological decision making. The battery's authors defended the process approach in a published update (Delis, Kramer, Kaplan, and Holdnack, 2004, Journal of the International Neuropsychological Society), and the practical synthesis is stable: lead with primary achievement scaled scores, use contrast and error patterns to describe how a score was earned and to generate hypotheses, and state the caution inside the section. The empirical warning is concrete: in a pediatric ADHD sample, groups separated on primary measures while no contrast score distinguished them (Wodka et al., 2008).
Expect it and explain it. Across 20 studies, 68 of 286 correlations (24 percent) between performance-based executive measures and rating scales reached significance, with a median correlation of .19 (Toplak, West, and Stanovich, 2013), which is evidence that they measure related but different constructs, not evidence that one is wrong. The D-KEFS samples processing efficiency under optimal, structured, one-to-one conditions; the BRIEF-2 samples goal-directed behavior in everyday environments over months. A strong report states both findings, explains the difference in what was sampled, and uses the pattern diagnostically: average performance with elevated ratings often means the structure of the testing session supplied what the everyday environment does not. Averaging the two, or discarding one, is the error reviewers flag.
Not legitimately on any public page, including this one. Items, stimuli, record forms, norms tables, and the manuals are protected test materials sold under Pearson qualification level C, and psychologists carry an ethical duty to maintain test security under APA Ethics Standard 9.11. Sites circulating scoring-manual content undermine the norms every report depends on, and some of it is flatly wrong, like composite-scoring instructions for a battery that has no composite. What a report can include: scaled scores, your own prose describing what each test measures, and fictional illustrations like the sample on this page. What it cannot: item text, stimuli, scoring rules, or reproduced norm tables. If you need the materials themselves, the publisher's product page and platforms are the legitimate route.
Yes. BastionGPT is trained and clinically tested on psychological and neuropsychological evaluation reports, the parent documents D-KEFS sections live inside. Paste a score summary (edition, tests given, scaled scores, observations) and it drafts the construct-organized results narrative with the edition statement and contrast-score caution in place for your review; it can also check a finished section for an invented composite, score-versus-descriptor mismatches, or a missing norms disclosure, and produce a plain-language summary for the family or referring provider. BastionGPT is HIPAA-compliant with a signed BAA on every plan, your data is never used to train models, and drafting from a score summary you paste means no protocol or item content ever needs to leave your records.
The instrument facts and compliance claims on this page trace to these sources, last verified July 2026:
Educational content, not legal or billing advice. Sample notes are fictional. Follow your organization's policies and your board, payer, and jurisdiction requirements.