The Boston Naming Test (BNT-2) is a 60-item confrontation naming test: the examinee names line drawings that run from common to rare words. Neuropsychologists and speech-language pathologists use it in aphasia, dementia, and epilepsy surgery evaluations. A total is uninterpretable without the form, edition, cue accounting, and norm source. This page covers how to document BNT results, with a fictional sample.
Neuropsychologists, psychologists, and speech-language pathologists; US publisher qualification level C for the kit (listed levels differ by seller and territory)
Neuropsychologists and SLPs, neurologists and epilepsy surgery teams, dementia and memory clinics, payers reviewing testing claims, attorneys in medicolegal work
3 to 8 chart lines or a short report paragraph · administration about 10 to 20 minutes for the full form, plus scoring and integration
Confrontation naming test (60 pictured items in a difficulty gradient; short forms are separate score systems)
Aphasia characterization, dementia and MCI evaluations, epilepsy presurgical and postoperative workups, broader neuropsychological batteries, serial monitoring of language change
BNT-2 published by PRO-ED (US) with regional distributors; no authority mandates it and no universal cutoff exists; described for documentation purposes, no test content reproduced
The Boston Naming Test is the most widely used confrontation naming measure in North American practice: 60 black-and-white line drawings, ordered roughly from high-frequency to rare words, that the examinee names one at a time. It began as a 1978 Boston VA research edition by Edith Kaplan, Harold Goodglass, and Sandra Weintraub, was first published commercially by Lea & Febiger in 1983, and exists today as the Boston Naming Test, Second Edition (BNT-2, 2001), published in the United States by PRO-ED, which also handles permissions; PAR sells the BDAE-3 package that contains it, Pearson Clinical distributes it in Australia and New Zealand, and the Lippincott name is the historical imprint. Two currency facts matter for records: no third edition exists ("BDAE-3" modifies the aphasia battery, not the naming test), and PRO-ED now offers a no-cost stimulus upgrade replacing Item 48, a picture the publisher itself describes as racially and culturally inappropriate, while keeping the 60-point maximum, so a report should say which stimulus set was used.
The load-bearing facts are the cue asymmetry and the form question. The countable total is spontaneous correct responses plus responses correct after a stimulus (semantic) cue, the clarification used when the drawing was visually misperceived; responses retrieved only after a phonemic cue, which supplies the word's sound, are recorded separately and never added to the total. And "the BNT" is really a family of forms: the full 60, the publisher's own 15-item short form, the separately permissioned CERAD 15-item set, odd and even 30-item halves, and several research-built short forms, whose scores are not interchangeable; the CERAD 15 in particular cannot be validly prorated to a 60-item score. A record that says "BNT = 42, impaired" is therefore uninterpretable. Naming results usually sit inside a neuropsychological report; the RBANS page covers the screening battery whose brief naming component is not a substitute for this test.
Neuropsychologists administer the BNT inside dementia and mild cognitive impairment workups (where the error and cue pattern helps separate retrieval difficulty from semantic loss), in aphasia and stroke characterization alongside a full language battery, and in epilepsy surgery programs, where naming is a core presurgical and postoperative measure: left temporal resections carry a documented naming-decline risk, so the baseline form, norms, and a reliable-change plan get set before surgery. Speech-language pathologists document confrontation naming within evaluations billed under the 92521 to 92524 family and use the cue profile to shape therapy targets. Geriatric and memory clinics meet it inside CERAD-derived protocols, which use the separately permissioned CERAD 15-item form; US Alzheimer's research centers replaced the BNT with the Multilingual Naming Test in their uniform dataset back in 2015, a fact worth knowing when research records cross into clinical charts. Readers are rarely just the author: the next evaluator needs the form and norms to test change, payers reviewing 96130-series claims need the medical-necessity chain, and the NEPSY-II page covers the pediatric language instruments used instead of adult naming norms. Executive-function measures like those on the D-KEFS page assess fluency, which is not confrontation naming.
No US, Canadian, or Australian authority prescribes a BNT note format, mandates the test, or sets a universal cutoff. What survives review is a record from which a reader can reconstruct exactly what was administered and how the total was assembled: the form and edition, the administration route and stop rule, the cue accounting, the named norms, and the language context. Each element below carries the pitfall that most often undermines it.
Form, edition, and stimulus set. Name the exact form: 60-item BNT-2, the publisher's 15-item short form, the CERAD 15-item set, a named 30-item half, or a named research short form, plus the edition and whether the original or upgraded stimulus set (the Item 48 replacement) was used. The maximum stays 60 across stimulus sets, so only the note preserves the difference. Pitfall: "BNT-15" or "BNT-30" with no source. The publisher 15 and the CERAD 15 are different forms, the halves differ, and none of their scores interchange with the full form or each other.
Administration route and stop rule. State whether the classic basal-and-discontinuation route was used (starting mid-list, establishing a basal, stopping after a run of consecutive failures) or the all-items revised procedure that administers all 60 with defined response intervals. Name the discontinuation rule as applied: the failure-run length differs by edition (six in 1983, eight in the 2001 edition), and say whether phonemic-cued successes interrupted the failure run. Pitfall: Silence on the stop rule. The two published interpretations of the failure run changed final scores in 31 percent of an Alzheimer's sample, by as much as 16 points in older adults, so an unstated convention makes serial comparison unsafe.
Cue accounting, made explicit. Report the arithmetic: spontaneous correct plus stimulus-cued correct equals the countable total, with phonemic-cued successes counted separately and excluded. The stimulus cue exists to repair a visual misperception and restore a valid naming opportunity; the phonemic cue hands over the word's sound, so counting it would inflate the construct. Pitfall: "Answers after cues count." Collapsing the two cue types, or crediting phonemic-cued responses, changes the score and destroys comparability with every properly scored administration.
Error pattern, in words. Describe the errors: semantic paraphasias, phonemic paraphasias, circumlocutions, visual misperceptions, omissions, perseverations. The pattern carries the clinical meaning: strong phonemic-cue benefit suggests retrieval difficulty with the lexicon intact, while semantic errors and no cue benefit point toward semantic degradation. Pitfall: A total with no error description. Two identical totals can represent different syndromes, and the treating SLP or the next evaluator cannot use a bare number.
Norm source and corrections. Name the normative system and the corrections actually applied: the publisher manual, Tombaugh and Hubley (1997), Mayo's older-adult systems (MOANS, MOAANS) and the 2024 Mayo regression norms (4,428 adults aged 30 to 91, with age, sex, and education terms), Zec (2007), Heaton's demographically corrected system, or CERAD norms for the CERAD form only. Record the exact age and education inputs. Pitfall: "Below cutoff" with no system. No universal BNT cutoff exists, systems correct for different variables, and race-stratified historical norms are sample calibrations, not biological claims; mismatched norms are a standing cross-examination target.
Language, culture, and familiarity. Document the testing language, the examinee's language history (dominance, age of acquisition, education language, daily use), and cultural and generational familiarity considerations. Record other-language correct responses separately rather than adding them to the English total, and consider a validated adaptation or the Multilingual Naming Test for bilingual examinees; monolingual English norms overestimate impairment in bilinguals. Pitfall: Anomia diagnosed from an English administration of a culturally loaded stimulus set. Twelve items showed differential functioning between demographically matched groups, and matched young adults hit "impaired" cutoffs at 18 versus 4 percent, so the confound analysis belongs in the note.
Interpretation and the change plan. Interpret the naming result inside the evaluation: what it converges with, what it does not establish (no naming score is independently diagnostic of dementia, aphasia, or seizure laterality), and the follow-up or comparison plan. For serial and surgical work, keep the same form, stimulus set, procedure, and norms, and use a reliable-change frame; post-left-temporal-resection decline is common and expected to exceed chance variation. Pitfall: Change computed across forms or conventions: a CERAD 15 subtracted from a prior 60-item total, a short form doubled into a full-form estimate, or a decline claimed without the stop rule and norms held constant.
BOSTON NAMING TEST DOCUMENTATION BLOCK Date: [ ] Setting: [ ] Administered by: [ ] Interpreted by: [ ] Referral question / battery context: [ ] Form + edition: [60-item BNT-2 / publisher 15-item / CERAD 15-item / named 30-item half / named research short form] Stimulus set: [original / upgraded (Item 48 replacement)] Administration route: [classic basal + discontinuation (rule length as applied; phonemic-cued successes interrupt run: Y/N) / all-items revised procedure] Language of testing: [ ] Cue accounting: Spontaneous correct: [ ] Stimulus-cued correct (counted): [ ] Countable total: [ ]/[60 or form maximum] Phonemic-cued correct (recorded, NOT counted): [ ] Error pattern: [semantic paraphasias / phonemic paraphasias / circumlocutions / visual misperceptions / omissions / perseverations] Norm source + corrections: [system + year; age, education, sex, other inputs actually used] Standardized result: [ ] Language + culture: [dominance, acquisition, education language, daily use; other-language responses recorded separately; adaptation or MINT considered: Y/N] Validity: [vision, hearing for cues, motor speech, familiarity factors] Conclusion: [interpretable / with caution / not interpretable] Interpretation: [pattern within the evaluation; retrieval vs semantic shape; what the score does not establish] Prior testing / change plan: [same form, set, procedure, norms; reliable change frame; no cross-form arithmetic] Plan / integration: [ ] Clinician signature / credentials: Date:
Free to use and share, no signup. The PDF includes a one-page cheat sheet with element-by-element pitfalls and a pre-sign checklist; the DOCX is the blank documentation block, ready to adapt. Neither reproduces pictures, target words, record forms, or norm tables.
Scenario: a full-form administration inside a memory-clinic dementia evaluation, documented with the cue arithmetic explicit, the norms named, and the retrieval-versus-semantic reasoning shown. All details are fictional.
Patient: E.S., 74 · Setting: Outpatient neuropsychology, memory clinic evaluation · Clinician: K. Whitfield, PhD · Note date: 08/19/2026
Measure and administration: Boston Naming Test, Second Edition, 60-item form, upgraded stimulus set, administered in English (the patient's first and dominant language) using the all-items procedure with standard response intervals and the authorized cue sequence. Corrected vision worn and adequate for the drawings; hearing adequate for cues; no motor-speech limitation. Administered and scored by this examiner within a three-hour dementia evaluation battery.
Results: Spontaneous correct 40; correct after stimulus clarification 3 (each following an initial visual misperception); countable total 43/60. Seven additional targets were retrieved only after phonemic cues and are recorded here but excluded from the total. Errors were predominantly circumlocutions conveying knowledge of the object without its name, with four semantic paraphasias on lower-frequency items and two visual misperceptions; no perseverations. Scored against the 2024 Mayo regression norms with age, sex, and education entered as that model specifies (age 74, 14 years of education, documented in the scoring record): performance fell below the expected range.
Interpretation: Clinically significant naming weakness for age and education. The shape of the performance matters as much as the level: the strong phonemic-cue benefit and knowledge-revealing circumlocutions indicate that lexical retrieval, more than semantic storage, is failing, a pattern consistent with the amnestic-plus-language presentation elsewhere in the battery and distinct from the semantic degradation profile in which cues stop helping. The naming score is one language marker within the evaluation and is not independently diagnostic of a neurodegenerative disorder.
Validity and context: Standard administration, dominant language, adequate sensory function, engaged effort with unremarkable validity indicators across the battery: result interpretable. Lifelong US resident with schooling in English; no cultural or generational familiarity concern was evident for this stimulus set, and the upgraded set was used and is named here so any future comparison holds it constant.
Comparison and plan: No prior naming testing exists. For future monitoring, the same 60-item form, stimulus set, all-items procedure, and norm system should be repeated with a reliable-change frame; a short form, and any CERAD 15-item administration inside a research protocol, would be reported as a separate measure rather than compared numerically with today's total. Findings are integrated with the memory, fluency, and functional results in the evaluation summary, with feedback scheduled 08/31/2026.
This sample is fictional and for educational purposes. It does not describe a real patient or record; the scores, dates, and details are invented to show documentation structure and are not clinical guidance. No test pictures, target words, or norm-table values are reproduced.
Writing these after every session? BastionGPT drafts complete notes from bullets, dictation, or a transcript.
Generate a note from bulletsUnited States: requirements attach to the evaluation, not the test. Neuropsychological and psychological testing is billed under the 96130 to 96139 family, with CMS contractor policy requiring the record to support medical necessity and show the referral question, tests administered, scoring and interpretation work, findings, and recommendations (PAYER POLICY); speech-language evaluations run under 92521 to 92524 with their own documentation expectations, and no code names the BNT. In epilepsy surgery programs, presurgical neuropsychology is standard of care by professional guidance (CONVENTION), and naming carries specific weight: left temporal lobe epilepsy depresses BNT performance, postoperative decline after dominant-side resection is common (a systematic review found a mean 5.8-point decline exceeding reliable-change thresholds, and one series found decline in 56 percent of left-sided cases, correlated with fMRI language lateralization), so the defensible record fixes the baseline form, stimulus set, norms, language-dominance evidence, and the reliable-change method before surgery, and reports change only within that frame. In dementia work, CERAD-derived protocols use the separately permissioned CERAD 15-item naming form under CERAD norms, and the note names it as such rather than "BNT" (CONVENTION).
Canada and Australia change the funding and the documentation registers, not the test. Canadian provincial plans generally do not cover outpatient psychologist-administered testing, which runs hospital-based, insurer-funded, or private (PAYER POLICY, provincial); the BNT is Canada's most used naming measure, Canadian norms exist (Tombaugh and Hubley's age-and-education strata; Quebec-French adaptation work), and Ontario's SLP college documentation standards (effective September 2025) expect records that identify the instrument, results, interpretation, and limitations (REGULATORY STANDARD, province-specific). A 2025 Canadian study of 525 multicultural older adults found immigration history and region of origin shaped BNT performance, which is exactly the confound analysis a Canadian chart should show. In Australia, the Better Access items are treatment items and cannot be used for formal cognitive or neuropsychological assessment, so BNT-based assessment is private, NDIS, or otherwise funded (PAYER POLICY); speech pathology services run through chronic-condition management and under-25 neurodevelopmental pathways with narrow eligibility, and Speech Pathology Australia's standards expect contemporaneous, sufficient records (CONVENTION). One dated transition worth noting: arrangements created under pre-July-2025 GP management plans remain usable only through June 2027, after which the current chronic-condition pathway applies.
Rights, editions, and language are where BNT documentation is most often wrong. PRO-ED is the current US publisher and the permissions contact: reproduction of pictures, target words, record forms, or norm tables, in print, in an EHR, or on any public web tool, requires written permission, and the publisher's electronic-use terms (encrypted, access-restricted, non-downloadable) are incompatible with a public site, so a free web "BNT tool" is unauthorized on its face; the CERAD naming form is likewise excluded from the freely distributable CERAD packet and needs separate PRO-ED permission. What belongs in the chart is the patient's results and your interpretation, never scanned stimuli or completed copyrighted forms in portal-accessible locations. Edition facts to keep straight as of August 2026: the current edition is still the BNT-2 (no third edition exists), the discontinuation rule differs between the 1983 and 2001 editions, and the publisher's free stimulus upgrade replaced Item 48 while holding the maximum at 60, so legacy and upgraded administrations should be distinguished in longitudinal and forensic records. On language and culture, the test is not culture-fair and modern norms do not make it so: item-level differential functioning is documented, bilinguals underperform monolingual norms even in their dominant language, other-language responses are recorded separately absent a defined bilingual protocol, and US Alzheimer's research centers replaced the BNT with the Multilingual Naming Test in 2015 for exactly these reasons. The Boston Naming Test and the Boston Diagnostic Aphasia Examination are products of their publisher, PRO-ED, Inc., with regional distribution by PAR, Pearson Clinical, and others. BastionGPT is not affiliated with, or endorsed by, PRO-ED or any distributor. This page reproduces no test items, stimuli, norms, or scoring materials.
The numbers behind these errors are specific. In Ferman and colleagues' 1998 study, the two published interpretations of the discontinuation rule changed final scores in 31 percent of an Alzheimer's sample and by up to 16 points in normal older adults; Pedraza and colleagues found 12 items functioning differently between demographically matched groups; in matched young adults, 18 percent of one group versus 4 percent of the other fell in the "impaired" range on monolingual norms; and after dominant temporal resection, naming declined in 56 percent of left-sided cases. The BastionGPT Clinical Advisory Board sees the same errors most often in Boston Naming Test documentation reviews:
BastionGPT is specifically trained, tuned, and clinically tested on psychological and neuropsychological evaluation reports.
See how clinicians use it day to day on the AI therapy notes page.
Many BastionGPT users report saving more than 90 minutes per day on documentation.
HIPAA-compliant with a signed BAA on every plan. Your data is never used to train models. BastionGPT drafts, you review and sign.
The countable total is spontaneous correct responses plus responses correct after a stimulus (often called semantic) cue, the clarification an examiner gives when the response shows the drawing was visually misperceived; it exists to restore a valid naming opportunity. Responses retrieved only after a phonemic cue, which supplies the beginning sound of the word, are recorded separately and never added to the total, because the cue hands over part of the answer. Self-corrections count when the correct name is the final response within the allowed interval. The defensible reporting frame makes the arithmetic visible: spontaneous correct, plus stimulus-cued correct, equals the countable total out of the form maximum, with phonemic-cued successes stated separately, because the cue profile itself is clinical information: strong phonemic-cue benefit suggests retrieval difficulty with the lexicon intact, while absent cue benefit with semantic errors points toward semantic degradation.
No universal cutoff exists, and "below cutoff" with no named system is not auditable. The same raw total means different things depending on the form, the norm system, and the corrections it applies: the publisher manual norms, Tombaugh and Hubley's age-and-education strata (adults 25 to 88), Mayo's older-adult systems (MOANS, and MOAANS for African American elders), Zec's large older-adult sample, Heaton's demographically corrected system, CERAD norms for the CERAD form only, and the 2024 Mayo regression norms built from 4,428 cognitively unimpaired adults aged 30 to 91 with age, sex, and education terms. A defensible note names the system and year, records the exact inputs used, and confirms the examinee's age, education, language, and background are actually represented in that sample; the freely circulating one-line cutoffs on scoring websites do none of this.
Materially. The classic route starts mid-list, establishes a basal, and stops after a run of consecutive failures, and that run's length differs by edition: six failures in the 1983 edition, eight in the 2001 BNT-2, a change made without published explanation. Worse, the six-failure rule itself has two published interpretations, turning on whether an item retrieved after a phonemic cue interrupts the failure run: in Ferman and colleagues' study the choice changed final scores in 3 percent of 655 normal older adults but 31 percent of 140 Alzheimer's patients, with differences up to 16 points, largest in people over 80. The all-items revised procedure avoids the problem by administering all 60 items with defined intervals. So the note names the edition, the route, and the convention as applied, and any change claim holds all three constant; the publisher's stimulus upgrade (the Item 48 replacement) gets named for the same reason.
No, and one of them is the documented trap. "BNT-15" is not a single form: the BNT-2 kit's own 15-item short form, the CERAD 15-item set used in dementia protocols, and several research-derived 15-item forms coexist, and the CERAD 15 is the specific case shown not to behave like the others and not to support extrapolation to a 60-item estimate. Thirty-item halves can be prorated with small error in some studies, but only with a validated equation for that exact half and population, labeled as an estimate. The safe rules: name the form's source every time, report each form against its own norms, never subtract a short-form score from a prior full-form score, and state in the record that the results are not directly numerically comparable, describing convergence qualitatively instead.
Name a system that actually contains an 80-year-old with your patient's education and background, and say why you chose it. Candidates include Mayo's older-adult systems (MOANS above 55; MOAANS for African American elders aged 56 to 94), Tombaugh and Hubley through 88, Zec's sample through 95, and the 2024 Mayo regression norms through 91; Heaton's system tops out around 85, which puts the oldest patients at its boundary. Two cautions specific to this age band: the stop-rule discrepancies are largest in people over 80, so state the convention, and generational familiarity with the older stimulus set is a live research question, so an atypical error on a dated object is worth describing rather than just scoring. "Older-adult norms" with no source and year is not sufficient documentation.
Only with the language analysis in the record, and often something else is better. Bilinguals underperform monolingual English norms even in their dominant language, item difficulty shifts across languages (frequency, cognate status, and familiarity all move), and a name produced correctly in the other language demonstrates intact conceptual knowledge but is not silently added to the English total; either-language scoring requires a predefined bilingual protocol. The Multilingual Naming Test was built for this problem, aligns better with bilingual dominance measures, and replaced the BNT in the US Alzheimer's centers' uniform dataset in 2015; validated adaptations and cross-cultural naming measures are the other options. If the BNT is used anyway, the note documents dominance, age of acquisition, education language, daily use, and immigration history, records other-language responses separately, and declines to diagnose anomia from monolingual norms the patient is not represented in.
Not without written permission from PRO-ED, the current US publisher and permissions contact. Reproduction of stimuli, target words, record forms, or norm tables is a case-by-case written-permission matter, and the publisher's electronic-use terms (encryption, access restriction, non-downloadable, non-printable) are incompatible with an ordinary EHR template or any public website, so a free web tool displaying BNT items or applying its norms is unauthorized on its face; the CERAD 15-item naming form likewise sits outside the freely distributable CERAD packet and needs separate PRO-ED permission, and the publisher no longer grants certain overseas translation requests. What belongs in the chart is the patient's scores, cue accounting, error pattern, and your interpretation; purchased materials stay in the authorized test-record workflow, not scanned into portal-accessible locations. Test names are used here for identification only, and this page reproduces no items or norms.
Yes. Give it the facts (form, edition, stimulus set, administration route and stop rule, spontaneous and cued counts, error types, norm system and inputs, language history, and prior testing) and it drafts the documentation block or report paragraph: cue arithmetic explicit, norms named, the retrieval-versus-semantic reasoning drafted from your error pattern, and the change rules stated, ready for your review. It can also cross-check a finished note for phonemic-cued responses in the total, short-form arithmetic, an unstated stop rule, unnamed norms, or a missing language analysis. BastionGPT is HIPAA-compliant with a signed BAA on every plan, and your data is never used to train models.
The instrument facts and compliance claims on this page trace to these sources, last verified August 2026:
Educational content, not legal or billing advice. Sample notes are fictional. Follow your organization's policies and your board, payer, and jurisdiction requirements.