CELF-5 Report Write-Up: Structure, Sample Language & Common Errors

The CELF-5 (Clinical Evaluation of Language Fundamentals, Fifth Edition) is a norm-referenced language battery for ages 5:0 through 21:11 that yields a Core Language Score and index scores from age-selected tests. Speech-language pathologists use it in school eligibility and clinical language evaluations. The write-up must name the tests behind each composite and the eligibility rule applied. This page covers how to write up CELF-5 results, with a fictional sample.

Free to use and share. No signup required.
Already have session bullets or a transcript? Generate a structured draft with BastionGPT — you review and sign it.
Who writes it

Speech-language pathologists (SLPs) in schools, clinics, and hospitals; publisher qualification level B

Audience

IEP and eligibility teams, school psychologists and psychoeducational evaluators, referring physicians and pediatricians, payers, NDIS and provincial program reviewers, families

Typical length

500 to 900 words for the results section · administration 30 to 45 minutes for the Core Language Score, 90 to 120 minutes for a full battery

Format family

Norm-referenced omnibus language battery (16 standalone tests, about 10 per age band; observational and pragmatic components)

When it's used

School speech or language impairment eligibility evaluations and reevaluations, clinical evaluations for developmental language disorder, reading and written-expression referrals needing an oral-language profile, bilingual evaluations with converging evidence, NDIS and school-support applications

Standards context

Published by NCS Pearson (2013; CELF-6 standardization research announced 2025, unreleased as of September 2026); described here for write-up purposes, no test content reproduced

What is the CELF-5?

The CELF-5 (Clinical Evaluation of Language Fundamentals, Fifth Edition; Wiig, Semel, and Secord; NCS Pearson, fall 2013) is a norm-referenced battery of spoken and written language for ages 5:0 through 21:11. It is built as 16 standalone tests, about 10 of which apply at any age, each scored as a scaled score (mean 10, SD 3). Four age-selected tests form the Core Language Score, and other combinations form the Receptive Language, Expressive Language, and Language Content indexes across the range, with a Language Structure Index at ages 5 to 8 replaced by a Language Memory Index from age 9; every composite is a standard score (mean 100, SD 15) reported with a confidence interval, a percentile rank, a Growth Scale Value for tracking change, and an age equivalent. Around the norm-referenced core sit the components that make a report defensible: the Observational Rating Scale (parent, teacher, and student observations of language at home and in class), the Pragmatics Profile (which in this edition yields a norm-referenced scaled score), the criterion-based Pragmatics Activities Checklist, and optional Reading Comprehension and Structured Writing tests from age 9. The family also includes the CELF Preschool-3 (2020; ages 3:0 to 6:11), CELF-5 Metalinguistics (2014; ages 9:0 to 21:11; a revision of the Test of Language Competence-Expanded), the CELF-5 Screening Test, a Spain-normed Spanish edition (2018; ages 5:0 to 15:11), an Australian and New Zealand edition (2017) with local census-based norms, and a Canadian French version on Q-interactive. As of September 2026 the CELF-5 remains the marketed edition; Pearson announced CELF-6 standardization field research in 2025 for ages 5 to 21, with no release date published.

The load-bearing fact for the write-up is that three different decision systems read the same number. Pearson's own severity guidance describes composite scores within 1 SD of the mean (86 to 114) as average, reports its best classification balance at a standard score of 80 (1.33 SD below the mean, with sensitivity and specificity of .97 in the publisher's study), and in the same document acknowledges that agencies qualify students at 1, 1.5, or 2 SD below the mean and that a student with real deficits may not meet a program's criterion. Federal law sets no CELF-5 cutoff at all and prohibits any single measure from being the sole criterion for a disability determination (34 CFR 300.304(b)(2)); states then write their own rules, which differ in threshold, in the number of measures required, and in whether a language sample can stand in for a second test. A second fact hides inside the score labels: the Core Language Score and the indexes are composed of different tests at ages 5 to 8, 9 to 12, and 13 to 21, so an index name without its constituent tests tells the next reader nothing about the evidence. Speech sound production belongs to the GFTA-3 page, children below the CELF-5 floor to the PLS-5 page, and the whole-report architecture to the psychoeducational report page.

Who uses the CELF-5 and when

School-based speech-language pathologists are the core users, administering the CELF-5 inside IDEA evaluations for the speech or language impairment category (34 CFR 300.8(c)(11)) and at reevaluation, where the write-up feeds the eligibility team, the IEP input, and any Section 504 statement. Clinic and hospital SLPs use it to characterize developmental language disorder and to support medical necessity for treatment, billing the untimed evaluation codes 92523 or 92522 depending on whether speech sound production was evaluated alongside language. Psychoeducational and dyslexia teams use it for the oral-language side of a reading or written-expression referral, alongside the achievement measures covered on the WIAT-4 page, and autism teams use its pragmatic components as one input to a social-communication question that the autism evaluation report owns. Outside the United States, Canadian school boards use the US-normed English edition or the Canadian French version within provincial identification processes, and Australian clinicians use the Australian and New Zealand edition to supply age-standardized evidence for NDIS and school-support applications that are judged on functional impact. The readers are therefore a team applying a rule, a psychologist integrating the profile, a payer or funder checking necessity and function, and a family entitled to a plain-language account; the write-up has to survive all four.

How to structure a CELF-5 results section

No regulation prescribes a CELF-5 report format. What the federal evaluation rules, the state criteria, and the instrument's own design dictate is the content: the edition and norm set, the tests behind every composite, uncertainty around every number compared with a threshold, the observational and pragmatic evidence in its correct category, and an eligibility statement written against the named rule. Each section below carries the pitfall that most often undermines it.

Identification, norm set, and tests administered. Open with the full test name and edition, the norm set (US, Australian and New Zealand, or Spain), the platform (paper, Q-global scoring, or Q-interactive administration), the language of testing, the evaluation dates and the student's age in years and months, and the age band that governs which tests compose each composite. List every test administered and, briefly, why: the CELF-5 is a flexible battery of standalone tests, so the report explains the selection rather than implying the whole kit was given. Pitfall: "CELF-5 administered" with no norm set and no test list. A Canadian reader assumes Canadian norms that do not exist for the English edition; a psychologist cannot tell what produced the index.

Behavioral observations and validity conditions. Document attention, effort, and fatigue; hearing and vision status; language history, English exposure, and dialect; interpreter use; any extension testing or dynamic assessment procedures; and telepractice conditions if applicable. State plainly whether the normative comparison is appropriate for this student. When it is not, say that the standardized scores are reported descriptively or not at all, and shift the evidence to language samples and converging measures. Pitfall: Second-language or dialect patterns scored as disorder. Federal rules require nondiscriminatory assessment in the language most likely to yield accurate information, and state rules (Georgia's, for one) exclude normal second-language acquisition from the impairment definition.

Composite results with confidence intervals. Report the Core Language Score and each relevant index as a standard score (mean 100, SD 15) with its 95 percent confidence interval, its percentile rank, the publisher's descriptor in prose (within 1 SD of the mean, 86 to 114, is average), and the tests that composed it at this age band. Report the indexes that answer the referral question; an index built from tests that were not administered does not exist for this student and is not estimated. Pitfall: A bare 78 compared with a 1.5 SD threshold of 77.5. The interval is the honest sentence, and the publisher itself ties confidence intervals to eligibility and placement decisions.

Test-level pattern. Describe scaled-score patterns (mean 10, SD 3) that converge, in prose, rather than listing every test as high or low. When an index discrepancy will be interpreted, use the manual's significance and base-rate procedures and say so; the same test contributes to more than one index, so a visible numerical gap is not automatically a clinical finding. Age equivalents, if an agency requires them, appear here with an explicit limitation and never as the headline. Pitfall: A low Language Memory Index written up as a memory disorder, or ordinary scatter narrated as a diagnostic profile. The indexes are descriptive groupings of overlapping tasks, not separate disorders.

Contextual and pragmatic evidence. Summarize the Observational Rating Scale by rater and by context (where communication breaks down and with what classroom consequence), report the Pragmatics Profile as the norm-referenced scaled score it is in this edition, and report the Pragmatics Activities Checklist, when given, as a criterion-based observation of authentic interaction. Add the language sample (many US rules expect at least 50 utterances) and any dynamic assessment response, because these are what carry the report when an average composite meets a struggling student. Pitfall: The Pragmatics Profile called criterion-referenced, or the three components treated as interchangeable measures of pragmatics. They are three kinds of evidence, and a rule that counts measures may treat them differently.

Eligibility or diagnostic statement against the named rule. Quote the governing criterion (state regulation, board policy, funder guidance, or payer policy), compare each element of the evidence with it, and state whether the criteria are met, whether adverse educational effect or functional impact is documented, and that the team makes the determination. In a clinic, the same paragraph states the diagnosis and the functional limitations that establish medical necessity. Pitfall: "A Core Language Score below 85 qualifies the student." No federal rule, no state rule quoted, and not the publisher's own cut, which is 80 in its classification study.

Recommendations and progress plan. Link each documented weakness to a service, accommodation, or classroom strategy, and name how progress will be measured: Growth Scale Values for change over repeated CELF-5 administrations, language-sample measures, and classroom data, with a reevaluation point. Close with the referrals the profile implies (audiology, reading evaluation, psychology) and a plain-language summary the family can use. Pitfall: Recommendations copied from a template with no thread back to the profile, or age-normed standard scores used to claim progress when the question was within-student growth.

Blank template (copy and adapt)

CELF-5 RESULTS SECTION SKELETON
Student: [initials]   Age: [y:m]   Grade: [ ]   Evaluation date(s): [ ]
Evaluator: [name, credentials]   Referral question: [ ]
Instrument: Clinical Evaluation of Language Fundamentals, Fifth Edition
   Norm set: [US / Australian and NZ / Spain]   Platform: [paper / Q-global
   scoring / Q-interactive]   Language of testing: [ ]   Age band: [5 to 8 /
   9 to 12 / 13 to 21]
Tests administered + rationale: [each test by name; why selected]
Other measures: [language sample: context, utterance count, analysis;
   dynamic assessment; records; second standardized measure if required]
Behavioral observations + validity: [attention, effort, hearing and vision,
   language history and dialect, interpreter, extension testing;
   normative comparison appropriate: yes / no, and why]
Composite results (each with uncertainty):
   Core Language Score: SS [ ] (95% CI [ ] to [ ]), percentile [ ],
   descriptor [ ]; composed of [tests at this age band]
   [Index]: SS [ ], CI [ ], percentile [ ], descriptor [ ]; composed of [ ]
Test-level pattern: [converging strengths and weaknesses in prose;
   discrepancy analysis per the manual; age equivalents only if required,
   with limitation stated]
Contextual + pragmatic evidence: [ORS by rater and context; Pragmatics
   Profile scaled score; PAC criterion result if given; language sample
   findings; dynamic assessment response]
Eligibility / diagnostic statement: [governing rule quoted; each element
   compared; adverse effect or functional impact; team determination]
Recommendations + progress plan: [weakness to service or accommodation;
   progress measure (GSVs, sample measures, classroom data); reevaluation;
   referrals; family summary]
Evaluator signature / credentials:            Date:

Free to use and share, no signup. The PDF includes a one-page cheat sheet with section-by-section pitfalls and a pre-sign checklist; the DOCX is the blank results-section skeleton, ready to adapt. Neither reproduces test items, stimuli, record forms, norms, or conversion tables.

Sample CELF-5 write-up (fictional)

Scenario: a second-grader referred for reading difficulty and trouble following classroom directions, evaluated by a school SLP, with composites reported with intervals and constituent tests, a language sample as the second measure, and the eligibility statement written against a named state-style criterion. All details are fictional.

Patient: J.M., 7 years 4 months, grade 2  ·  Setting: Public school speech-language eligibility evaluation  ·  Clinician: A. Okafor, MS, CCC-SLP  ·  Note date: 09/16/2026

Measures and conditions: Clinical Evaluation of Language Fundamentals, Fifth Edition (CELF-5), US edition and norms, paper administration scored on Q-global, in English, across two sessions on 09/10/2026 and 09/14/2026. J.M. was 7 years 4 months at testing, so the 5 to 8 age band governs test composition. Tests administered: Sentence Comprehension, Linguistic Concepts, Word Structure, Word Classes, Following Directions, Formulated Sentences, and Recalling Sentences, selected to obtain the Core Language Score and the Receptive Language, Expressive Language, Language Content, and Language Structure indexes; the Pragmatics Profile was completed by the classroom teacher, and Observational Rating Scale forms were returned by the teacher and a parent on 09/11/2026. A 62-utterance language sample (conversation and a story retell) was recorded and transcribed on 09/14/2026 as the second measure. J.M. is a monolingual English speaker of the local community dialect; hearing was screened and passed at school on 09/08/2026, and vision is corrected with glasses that were worn. Attention was adequate with one movement break per session, and the normative comparison is judged appropriate.

Composite results: The Core Language Score was 76 (95 percent confidence interval 70 to 82; 5th percentile), below the average range, composed at this age of Sentence Comprehension, Word Structure, Formulated Sentences, and Recalling Sentences. The Expressive Language Index was 74 (68 to 80; 4th percentile), from Word Structure, Formulated Sentences, and Recalling Sentences. The Receptive Language Index was 84 (77 to 91; 14th percentile), from Sentence Comprehension, Word Classes, and Following Directions, at the lower edge of average with an interval that spans it. The Language Content Index was 88 (81 to 95; 21st percentile), from Linguistic Concepts, Word Classes, and Following Directions, within the average range. The Language Structure Index, which at ages 5 to 8 draws on the same four tests as the Core Language Score, was 76 (70 to 82; 5th percentile). Descriptors follow the publisher's convention that composite scores within 1 SD of the mean are average.

Test-level pattern: Scaled scores (mean 10, SD 3) converged on a structural-language weakness: Word Structure 5, Recalling Sentences 5, Formulated Sentences 6, and Following Directions 6 were all more than 1 SD below the mean, while Linguistic Concepts 9, Word Classes 8, and Sentence Comprehension 7 sat at or near the average range. The content-versus-structure contrast (Language Content Index 88 against Language Structure Index 76) was tested with the manual's index-comparison procedure and met its significance criterion; it is interpreted as weak morphosyntax and sentence-level formulation and recall with relatively preserved vocabulary knowledge, not as a discrete disorder of any single index. Age equivalents are not reported; standard scores, intervals, and percentiles carry the comparison.

Contextual and pragmatic evidence: On the Observational Rating Scale, the teacher and the parent both marked frequent difficulty following multistep spoken directions, retelling events in order, and producing complete sentences in writing; the teacher tied these to missed instructions during whole-class lessons and incomplete written responses. The Pragmatics Profile, a norm-referenced scaled score in this edition, was 9, within the average range: conversational turn-taking, topic maintenance, and nonverbal communication are age-appropriate, so the concern is structural language, not social communication. The 62-utterance sample showed a mean length of utterance below expectations for age, inconsistent regular past-tense and plural marking, and few complex sentences; in a brief dynamic teaching trial, J.M. followed two-step directions accurately when they were paired with a visual sequence and repeated once.

Eligibility analysis: Our state criterion requires a score at least 1.5 standard deviations below the mean, or below the 7th percentile, on two or more standardized language tests in morphology, syntax, semantics, or pragmatics, or on one such test together with a representative language sample of at least 50 utterances. The CELF-5 is one standardized test; on it the Expressive Language Index (74; interval 68 to 80) sits entirely below the 1.5 SD line of 77.5, and the Core Language Score (76) sits below it with an interval whose upper bound crosses it, which is why this statement rests on the convergence of the expressive index, the language sample, and the classroom evidence rather than on the Core score alone. The 62-utterance sample documents morphological and syntactic use below age expectations, satisfying the second route. Adverse educational effect is documented by the teacher's observations and by the reading evaluation in progress with the school psychologist; these language findings contribute to that evaluation but do not by themselves establish a specific learning disability. No single measure has been used as the sole criterion, and eligibility under the speech or language impairment category is the team's determination.

Recommendations and progress plan: Recommended: direct language intervention targeting sentence formulation, grammatical morphology, and recall of spoken directions, coordinated with the reading evaluation so that oral-language and decoding goals are written together; classroom accommodations of chunked and visually sequenced directions, a repeat-and-check routine before independent work, and sentence starters for written responses. Progress will be measured with CELF-5 Growth Scale Values at reevaluation, quarterly language-sample measures on the same tasks, and teacher data on direction-following, with reevaluation planned within the state timeline. Reviewed with J.M.'s parents on 09/16/2026 in plain language: J.M. understands words and ideas well, and what is hard right now is putting sentences together, using word endings, and holding on to long spoken directions, all of which are teachable.

This sample is fictional and for educational purposes. It does not describe a real student or record; the scores, dates, and details are invented to show write-up structure and are not clinical guidance. Scores are invented for illustration and correspond to no real child or record, and no norm-table values are reproduced.

↑ Back to the template and downloads

Why this sample works

  • The edition, norm set, platform, age band, and every test administered are stated before any score, so a psychologist, a receiving district, or a Canadian or Australian reader can tell exactly what produced each number.
  • Every composite carries its confidence interval, percentile, and constituent tests, and the one interval that crosses the threshold is named rather than hidden behind the integer.
  • The Pragmatics Profile is reported as the norm-referenced scaled score it is, the Observational Rating Scale as rater-by-context observation, and the language sample as the second measure, each in its own evidentiary category.
  • The eligibility statement quotes the rule, walks through each element, keeps the reading question with the reading evaluation, and leaves the determination to the team, which is what 34 CFR 300.304 requires.
  • Recommendations trace back to the documented pattern, and the progress plan names Growth Scale Values for within-student change instead of re-comparing age-normed scores.

Writing these after every session? BastionGPT drafts complete notes from bullets, dictation, or a transcript.

Generate a note from bullets

Documentation and compliance considerations

United States: the write-up sits inside three decision systems, and the report should say which one each sentence serves. Under IDEA, speech or language impairment is a disability category defined as a communication disorder that adversely affects educational performance (34 CFR 300.8(c)(11)), and the evaluation rules require a variety of tools and strategies including parent information, prohibit any single measure as the sole criterion, require nondiscriminatory instruments administered in the child's native language or the mode most likely to yield accurate information, and require assessment in all areas related to the suspected disability (34 CFR 300.304(b) and (c)); no federal CELF-5 cutoff exists (LAW). States supply the operational rule and they differ: California requires at least 1.5 SD below the mean or below the 7th percentile on two or more standardized tests in morphology, syntax, semantics, or pragmatics, or on one test plus a representative language sample of at least 50 utterances (5 CCR 3030(b)(11)(D)); New Jersey requires functional assessment outside the testing situation plus performance below 1.5 SD or the 10th percentile on at least two standardized language tests, one a comprehensive receptive and expressive measure, and its April 17, 2024 clarification states that index, subtest, and standard scores may all be considered (N.J.A.C. 6A:14-3.5(c)4 and 3.6(a)); Georgia requires at least two measures or procedures, at least one formal, plus documented adverse effect, and excludes normal second-language acquisition and dialect difference from the impairment definition (Rule 160-4-7-.05) (LAW, state by state). The publisher's guidance is a third voice: its best classification balance sits at a standard score of 80, with sensitivity and specificity of .97 in its own study, and it acknowledges that programs use 1, 1.5, or 2 SD (publisher guidance), while the familiar "below 85" threshold is a local CONVENTION wherever it is not written into a rule. In clinics, coverage turns on documented clinical history, objective findings, functional limitation, and medical necessity under Medicare's therapy documentation policy and plan-specific commercial rules, billed through the untimed evaluation codes 92523 or 92522; a standard score supports the case and never makes it (PAYER POLICY).

Canada and Australia change both the norm set and the frame. Pearson Canada sells the English CELF-5 with US norms (based on the 2010 US Census) and a Canadian French version on Q-interactive, so a Canadian report names which one was used and, for the English edition, discloses the normative population; identification of a communication exceptionality is a provincial and board process (Ontario's identification, placement, and review committee route, British Columbia's designation and IEP orders) with no national numeric criterion, so any local threshold is labeled a board CONVENTION unless a provincial policy source is cited (LAW and PROGRAM POLICY). Australia has its own edition, the CELF-5 Australian and New Zealand (2017), normed on the 2011 Australian and 2013 New Zealand censuses, and NDIS evidence is judged on functional impact across communication, social interaction, and learning rather than on a cutoff (PAYER POLICY); state school-support criteria could not be verified from primary sources for this page and should be taken from the funder's current document. Edition and norms: the CELF-5 (2013) is the current marketed edition as of September 2026, Pearson has publicly recruited for CELF-6 standardization since 2025, and no release date is published, so a report names the edition and platform and dates itself. The Spanish product on the US store is the Spain-normed edition (2018; ages 5:0 to 15:11; Spanish, not Catalan, as the home language), not a US Hispanic norm set, and English-edition standard scores are not diagnostic for a student the norming sample does not represent. The publisher classifies the CELF-5 at qualification level B; record forms and stimulus materials are protected and are not scanned, photographed, or reproduced in a report, transferring results into an electronic record and placing scoring on another platform are permission and license matters with Pearson, and a free public web scorer has no authorization. A report may always contain the student's derived scores and your interpretation in your own words.

CELF-5 is a trademark of NCS Pearson, Inc. (the CELF word mark is US registration 1776508, held by NCS Pearson, Inc., of Bloomington, Minnesota, and renewed in 2023). BastionGPT is not affiliated with, or endorsed by, the publisher. This page reproduces no test items, stimuli, norms, or scoring materials.

↑ Back to the template and downloads

Common CELF-5 write-up errors reviewers flag

The numbers behind these errors are specific. Pearson's own severity guidance puts its best classification balance at a standard score of 80 (sensitivity and specificity of .97 in the publisher's study) while acknowledging that programs qualify students at 1, 1.5, or 2 SD below the mean; California requires 1.5 SD or the 7th percentile on two standardized tests, or one plus a 50-utterance language sample; New Jersey requires 1.5 SD or the 10th percentile on two tests, one of them a comprehensive receptive and expressive measure; Georgia requires at least two measures or procedures, at least one formal; the battery itself is 16 tests of which about 10 apply at any age, with the Core Language Score built from four different tests in each of three age bands; and 34 CFR 300.304(b)(2) prohibits any single measure from deciding. The BastionGPT Clinical Advisory Board sees the same errors most often in CELF-5 documentation reviews:

  • The 85 rule written as the test's rule. "A Core Language Score below 85 indicates a language disorder and qualifies the student." No federal rule says so, the publisher's own studied cut is 80, and the state may require 1.5 or 2 SD on two measures. Write the number, its interval, and the named rule, and let the team apply it.
  • Composites without their constituent tests. "Receptive Language Index 84" with no tests named. The index is built from different tests at ages 5 to 8, 9 to 12, and 13 to 21, and the Language Structure Index gives way to a Language Memory Index at 9, so the label alone hides what evidence produced the number. Name the tests behind every composite at the student's age band.
  • Point scores at a threshold with no interval. A 78 written against a 77.5 line as if the integer were exact. The publisher reports confidence intervals and ties them to classification and placement decisions; a report says where the interval falls relative to the criterion and, when it straddles the line, rests the conclusion on converging evidence.
  • Index scatter narrated as a diagnosis. A low Language Memory Index becomes a "memory disorder," and a receptive-expressive gap becomes "the diagnosis." The same test feeds more than one index, gaps need the manual's significance and base-rate procedures before they mean anything, and the indexes are descriptive groupings, not separate conditions.
  • The Pragmatics Profile mislabeled. Called criterion-referenced, or lumped with the Observational Rating Scale and the Pragmatics Activities Checklist as one "pragmatics measure." In this edition the Profile yields a norm-referenced scaled score, the Checklist is the criterion-based authentic observation, and the Rating Scale is rater-by-context observation; a rule that counts measures may treat each differently, so the report labels each correctly.
  • Language age doing decisional work. "J. is functioning at the level of a five-year-old." Age equivalents show no rank among peers, move sharply with small raw-score changes, and are not comparable across tests, and the publisher advises against using them as the basis for diagnosis or placement. Report them only where a form requires them, with the limitation stated.
  • English norms applied to an unrepresented student. Standard scores from the English edition reported as diagnostic for a bilingual or dialect-speaking student, or the Spain-normed Spanish edition treated as a US norm set. Federal rules require assessment in the language most likely to yield accurate information, some states exclude normal second-language acquisition from the definition outright, and the defensible report presents such tasks descriptively while resting the conclusion on the dominant language, a language sample, and converging evidence.
How BastionGPT helps

BastionGPT is specifically trained, tuned, and clinically tested on speech-language evaluation reports.

  • Give it the facts (age band and tests administered, scaled and standard scores with intervals and percentiles, Observational Rating Scale and Pragmatics Profile results, language-sample findings, the governing rule) and it drafts the results section: constituent tests named per composite, intervals compared with the threshold, the pragmatic components in their correct categories, and the eligibility statement written against the rule, ready for your review.
  • Cross-check a finished report for the gaps reviewers flag: a composite without its tests, a bare score at a threshold, the Pragmatics Profile called criterion-referenced, an age equivalent doing decisional work, or an eligibility sentence that names no rule.
  • Draft the companion paragraphs: the plain-language family summary, the bilingual-evaluation validity statement, or the progress paragraph that uses Growth Scale Values instead of re-comparing age-normed scores.

See how clinicians use it day to day on the AI therapy notes page.

Many BastionGPT users report saving more than 90 minutes per day on documentation.

HIPAA-compliant with a signed BAA on every plan. Your data is never used to train models. BastionGPT drafts, you review and sign.

Frequently asked questions

Each test yields a scaled score with a mean of 10 and SD of 3, and the composites (the Core Language Score and the Receptive Language, Expressive Language, Language Content, and Language Structure or Language Memory indexes) are standard scores with a mean of 100 and SD of 15, each reported with a confidence interval, a percentile rank, a Growth Scale Value, and an age equivalent. In the publisher's own prose, composites within 1 SD of the mean (86 to 114) are average and scores below that are below average to very low relative to age peers, which may or may not affect classroom participation; a scaled score of 7 sits 1 SD below the test mean. The interval matters more than the label whenever a score is compared with a rule, and the constituent tests matter more than the index name, because composition changes at ages 9 and 13. Growth Scale Values answer a different question (how much the student's own ability changed) and are not peer rankings.

Whatever the governing rule says, and there are three voices to keep apart. Federal law sets no cutoff and prohibits any single measure from being the sole criterion (34 CFR 300.304(b)(2)) (LAW). States set the operational rule: California uses 1.5 SD or the 7th percentile on two standardized tests, or one plus a 50-utterance language sample; New Jersey uses 1.5 SD or the 10th percentile on two tests, one comprehensive; Georgia requires two measures or procedures, one formal, plus adverse effect (LAW, state by state). The publisher reports its best classification balance at a standard score of 80 and acknowledges that programs use 1, 1.5, or 2 SD (publisher guidance), and the familiar 85 is a CONVENTION unless a rule adopts it. So the report quotes the rule, compares the interval with it, documents adverse educational effect, and leaves the determination to the team.

Norm-referenced. Pearson lists norm-referenced Pragmatics Profile scores among the fifth-edition changes, and the product page states that the Profile now produces scaled scores. The criterion-based pragmatic component is the Pragmatics Activities Checklist, an authentic interaction observation, and the Observational Rating Scale is a rater-by-context observation of listening, speaking, reading, and writing at home and school. Getting the labels right matters when a state rule counts measures or formal measures: a norm-referenced Profile score may count differently from a checklist result, and the report should not present the three as interchangeable measures of pragmatics.

Not as an interpretive result. The publisher's own guidance is that age equivalents show no rank among age peers, can shift substantially with small raw-score differences, are not comparable across tests, and do not mean the student functions overall like a child of that age, and it advises against using them as the primary basis for diagnosis or placement. Standard scores with confidence intervals and percentiles carry the peer comparison; Growth Scale Values carry change over time. If a form or district requires an age equivalent, report it with that limitation written next to it, and never write that the student's language is at the level of a younger child.

Start with language history and the normative question: the English edition's norms assume English-dominant experience, so if the student is not represented, the standard scores are reported descriptively or withheld, and the conclusion rests on assessment in the dominant language, a language sample in each language where feasible, dynamic assessment, and converging parent and teacher evidence. Federal rules require instruments that are not racially or culturally discriminatory and administration in the native language or mode most likely to yield accurate information (34 CFR 300.304(c)(1)), and states such as Georgia exclude normal second-language acquisition and dialect difference from the impairment definition. Do not reach for the Spanish CELF-5 on the US store as a fix: it is the Spain-normed edition (2018; ages 5:0 to 15:11), a different normative population rather than a US bilingual norm set.

As of September 2026 the CELF-5 (2013) is the edition Pearson markets in the US, Canada, and Australia. Pearson published a CELF-6 standardization field-research notice in 2025 recruiting English-speaking participants aged 5 to 21, which means a sixth edition is in development, and no release date has been published. A report therefore names the edition, the norm set, and the platform, and carries its own date. When a new edition arrives, the defensible practice is to state which edition and norms were used and, for reevaluations, to note that scores across editions are not directly comparable and that Growth Scale Values are edition-specific.

By referral question and developmental level, not birthday. The CELF Preschool-3 (2020) covers 3:0 to 6:11 and the CELF-5 begins at 5:0, so five- and six-year-olds fall in both: a child already facing school-age eligibility questions usually matches the CELF-5's Core and index architecture, while a developmentally younger child with broad delay matches the Preschool-3 or the play-based PLS-5, covered on the PLS-5 page. The CELF-5 assesses language, not speech sound production; an articulation or phonology question needs the GFTA-3, covered on the GFTA-3 page, and the two are written up in separate sections even when billed in one evaluation. Reading and written-expression questions add achievement measures such as the WIAT-4; the CELF-5 explains the oral-language substrate but does not establish a specific learning disability.

No, and the report does not need them. Record forms, stimulus materials, scoring keys, and norms tables are the publisher's protected content at qualification level B; a report contains the student's derived scores, the tests administered by name, your observations, and your interpretation in your own words, which is exactly the content downstream readers use. Pearson's permissions process covers transfer of results into an electronic record and any placement of an assessment on another platform, and describing a task at the topic level (for example, recalling sentences of increasing length) is fine while quoting an item is not. Test security is also a credibility matter: a team that sees leaked stimuli in a report has reason to doubt the scores that followed.

Yes. Give it the facts (age band and tests administered, scaled and standard scores with intervals and percentiles, Observational Rating Scale and Pragmatics Profile results, language-sample findings, observations, and the governing rule) and it drafts the results section: constituent tests named per composite, intervals compared with the threshold, the pragmatic components labeled correctly, the eligibility statement written against the rule, and a plain-language family summary, ready for your review. It can also cross-check a finished report for a composite without its tests, a bare score at a threshold, a mislabeled Pragmatics Profile, an age equivalent doing decisional work, or an eligibility sentence with no rule. BastionGPT is HIPAA-compliant with a signed BAA on every plan, and your data is never used to train models.

Primary sources

The instrument facts and compliance claims on this page trace to these sources, last verified September 2026:

  1. Publisher record, accessed September 2026: NCS Pearson, CELF-5 product page (fall 2013; ages 5:0 to 21:11; 30 to 45 minutes for the Core Language Score; qualification level B; 16 tests, about 10 per age; scores and platforms; Pragmatics Profile scaled scores), What's New About CELF-5 (norm-referenced Pragmatics Profile; Pragmatics Activities Checklist; Growth Scale Values; optional tests), Determining the Severity of a Language Disorder (average range; optimal cut of 80; agency criteria of 1, 1.5, or 2 SD), and CELF-6 standardization field-research notice (2025; ages 5 to 21).
  2. Family and international editions, accessed September 2026: Pearson, CELF Preschool-3 (2020; ages 3:0 to 6:11), CELF-5 Metalinguistics (2014; ages 9:0 to 21:11), and CELF-5 Spain version (2018; ages 5:0 to 15:11; Spain norms); Pearson Canada, CELF-5 (US norms; French Canadian version on Q-interactive); Pearson Australia and New Zealand, CELF-5 Australian and New Zealand edition (2017; 2011 Australian and 2013 New Zealand census-based norms).
  3. Federal law: 34 CFR 300.8(c)(11) (speech or language impairment definition) and 34 CFR 300.304 (variety of tools; no single measure as sole criterion; nondiscriminatory assessment in the native language; all areas of suspected disability), via Cornell LII.
  4. State rules: California, 5 CCR 3030(b)(11)(D); New Jersey Department of Education, Speech-Language Services Eligibility broadcast, April 17, 2024 (N.J.A.C. 6A:14-3.5(c)4 and 3.6(a)); Georgia Department of Education, Rule 160-4-7-.05, Appendix (j).
  5. Independent reviews: Coret MC and McCrimmon AW, 2015, Journal of Psychoeducational Assessment 33(5), 495 to 500, CELF-5 test review (doi 10.1177/0734282914557616; 30 to 45 minutes for the core, 90 to 120 minutes for a full battery); Columbia University LEADERS Project, CELF-5 test review (2014; classification-sample and cut-score critique; second-language misidentification risk).
  6. Payer and professional context: CMS, Medicare Benefit Policy Manual, Chapter 15, section 220.3 (therapy documentation requirements); ASHA Practice Portal, Spoken Language Disorders (assessment components including language sampling and dynamic assessment).
  7. Rights and trademark: Pearson, Permissions and Licensing (electronic-record transfer and commercial platform licensing); USPTO TSDR, CELF, US registration 1776508 (NCS Pearson, Inc.; renewed 2023).
  8. Canada and Australia: Ontario Ministry of Education, Special Education in Ontario, Kindergarten to Grade 12: Policy and Resource Guide; NDIS, supporting-evidence guidance for health professionals (functional impact across communication, social interaction, and learning).

Educational content, not legal or billing advice. Sample notes are fictional. Follow your organization's policies and your board, payer, and jurisdiction requirements.