GFTA-3 Report Write-Up: Structure, Sample Language & Common Errors

The GFTA-3 (Goldman-Fristoe Test of Articulation, Third Edition) is a norm-referenced single-word articulation test for ages 2:0 through 21:11 that scores consonant production against same-age, same-sex norms. Speech-language pathologists use it in school eligibility and clinical speech sound evaluations. The write-up must pair the score with pattern, stimulability, and connected-speech evidence. This page covers how to write up GFTA-3 results, with a fictional sample.

Free to use and share. No signup required.
Already have session bullets or a transcript? Generate a structured draft with BastionGPT — you review and sign it.
Who writes it

Speech-language pathologists (SLPs) in schools, clinics, hospitals, and early-intervention programs; publisher qualification level B

Audience

IEP and eligibility teams, school psychologists and psychoeducational evaluators, referring physicians and pediatricians, payers, NDIS and provincial program reviewers, families

Typical length

400 to 800 words for the results section · administration 5 to 15 minutes for Sounds-in-Words, longer with Sounds-in-Sentences, stimulability, and a connected-speech sample

Format family

Norm-referenced single-word articulation test (Sounds-in-Words core; separately normed Sounds-in-Sentences; criterion-referenced stimulability; KLPA-3 phonological-process companion)

When it's used

School speech or language impairment eligibility evaluations and reevaluations for articulation and phonology concerns, clinical speech sound disorder evaluations, preschool and early-intervention speech referrals, bilingual evaluations alongside the Spanish edition, progress reevaluations

Standards context

Published by NCS Pearson (2015; Spanish edition 2017; no fourth edition announced as of September 2026); described here for write-up purposes, no test content reproduced

What is the GFTA-3?

The GFTA-3 (Goldman-Fristoe Test of Articulation, Third Edition; Goldman and Fristoe; NCS Pearson, 2015) is a norm-referenced test of consonant articulation for ages 2:0 through 21:11. Its core, Sounds-in-Words, elicits single consonants in initial, medial, and final positions plus consonant clusters through picture naming in 5 to 15 minutes; the raw score is the number of errors (substitutions, omissions, and distortions count; recognized dialectal variations do not), converted through separate male and female norm tables into a standard score (mean 100, SD 15) with a 95 percent confidence interval, a percentile rank, a test-age equivalent, and a Growth Scale Value. Sounds-in-Sentences is a structured sentence-imitation task (the examiner reads a short story, then the child repeats each sentence) with its own standard score on its own scale from age 4, and the examiner's rating of each repeated sentence produces the record-form intelligibility rating. The stimulability section records whether error sounds can be imitated in syllables, words, and sentences; the third edition added vowel and R error analyses, two picture sets (ages 2 to 6 and 7 to 21), and dialect-sensitive scoring for American English dialects and English influenced by other languages; and it runs on paper, Q-global, and Q-interactive. The GFTA-3 Spanish (2017) is an adaptation with its own norms based on Spanish speakers living in the US and Puerto Rico, not a translation. As of September 2026 the 2015 edition remains current on Pearson's US, Canadian, and Australian sites, with no fourth edition announced.

The load-bearing distinction for the write-up is articulation versus phonology, and the GFTA-3 sits on one side of it. Its Sounds-in-Words score counts how many consonants were wrong in single words against same-age, same-sex norms; it does not say whether the errors are organized. Whether they form rule-governed patterns (velar fronting, cluster simplification, deletion of final consonants) is a phonological question, and Pearson answers it with a companion product, the KLPA-3 (Khan-Lewis Phonological Analysis, Third Edition; Khan and Lewis, 2015), which re-analyzes the same Sounds-in-Words responses into core developmental and supplemental atypical processes, each by percent of occurrence, with its own standard score and percentile by age and sex; no second administration is involved, and the companion is English only. So a report that gives a GFTA-3 score and then diagnoses a phonological disorder has skipped the construct that defines the diagnosis, and a report that treats the KLPA-3 as a second test has double-counted one set of responses. The defensible write-up keeps two sections with two jobs: the articulation section (sounds and positions in error, error type, consistency, stimulability, and the norm-referenced score) and the phonology section (which patterns, whether they are developmental at this age, and how much of the intelligibility problem they explain), and it derives the treatment approach from whichever mechanism the evidence supports. Two further things the GFTA-3 does not measure need their own sections: intelligibility (the record-form rating is the examiner's judgment of imitated sentences summarized as a percentage, not a norm-referenced score) and spontaneous connected speech (the sentence task is imitation, and Pearson states its raw score is not comparable with Sounds-in-Words). Language belongs to the CELF-5 page and the PLS-5 page, and the whole-report architecture to the psychoeducational report page.

Who uses the GFTA-3 and when

School-based speech-language pathologists are the core users, giving the GFTA-3 inside IDEA evaluations for the speech or language impairment category, whose federal definition expressly includes impaired articulation (34 CFR 300.8(c)(11)), and at reevaluation, where the write-up feeds the eligibility team, the IEP input, and any Section 504 statement. Clinic, hospital, and private-practice SLPs use it to characterize a speech sound disorder and to support medical necessity for treatment, billing the untimed evaluation codes 92522 (speech sound production) or 92523 (speech sound production with language). Early-intervention and preschool teams use it from age 2:0 alongside the broader picture covered on the developmental assessment and ASQ-3 pages, and bilingual evaluators pair the English edition with the GFTA-3 Spanish. Outside the United States, Canadian clinicians use the same US-normed edition inside provincial preschool programs and school-board processes, and Australian clinicians use it, with no Australian normative study, to supply standardized evidence for NDIS and school-support applications that are judged on functional impact. Each of those readers asks a different question of the same paragraph: does the criterion hold, does the speech section agree with the language section, is treatment necessary, and what does it mean for this child. The write-up is defensible when it answers all four without changing the facts.

How to structure a GFTA-3 results section

No regulation prescribes a GFTA-3 report format. What the federal evaluation rules, the state criteria, and the instrument's own design dictate is the content: the sections administered and the norm group, uncertainty around every number compared with a threshold, an articulation section and a phonology section built from one set of responses, intelligibility by a named method, a developmental comparison against a named source, and an eligibility statement written against the named rule. Each section below carries the pitfall that most often undermines it.

Identification, sections administered, and norm group. Open with the full test name and edition, the sections administered (Sounds-in-Words, Sounds-in-Sentences, stimulability, the vowel analysis) and whether the KLPA-3 analysis was run, the platform (paper, Q-global scoring, or Q-interactive administration) and mode (in person or telepractice, following Pearson's guidance), the evaluation dates, the child's age in years and months, the norm table used (male or female) with the reason, and the language and dialect of testing. Name the rule the eligibility statement will apply. Pitfall: "GFTA-3 administered" with no sections, no norm group, and no rule. A reader cannot tell whether the sentence task was given, which table produced the percentile, or what criterion the conclusion answers.

Validity conditions, language background, and oral mechanism. Document attention and effort, hearing status (screened and passed, or referred), oral structure and function relevant to speech, any motor-speech signs, the language history (languages, dominance, simultaneous or sequential learning, community dialect), interpreter use, and whether the normative comparison is appropriate: the English norm sample's bilingual members were simultaneous learners who used English most, so a sequential or English-learning child is not represented, and ASHA's position is that standard scores are not reported for a child the norm group does not represent. Pitfall: Dialect or transfer features scored as errors. Pearson's scoring counts recognized dialectal variations as correct, and its own FAQ retains the correction of a 2016 webinar slide that said otherwise.

Standard scores with uncertainty. Report the Sounds-in-Words standard score (mean 100, SD 15) with its 95 percent confidence interval, percentile rank, and named norm group, and describe the standing in prose (within 1 SD of the mean is conventionally average). Report Sounds-in-Sentences, when given from age 4, as a separate standard score on its own scale, and never compare the two raw scores, which Pearson states are not comparable. Age equivalents, if a form requires them, carry their limitation and never lead. Pitfall: A bare 78 written against a threshold as if the integer were exact, or a Sounds-in-Sentences raw score set beside the Sounds-in-Words raw score as evidence of a context effect.

Error inventory: the articulation section. Describe the errors by sound class and position (which consonants and clusters; initial, medial, final), by type (substitution, omission, distortion, including lateral or dentalized sibilants), and by consistency, adding the vowel and R analyses when relevant as descriptive findings. This is where a motor-based error such as a lateral lisp lives. Never list target words or item-level responses; the pattern description says more and respects test security. Pitfall: An item-by-item word list copied from the record form. It leaks protected content, reads like a score report, and still does not say whether the errors are organized.

Pattern analysis: the phonology section (KLPA-3). Report the KLPA-3 as a re-analysis of the same Sounds-in-Words responses: its standard score with interval and percentile, then each core or supplemental process by percent of occurrence in prose, stating whether the pattern is developmental or atypical at the child's age and how much of the intelligibility problem it explains. This section, not the articulation score, carries a phonological diagnosis, and it decides between a motor-based and a pattern-based treatment approach. Pitfall: "Phonological disorder" diagnosed from an articulation standard score with no pattern analysis, or the KLPA-3 written up as a second test that double-counts one set of responses.

Stimulability, intelligibility, and connected speech. Report stimulability descriptively (which error sounds became accurate, at which level, with what support) as prognostic and target-selection information, never as a score. Report intelligibility by a named method: the Intelligibility in Context Scale, a Percentage of Consonants Correct from a connected-speech sample of stated size and context, or a listener-based percentage of words understood; the record-form rating is the examiner's judgment of imitated sentences and is reported as descriptive information. Then compare single-word accuracy with the conversational sample, because the two diverge in children with speech delay. Pitfall: "The GFTA-3 indicates 60 percent intelligibility." It produces no norm-referenced intelligibility score, and an average single-word score does not rule out a child nobody can understand.

Developmental comparison, eligibility statement, and recommendations. Compare each error sound or pattern with a named contemporary source (Crowe and McLeod, 2020, for US English; McLeod and Crowe, 2018, cross-linguistically), quote the governing criterion, walk through each of its elements with the evidence, document adverse educational effect or functional impact, and leave the determination to the team. Then match recommendations to the mechanism: pattern-based intervention for rule-governed processes, motor-based work for isolated distortions, targets ordered by stimulability and developmental expectation, a language evaluation when the profile suggests one, and progress measures that repeat the same probes and sample. Pitfall: "A standard score below 85 qualifies." No federal rule says so, California's articulation subsection names no number, and Wisconsin's turns on intelligibility and stimulability.

Blank template (copy and adapt)

GFTA-3 RESULTS SECTION SKELETON
Child: [initials]   Age: [y:m]   Grade: [ ]   Evaluation date(s): [ ]
Evaluator: [name, credentials]   Referral question: [ ]
Instrument: Goldman-Fristoe Test of Articulation, Third Edition
   Sections given: [Sounds-in-Words / Sounds-in-Sentences / stimulability /
   vowel analysis]   Platform + mode: [paper / Q-global / Q-interactive;
   in person / telepractice]   Norm group: [age; male or female table]
   Language + dialect of testing: [ ]   KLPA-3 analysis: [yes / no]
Other measures: [connected-speech sample: context, size, PCC; intelligibility
   measure (ICS, listener estimate); oral mechanism; hearing; language]
Validity conditions: [attention, hearing status, language history and
   bilingual status, interpreter; norm sample representative: yes / no, why]
Standard scores (each with uncertainty):
   Sounds-in-Words: SS [ ] (95% CI [ ] to [ ]), percentile [ ], norm group [ ]
   Sounds-in-Sentences (own scale, from age 4): SS [ ], CI [ ], percentile [ ]
   Age equivalent: [only if required, with limitation stated]
Error inventory (articulation): [sounds and positions in error; substitution,
   omission, distortion; consistency; vowel or R findings; no target words]
Pattern analysis (phonology): [KLPA-3 SS [ ], CI [ ], percentile [ ]; each
   process by percent of occurrence; developmental or atypical for age]
Stimulability: [sound; level (syllable, word, sentence); support needed]
Intelligibility + connected speech: [method named; sample size and context;
   PCC [ ]%; single-word versus conversational accuracy compared]
Developmental comparison: [each error against Crowe and McLeod (2020)]
Eligibility / diagnostic statement: [governing rule quoted; each element
   compared; adverse effect or functional impact; team determination]
Recommendations + progress plan: [approach matched to mechanism; targets
   ordered by stimulability; language referral if needed; progress measures]
Evaluator signature / credentials:            Date:

Free to use and share, no signup. The PDF includes a one-page cheat sheet with section-by-section pitfalls and a pre-sign checklist; the DOCX is the blank results-section skeleton, ready to adapt. Neither reproduces test items, target words, stimuli, record forms, norms, or conversion tables.

Sample GFTA-3 write-up (fictional)

Scenario: a first-grader referred by his teacher because classmates cannot understand him, evaluated by a school SLP, with articulation and phonology written as separate sections from one set of responses, intelligibility measured by named methods, and the eligibility statement written against a state-style criterion that names no cutoff. All details are fictional.

Patient: D.R., 6 years 2 months, grade 1  ·  Setting: Public school speech-language eligibility evaluation  ·  Clinician: M. Delgado, MS, CCC-SLP  ·  Note date: 09/21/2026

Measures and conditions: Goldman-Fristoe Test of Articulation, Third Edition (GFTA-3), English edition with US norms, administered in person on paper and scored on Q-global on 09/14/2026; Sounds-in-Words, Sounds-in-Sentences, and the stimulability section were given, and the Sounds-in-Words responses were analyzed with the Khan-Lewis Phonological Analysis, Third Edition (KLPA-3). D.R. was 6 years 2 months at testing and was scored against the male norm table. He is a monolingual English speaker of the local community variety; no dialectal productions were counted as errors, and the normative comparison is judged appropriate. Hearing was screened and passed at school on 09/09/2026. Oral structure and function were adequate for speech, diadochokinetic rates were age-appropriate, and no signs of a motor speech disorder were observed. A 118-word conversational and story-retell sample was recorded on 09/17/2026 and transcribed for a Percentage of Consonants Correct; the Intelligibility in Context Scale was completed by his mother on 09/14/2026, and the classroom teacher returned an observation form on 09/15/2026. Language was evaluated separately with the CELF-5 and is reported in the language section, where the Core Language Score fell within the average range. Attention was good across both sessions.

Standard scores: On Sounds-in-Words, D.R. obtained a standard score of 72 (95 percent confidence interval 66 to 78; 3rd percentile) compared with same-age peers in the male norm group, well below the average range. On Sounds-in-Sentences, scored on its own scale, the standard score was 70 (63 to 77; 2nd percentile); the two raw scores are not compared because Pearson states they are not comparable across tasks. The KLPA-3 standard score, derived from the same Sounds-in-Words responses, was 67 (60 to 74; 1st percentile). The test-age equivalent is not reported; standard scores, intervals, and percentiles carry the comparison. The sentence-task intelligibility rating (the examiner's rating of each imitated sentence, summarized as a percentage) was 40 percent and is reported here as descriptive information, not as a normed score.

Error inventory and pattern analysis: Articulation: the velar stops /k/ and /g/ were replaced by alveolar stops in initial, medial, and final positions on every opportunity; clusters beginning with /s/ and clusters ending in /r/ or /l/ were reduced to a single consonant in most opportunities; and /s/ and /z/ were produced with lateral airflow in all positions, a distortion rather than a substitution. Singleton /r/ and /l/, the non-sibilant fricatives, and all vowels were produced correctly, so the phonetic inventory is otherwise complete for age. Phonology (KLPA-3): velar fronting occurred in roughly four of five opportunities and cluster simplification in about two thirds, both core processes; deletion of final consonants was occasional in single words but frequent in conversation; no supplemental (atypical) process was observed. Interpretation: the fronting and cluster patterns are rule-governed simplifications that are expected to have resolved well before this age, while the lateralized sibilants are a separate, motor-based articulation error that is not a developmental pattern at any age.

Stimulability, intelligibility, and connected speech: Stimulability: /k/ was produced accurately in syllables and words with a model and a light tactile cue, /g/ in syllables only, and the reduced clusters when both consonants were modeled slowly; the lateral /s/ was not stimulable at any level on this date. Intelligibility: on the Intelligibility in Context Scale his mother's ratings averaged 2.7 of 5, describing speech usually understood by parents, sometimes by his teacher, and rarely by unfamiliar adults; the teacher estimated that she understood about half of his classroom speech. The conversational sample yielded a Percentage of Consonants Correct of 61 percent, which the Shriberg and Kwiatkowski convention describes as moderate to severe; single-word accuracy was clearly higher than conversational accuracy, and final consonants present in single words were often omitted in connected speech. Developmental comparison: Crowe and McLeod (2020) place plosives at the 90 percent criterion by 3;11, so the velar substitutions are more than two years past expectation; the cluster and final-consonant patterns are likewise well past the ages at which they typically resolve; the lateral sibilant distortion is not developmental.

Eligibility analysis: Our state's articulation criterion requires reduced intelligibility or an inability to use the speech mechanism that significantly interferes with communication and attracts adverse attention, speech sound production below developmental expectation for chronological age, and an adverse effect on educational performance; it names no standard-score cutoff, and the 1.5 SD language rule in the same regulation does not apply to articulation. Element one is met by the conversational PCC, the parent and teacher ratings, and the peer reactions recorded by the teacher; element two by the pattern analysis against the 2020 acquisition data, corroborated by the norm-referenced scores; element three by the teacher's report of 09/15/2026 that classmates ask him to repeat, that he shortens his answers, and that he avoids reading aloud. Dialect and language background were considered and do not account for the errors. No single measure has been used as the sole criterion (34 CFR 300.304(b)(2)); eligibility under speech or language impairment is the team's determination.

Recommendations and progress plan: Recommended: direct speech intervention using a pattern-based phonological approach for velar fronting and cluster simplification, starting with /k/ because it is stimulable, with final-consonant production targeted in connected speech; the lateral sibilants addressed as a separate motor-based articulation target once the velar pattern is established; classroom supports of a repair routine (repeat, rephrase, show), teacher checks for understanding before independent work, and no cold-call reading aloud until intelligibility improves. Progress will be measured by percent of occurrence of each pattern on the same monthly probe set, a Percentage of Consonants Correct from a repeated conversational sample each quarter, and the Intelligibility in Context Scale at reevaluation; the GFTA-3 will not be re-administered for progress. Reviewed with D.R.'s parents on 09/21/2026 in plain language: he knows and uses almost all of the sounds of English, a few sounds follow a simplifying pattern that most children outgrow earlier, and one sound is made with the air going out the sides; both are teachable, and he can already make the first target with help.

This sample is fictional and for educational purposes. It does not describe a real child or record; the scores, dates, and details are invented to show write-up structure and are not clinical guidance. Scores are invented for illustration and correspond to no real child or record, and no norm-table values are reproduced.

↑ Back to the template and downloads

Why this sample works

  • The sections administered, the platform, the norm table, and the language of testing are stated before any score, so a reader knows which tasks were given and which normative group produced each percentile.
  • Articulation and phonology sit in separate sections built from one set of responses: the error inventory names sounds, positions, and error types without target words, and the KLPA-3 patterns carry the phonological interpretation and the treatment approach.
  • Intelligibility is measured by named methods (a parent-report scale, a percentage of consonants correct from a sized conversational sample, and teacher report), the record-form rating is reported as descriptive, and single-word accuracy is compared with connected speech.
  • Each error is compared with a named 2020 acquisition source, and the eligibility statement quotes the criterion, walks through its elements, documents adverse effect, and leaves the determination to the team, as 34 CFR 300.304 requires.
  • Recommendations follow the mechanism (pattern-based intervention for the processes, motor-based work for the distortion), targets are ordered by stimulability, and progress is measured by repeating the same probes and sample rather than re-administering the norm-referenced test.

Writing these after every session? BastionGPT drafts complete notes from bullets, dictation, or a transcript.

Generate a note from bullets

Documentation and compliance considerations

United States: the write-up sits inside three decision systems, and the report should say which one each sentence serves. Under IDEA, speech or language impairment is a communication disorder, expressly including impaired articulation, that adversely affects educational performance (34 CFR 300.8(c)(11)); the evaluation rules require a variety of tools and strategies, prohibit any single measure or assessment as the sole criterion, require instruments that are not racially or culturally discriminatory and are administered in the child's native language or mode of communication, and require assessment in all areas of suspected disability (34 CFR 300.304(b) and (c)); no federal GFTA-3 cutoff exists (LAW). States supply the operational rule, and for articulation the rules are mostly functional rather than numeric: California defines an articulation disorder by reduced intelligibility or an inability to use the speech mechanism that significantly interferes with communication and attracts adverse attention, with sound production below developmental expectation and adverse educational effect, and it reserves its 1.5 SD or 7th percentile rule for the language-disorder subsection (5 CCR 3030(b)(11)(A) and (D)); Wisconsin requires, after consideration of age, culture, language background, and dialect, delayed production documented in a natural environment and on a criterion-referenced or norm-referenced measure, intelligibility below the expected range that is not due to home language or dialect and is rated across environments, and production less than 30 percent stimulable for incorrect sounds (PI 11.36(5)(b)1) (LAW, state by state). Texas has no state number either: the Texas Speech-Language-Hearing Association's 2020 articulation guidelines, which carry no regulatory authority, build the determination on a communication-disorder finding plus adverse effect, draw their acquisition table from McLeod and Crowe (2018), use a 100-consecutive-word intelligibility procedure, a 50 to 100 utterance connected-speech sample, and stimulability, and warn that the skewed distribution of articulation errors across ages should not be used to determine eligibility (CONVENTION). The familiar threshold of 85 (1 SD below the mean) is a descriptive convention; where a publisher reports classification statistics at a cutpoint, that statistic describes the test against a clinical reference group, not any agency's definition of disability, so "below 85 qualifies" is a local CONVENTION wherever no rule adopts it. In clinics, coverage turns on documented history, objective findings, functional limitation, and medical necessity under the payer's policy, billed through the untimed evaluation codes 92522 or 92523; a standard score supports the case and never makes it (PAYER POLICY).

Canada and Australia change the norm question and the frame. The English GFTA-3 carries US norms only: Pearson's Canadian and Australian catalogs sell the same edition, and no Canadian or Australian normative study exists, so a report written in either country names the American-English normative frame instead of writing "same-age peers" as if they were local. Canadian service is provincial and program-based: Ontario's Preschool Speech and Language Program, funded by the Ministry of Children, Community and Social Services, serves children from birth until they start school on self-referral and developmental concern with no test-score gate, school-aged children are served through school boards, and other provinces run their own pathways, so any local threshold is labeled a board or agency CONVENTION unless a provincial policy source is cited (PROGRAM POLICY). Australia's NDIS early childhood approach admits young children with developmental delay without a diagnosis and judges evidence on functional impact across communication, social interaction, learning, and daily activities (PAYER POLICY); state school-support criteria target broader functional disability rather than articulation alone; and the Australian-relevant alternative is the DEAP, which carries UK norms and was check-normed on 144 Queensland children aged 5:0 to 6:0, with Australian results reported separately in its manual. Edition and rights: the GFTA-3 (2015) is the current edition on Pearson's US, Canadian, and Australian sites as of September 2026, with no fourth edition announced; the GFTA-3 Spanish (2017) is an adaptation with its own norms based on Spanish speakers living in the US and Puerto Rico, not a translation, and the KLPA-3 companion is English only. The publisher classifies the GFTA-3 at qualification level B; record forms and stimulus books are protected consumables, transferring results into an electronic record and placing scoring on another platform are permission and license matters with Pearson, and a free public web scorer of GFTA-3 conversions has no authorization. A report may always contain the child's derived scores and your interpretation in your own words.

GFTA-3 is a trademark of NCS Pearson, Inc. (Pearson's own materials mark GFTA and KLPA as trademarks of Pearson Education and its affiliates, and the 2015 test content is copyrighted by NCS Pearson, Inc.). BastionGPT is not affiliated with, or endorsed by, the publisher. This page reproduces no test items, stimuli, norms, or scoring materials.

↑ Back to the template and downloads

Common GFTA-3 write-up errors reviewers flag

The numbers behind these errors are specific. 34 CFR 300.304(b)(2) prohibits any single measure from deciding; Pearson's own FAQ states that dialectal variations are not counted as errors (and retains the correction of a webinar slide that said otherwise) and that Sounds-in-Words and Sounds-in-Sentences raw scores are not comparable; Crowe and McLeod's 2020 review of 18,907 US children places plosives, nasals, and glides at the 90 percent criterion by 3;11, affricates by 4;11, liquids by 5;11, and fricatives by 6;11; Morrison and Shriberg (1992) showed that articulation-test and conversational-sample results diverge in children with speech delay; Kirk and Vigeland (2014) found that norm-referenced phonological pattern tests failed many of the psychometric criteria expected of them; and Wisconsin's rule keys eligibility to intelligibility and a 30 percent stimulability criterion while California's articulation subsection names no number at all. The BastionGPT Clinical Advisory Board sees the same errors most often in GFTA-3 documentation reviews:

  • The 85 cutoff written as the rule. "A standard score below 85 indicates an articulation disorder and qualifies the student." No federal rule says so, California's articulation criterion is intelligibility, developmental expectation, and adverse effect, Wisconsin's adds a stimulability threshold, and a publisher's classification cutpoint describes the test, not the law. Write the score, its interval, and the named rule, and let the team apply it.
  • The articulation score standing in for intelligibility. "The GFTA-3 indicates 55 percent intelligibility." Sounds-in-Words counts consonant errors in single words, and the record-form rating is the examiner's judgment of imitated sentences. Intelligibility needs its own named method (the Intelligibility in Context Scale, a Percentage of Consonants Correct from a sized conversational sample, or a listener-based words-understood percentage), and an average single-word score with poor spontaneous intelligibility is a finding, not a dismissal.
  • A phonological diagnosis with no pattern analysis. A low GFTA-3 score followed by "phonological disorder" with no KLPA-3 or equivalent process analysis, or the KLPA-3 written up as a second test. The GFTA-3 says how many consonants were wrong; the KLPA-3 re-analyzes the same responses to say whether the errors are organized into developmental or atypical patterns, which is the construct the diagnosis and the treatment approach depend on.
  • Dialect and transfer features scored as errors. A community-dialect or Spanish-influenced production counted as a substitution, or English-edition standard scores reported as diagnostic for a sequential bilingual child. Pearson scores recognized dialectal variations as correct, the English norm sample's bilingual members were simultaneous learners who used English most, ASHA's position is that standard scores are not reported for a child the norm group does not represent, and the federal rules require nondiscriminatory assessment in the native language.
  • The two raw scores compared, or the sentence task scored out of range. "He made 12 more errors in sentences than in words, showing a connected-speech effect." Pearson states the raw scores are not comparable because the tasks offer different production opportunities, and Sounds-in-Sentences yields no standard score below age 4. Compare the two standard scores' standing and the clinical pattern, and get connected-speech evidence from a spontaneous sample.
  • The norm group left unstated. A percentile with no mention of the male or female table that produced it, or a nonbinary or transgender student scored silently against one table. The instrument offers only binary norms and the publisher has issued no guidance, so the report names the table, gives the reason, and rests the conclusion on criterion-referenced evidence rather than the choice.
  • An age equivalent or an old chart doing decisional work. "Articulation age of 3 years 6 months" as the headline, or an error excused or flagged by a decades-old acquisition chart that places /r/ years later than the 2020 review's 5;11. Pearson's own guidance is that age equivalents should not be used for diagnostic or placement decisions; standard scores with intervals carry the peer comparison, and each error is compared with a named contemporary source.
How BastionGPT helps

BastionGPT is specifically trained, tuned, and clinically tested on speech-language evaluation reports.

  • Give it the facts (sections administered and platform, age and norm table, standard scores with intervals and percentiles, the error inventory by sound, KLPA-3 patterns with percent of occurrence, stimulability, intelligibility measures and the connected-speech sample, language history, the governing rule) and it drafts the results section: articulation and phonology in separate sections, intelligibility by a named method, each error compared with a named acquisition source, and the eligibility statement written against the rule, ready for your review.
  • Cross-check a finished report for the gaps reviewers flag: a norm group left unstated, an intelligibility claim resting on the record-form rating, a phonological diagnosis with no pattern analysis, a dialect feature counted as an error, two raw scores compared, or an eligibility sentence built on a cutoff.
  • Draft the companion paragraphs: the plain-language family summary, the bilingual or dialect validity statement, the nonbinary norm-group note, or the progress paragraph that repeats the same probes and conversational sample instead of re-administering the test.

See how clinicians use it day to day on the AI therapy notes page.

Many BastionGPT users report saving more than 90 minutes per day on documentation.

HIPAA-compliant with a signed BAA on every plan. Your data is never used to train models. BastionGPT drafts, you review and sign.

Frequently asked questions

The Sounds-in-Words raw score is the number of consonant and cluster errors in single words, converted through age- and sex-specific norms into a standard score (mean 100, SD 15) with a confidence interval, a percentile rank, a test-age equivalent, and a Growth Scale Value; the comparison is with same-age peers in the male or female norm table used, and the report says which. Scores within 1 SD of the mean (85 to 115) are conventionally described as average and lower scores as below average to very low, in prose rather than a lookup table. Sounds-in-Sentences, given from age 4, has its own standard score on its own scale, and Pearson states the two raw scores are not comparable. Stimulability, the vowel and R analyses, and the sentence-task intelligibility rating are descriptive, not normed. Age equivalents show no rank among peers and, in Pearson's own guidance, should not be used for diagnostic or placement decisions. As of September 2026 the 2015 edition is current, with no fourth edition announced.

No score does by itself, and three voices need separating. Federal law sets no cutoff, prohibits any single measure as the sole criterion (34 CFR 300.304(b)(2)), and requires adverse educational effect for the speech or language impairment category (LAW). State articulation rules are mostly functional: California requires reduced intelligibility or a speech-mechanism limitation that significantly interferes with communication, production below developmental expectation, and adverse effect, with no number; Wisconsin requires documented delay, intelligibility below the expected range not due to home language or dialect, and less than 30 percent stimulability for incorrect sounds; Texas association guidance builds on developmental norms, an intelligibility procedure, stimulability, and adverse effect (LAW and guidance, state by state). "Below 85" is a CONVENTION unless a rule adopts it, and where a publisher reports classification statistics at a cutpoint, that describes the test, not the law. So the report quotes the rule, compares each element with the evidence, and leaves the determination to the team.

Not in the norm-referenced sense. The Sounds-in-Words score counts consonant errors in single words. The record form does carry an intelligibility rating, but it is the examiner's rating of each imitated sentence in the Sounds-in-Sentences task, summarized as a percentage: a clinician judgment of structured, repeated speech, not a normed score and not spontaneous conversation. A defensible intelligibility statement therefore names its method: the Intelligibility in Context Scale (a seven-item parent report across communication partners; McLeod, Harrison, and McCormack, 2012), a Percentage of Consonants Correct from a connected-speech sample of stated size (Shriberg and Kwiatkowski's 1982 convention reads above 85 percent as mild and below 50 percent as severe), or a listener-based percentage of words understood, as Texas guidance does with a 100-consecutive-word sample. An average single-word score with poor spontaneous intelligibility is a finding to investigate, not a reason to close the case.

They answer different questions from one set of responses. The GFTA-3 asks which consonants were produced accurately in the standardized single-word and sentence tasks; the KLPA-3 (Khan and Lewis, 2015) re-analyzes the same Sounds-in-Words responses to ask whether the errors are organized into phonological processes, scoring core developmental patterns and supplemental atypical ones by percent of occurrence, with its own standard score and percentile by age and sex. No second administration is involved, and the KLPA-3 is English only. The articulation section therefore reports sounds, positions, error types, consistency, and stimulability; the phonology section reports patterns and whether they are developmental for the child's age. Neither leads eligibility: the rule leads, and the report draws on whichever evidence answers it, with the treatment approach following the mechanism (motor-based work for an isolated distortion, pattern-based work for rule-governed processes).

Pearson provides separate male and female norm tables and, as of September 2026, no published guidance for a student who fits neither. The defensible write-up states that the instrument offers only binary norms, names the table used and why, reports the result as a comparison with that normative category rather than as a universal standing, notes that no eligibility conclusion rests on the choice, and leans on criterion-referenced evidence: the error inventory against a named acquisition source, KLPA-3 patterns, stimulability, and intelligibility in connected speech. If local policy permits scoring against both tables, label that as a sensitivity check rather than a publisher-approved procedure, and do not equate a student's gender identity with the test's norm variable.

Start with the language history: which languages, since when, which is dominant, which dialect, and whether the child is a simultaneous or sequential learner. Pearson's scoring counts recognized dialectal variations as correct (its FAQ retains the correction of a 2016 webinar slide that said the opposite), and the English norm sample's bilingual subgroup (13 percent) was limited to simultaneous learners who used English most, so the English standard score is not diagnostic for a sequential bilingual or English-learning child; ASHA's position is that standard scores cannot be reported for a child the norm group does not represent, and federal rules require nondiscriminatory assessment in the native language (34 CFR 300.304(c)(1)). Sample speech in both languages, treat cross-linguistic transfer as difference rather than error, consider the GFTA-3 Spanish (2017; its own norms from Spanish speakers in the US and Puerto Rico; not a translation), and keep whatever remains atypical in both languages as the clinical finding.

No universal rule requires it, and the report should state which sections were given rather than implying a complete kit. Sounds-in-Sentences is a structured sentence-imitation task (the examiner reads a short story, then the child repeats each sentence) normed on its own scale from age 4; for younger children it yields no standard score. It is useful when the referral concerns performance beyond single words, but it is not a substitute for a spontaneous connected-speech sample. Pearson states that the two raw scores are not comparable because the tasks offer different production opportunities; compare the two standard scores' standing and the clinical pattern, not the raw-error difference. If Sounds-in-Words is average and everyone says the child is hard to understand, the sentence task and a conversational sample are where the answer lives.

No, and the report does not need them. Record forms, stimulus books, and norms tables are the publisher's protected content at qualification level B, and a report that prints the word list leaks content while explaining nothing that a pattern description does not. Write the error inventory by sound, position, error type, and consistency (for example, velar stops fronted in all positions, clusters reduced, sibilants lateralized) and the KLPA-3 patterns by name and frequency. Pearson's permissions process covers transfer of results into an electronic record and placement of scoring on another platform, and unauthorized public scoring tools have no standing. Test security is also credibility: a team that sees stimuli in a report has reason to doubt the scores that follow.

Yes. Give it the facts (sections administered and platform, age and norm table, standard scores with intervals and percentiles, the error inventory, KLPA-3 patterns with percent of occurrence, stimulability, intelligibility measures and the connected-speech sample, language history, and the governing rule) and it drafts the results section: articulation and phonology in their own sections, intelligibility by a named method, each error compared with a named acquisition source, and the eligibility statement written against the rule, ready for your review. It can also cross-check a finished report for a norm group left unstated, an intelligibility claim resting on the record-form rating, a phonological diagnosis with no pattern analysis, a dialect feature counted as an error, or an eligibility sentence built on a cutoff. BastionGPT is HIPAA-compliant with a signed BAA on every plan, and your data is never used to train models.

Primary sources

The instrument facts and compliance claims on this page trace to these sources, last verified September 2026:

  1. Publisher record, accessed September 2026: NCS Pearson, GFTA-3 product page (2015; ages 2:0 to 21:11; 5 to 15 minutes for Sounds-in-Words; qualification level B; separate male and female norms; FAQ on dialect scoring, raw-score comparability, and the bilingual norm subgroup), GFTA-3 and GFTA-3 Spanish brochure (sections, sentence-imitation task, intelligibility rating, stimulability, vowel and R analyses, picture sets, KLPA-3 core and supplemental processes), KLPA-3 product page (2015; uses GFTA-3 responses; English only), GFTA-3 Spanish product page (2017; not a translation; norms from Spanish speakers in the US and Puerto Rico), interpretation problems of age and grade equivalents, and Permissions and Licensing.
  2. Federal law: 34 CFR 300.8(c)(11) (speech or language impairment, including impaired articulation, that adversely affects educational performance) and 34 CFR 300.304 (variety of tools; no single measure as sole criterion; nondiscriminatory assessment in the native language; all areas of suspected disability), via Cornell LII.
  3. State rules and guidance: California, 5 CCR 3030(b)(11)(A) and (D); Wisconsin, PI 11.36(5)(b)1 (observation, criterion- or norm-referenced delay, intelligibility across environments not due to home language or dialect, less than 30 percent stimulable); Texas Speech-Language-Hearing Association, SI Disability Determination Guidelines for Articulation Disorders, revised 2020 (McLeod and Crowe acquisition table; 100-consecutive-word intelligibility procedure; 50 to 100 utterance sample; stimulability; skewed error distributions not used for eligibility; no regulatory authority).
  4. Acquisition evidence: Crowe K and McLeod S, 2020, American Journal of Speech-Language Pathology 29(4), 2155 to 2169, Children's English consonant acquisition in the United States: a review (15 studies; 18,907 children; 90 percent criterion by class); McLeod S and Crowe K, 2018, American Journal of Speech-Language Pathology 27(4), 1546 to 1571, Children's consonant acquisition in 27 languages: a cross-linguistic review (64 studies; 26,007 children; most consonants acquired by 5;0).
  5. Single-word versus connected speech and intelligibility: Morrison JA and Shriberg LD, 1992, Journal of Speech and Hearing Research 35(2), 259 to 273, Articulation testing versus conversational speech sampling; Shriberg LD and Kwiatkowski J, 1982, Journal of Speech and Hearing Disorders 47(3), 256 to 270, Phonological disorders III: a procedure for assessing severity of involvement (Percentage of Consonants Correct); McLeod S, Harrison LJ, and McCormack J, 2012, Journal of Speech, Language, and Hearing Research 55(2), 648 to 656, The Intelligibility in Context Scale: validity and reliability of a subjective rating measure (seven items; five-point scale; alpha .93).
  6. Psychometric critique and practice survey: Kirk C and Vigeland L, 2014, Language, Speech, and Hearing Services in Schools 45(4), 365 to 377, A psychometric review of norm-referenced tests used to assess phonological error patterns, and 2015, 46(1), 14 to 29, Content coverage of single-word tests used to assess common phonological error patterns; Skahan SM, Watson M, and Lof GL, 2007, American Journal of Speech-Language Pathology 16(3), 246 to 259, Speech-language pathologists' assessment practices for children with suspected speech sound disorders (333 respondents; commercial single-word tests, intelligibility estimates, and stimulability among the most used tasks); Buros Center for Testing, GFTA-3 listing (reviewed in the Twentieth Mental Measurements Yearbook).
  7. Professional guidance: ASHA Practice Portal, Speech Sound Disorders: Articulation and Phonology (single-word and connected-speech sampling; stimulability; intelligibility; orofacial examination; error-pattern analysis; dialect and language difference versus disorder; standard scores not reported for an unrepresented child).
  8. Canada and Australia: Ontario, Preschool Speech and Language Program (birth until school entry; self-referral); NDIS, supporting-evidence guidance for health professionals (functional impact across communication, social interaction, learning, and daily activities); Pearson Australia, DEAP product page (UK norms; check-normed on 144 Queensland children aged 5:0 to 6:0).

Educational content, not legal or billing advice. Sample notes are fictional. Follow your organization's policies and your board, payer, and jurisdiction requirements.