The Bayley-4 (Bayley Scales of Infant and Toddler Development, Fourth Edition) is an individually administered developmental assessment for children 16 days to 42 months old, published by Pearson in 2019. Psychologists, therapists, and early-intervention teams use it to measure cognitive, language, and motor development and to inform eligibility decisions. This page covers how to write up Bayley-4 results, with a fictional sample and a results-section template.
Psychologists, physical and occupational therapists, and speech-language pathologists; publisher qualification level B
Early-intervention and IFSP teams, NICU high-risk follow-up clinics, pediatricians, courts and payers, parents
400 to 800 words for the results section · administration 30 to 70 minutes by age and cooperation
Norm-referenced infant and toddler developmental battery
Early-intervention (IDEA Part C) eligibility, NICU high-risk follow-up, developmental-delay and autism evaluations, re-evaluations
Published by Pearson (2019, current edition; norms revised 2023); described here for write-up purposes, no test content reproduced
The Bayley-4 (Bayley Scales of Infant and Toddler Development, Fourth Edition) is a norm-referenced developmental assessment for children 16 days to 42 months old, published by NCS Pearson in September 2019 as the successor to the Bayley-III (2006), with Nancy Bayley and Glen P. Aylward as authors of record. It measures development across five scales. Three are administered directly to the child through structured play interaction: Cognitive; Language, with Receptive Communication and Expressive Communication subtests; and Motor, with Fine Motor and Gross Motor subtests. Two are collected from a primary caregiver through a questionnaire: Social-Emotional and Adaptive Behavior. Directly administered subtests yield scaled scores (mean 10, standard deviation 3), and each scale yields a composite standard score (mean 100, standard deviation 15) with a percentile rank, a 95% confidence interval, and a descriptive classification; the reports also produce developmental age equivalents and growth scale values. It is often called simply the Bayley. Administration is paper or fully digital on Q-global, and completion time runs about 30 to 70 minutes depending on the child's age and cooperation. As of July 2026 the Bayley-4 is the current edition, and no Bayley-5 has been announced.
Two facts separate a defensible Bayley-4 write-up from the sample reports that rank in search. First, the norm version matters. In 2023 Pearson revised the Cognitive, Language, and Motor norms after its own postpublication research identified "70 cases...with performance concerns" that, in Pearson's words, "significantly impacted the normative information in the average to extremely high" score range; those cases were replaced, not deleted, with demographically matched cases from the Bayley-4 UK validation study, and Q-global scoring was updated to match. Any Cognitive, Language, or Motor score generated before that update derives from the original 2019 norms, so a report relied on for eligibility or longitudinal comparison should record which norm version produced it (the fastest check is the manual's copyright and normative-copyright dates). Second, the caregiver-report boundary is easy to blur: the Social-Emotional and Adaptive Behavior scales are questionnaire-based, not directly observed, and the write-up should label them that way and corroborate them against observation and history. The Bayley-4 also sits inside a family of measures a report frequently names together, the Vineland-3 or ABAS-3 for a fuller adaptive picture and the WPPSI-IV once a child can sustain a more cognitively demanding battery.
Psychologists, physical therapists, occupational therapists, and speech-language pathologists write up Bayley-4 results inside developmental evaluations, most often for early-intervention eligibility under IDEA Part C, where services run from birth to age 3 and physical therapists commonly administer the Motor scale. NICU high-risk follow-up is the other core setting: the Bayley is the standard developmental measure there, and many programs assess at fixed corrected ages such as 6, 12, 18, 24, and 36 months rather than on a fixed calendar interval. It also appears as a developmental-level adjunct in autism evaluations, sitting next to observation and interview measures such as the ADI-R and CARS-2 rather than diagnosing on its own, and as a neurodevelopmental endpoint in clinical research. Age drives instrument choice at the upper edge: the Bayley-4 stops at 42 months, so between roughly 30 and 42 months a team that needs an IQ estimate or a broader personal-social profile may move to the WPPSI-IV (from 2:6) or the BDI-3 (birth to 7:11). A Bayley-4 section rarely travels alone; it anchors the developmental portion of a broader evaluation next to adaptive data from the Vineland-3 or ABAS-3, caregiver history, and the behavioral observations that infant and toddler validity rests on.
No statute, payer, or publisher mandates a results-section format. The sequence below is the convention experienced infant and toddler evaluators converge on because it survives review: it states chronological and corrected age and the norm version before any score, gives behavioral observation the weight infant validity actually rests on, leads with standard scores over age equivalents, labels which data came from the caregiver, and phrases delay in the metric the eligibility rule uses. Each section carries the pitfall that most often undermines it.
Measures, edition, norm version, and ages. Name the instrument and edition, the norm version (original 2019 or the 2023 revision for the Cognitive, Language, and Motor scales), the administration format (paper or Q-global), the scales given, and both the child's chronological age and, for a preterm child, the corrected age used to derive scores. Pitfall: a silent norm version and a single age. A Cognitive, Language, or Motor score means different things under the 2019 and 2023 norms, and a score table that does not say whether corrected or chronological age was used is uninterpretable at re-evaluation.
Behavioral observations and validity. Document state and arousal, attention, cooperation, the role of caregiver participation, and any factor such as fatigue, hunger, or illness that could depress performance, then close with a validity statement tied to what you observed. Infant results are only as good as the state the child was in. Pitfall: a boilerplate validity sentence. "Results are considered valid" with no observational basis reads as template text, and with infants the observed state is the validity evidence.
Domain-by-domain composite results. Report the Cognitive, Language, and Motor composites, each with its standard score, 95% confidence interval, percentile rank, and one named descriptor set applied consistently, and note receptive-versus-expressive or fine-versus-gross patterns in prose. Pitfall: leading with age equivalents or a developmental quotient. Standard scores, intervals, and percentiles carry the interpretation; an age equivalent is not a standard score and must not stand in for one.
Caregiver-report domains. Report the Social-Emotional and Adaptive Behavior composites and label them explicitly as caregiver questionnaire data, corroborated against observation and history, with any observed-versus-reported discrepancy noted. Pitfall: presenting questionnaire scores as directly observed. These two scales come from a caregiver rating form, and a report that blurs the source overstates what was directly measured.
Age equivalents and growth scale values. If age equivalents or growth scale values appear, present them only with their published cautions and only for their intended purpose. Pitfall: comparing values that do not compare. Age equivalents are not an equal-interval scale and, in Pearson's words, growth scale values "cannot be meaningfully compared between subtests or subdomains"; they track one child over time, they do not rank a child against peers.
Corrected age and delay severity. State the corrected-age rule you applied and translate the composite into the exact metric the jurisdiction uses, standard-deviation cuts for SD-rule programs and percent-delay language for percent-delay programs. Pitfall: a single unstated correction rule. The 24-month default and the domain-differential alternative give different scores between 24 and 36 months, and a delay described in the wrong metric can qualify or disqualify a child by accident.
Interpretive summary and recommendations linkage. Answer the referral question, state what the scores do and do not establish, place the eligibility decision with the team, and tie each recommendation to a specific finding and a re-evaluation interval. Pitfall: a low score read as a fixed verdict. An early Bayley score is a current estimate, not a diagnosis of cognitive impairment, and recommendations unmoored from findings tell the reader the data did not drive them.
MEASURES, EDITION, NORM VERSION, AND AGES Instrument/edition: Bayley-4 (2019) Norms: [2019 / 2023 revision] Format: [paper / Q-global] Scales given: [Cog, Lang, Motor, SE, Adaptive] Chronological age: ____ Corrected age (if preterm): ____ Basis: [corrected / chronological] BEHAVIORAL OBSERVATIONS AND VALIDITY [State/arousal, attention, cooperation, caregiver participation, fatigue/hunger/illness; validity statement tied to observed state] DOMAIN-BY-DOMAIN COMPOSITE RESULTS (directly administered) [Cognitive / Language / Motor: standard score, 95% CI, percentile, one named descriptor set; receptive vs expressive, fine vs gross in prose] CAREGIVER-REPORT DOMAINS [Social-Emotional and Adaptive Behavior: labeled as caregiver questionnaire; corroborated vs observation; discrepancies noted] AGE EQUIVALENTS AND GROWTH SCALE VALUES (if reported) [With published cautions; not compared across subtests; not a substitute for standard scores] CORRECTED AGE AND DELAY SEVERITY [Correction rule applied (24-month default / domain-differential); delay stated in the jurisdiction's metric: SD cut or percent delay] INTERPRETIVE SUMMARY (answer the referral question) [What converges; what the scores do NOT establish; eligibility is the team's decision under jurisdiction criteria] RECOMMENDATIONS LINKAGE [Each recommendation tied to a finding; re-evaluation interval] Evaluator signature / credentials: Date:
Free to use and share, no signup. The PDF includes a one-page cheat sheet with section-by-section pitfalls and a pre-sign checklist; the DOCX is the blank results-section skeleton, ready to adapt.
Scenario: a 24-month-old born at 27 weeks, seen in a NICU high-risk follow-up clinic. Because the child was about 3 months premature, all scores use corrected age (roughly 21 months), stated prominently, the detail most sample reports bury. An internally consistent pattern shows Cognitive in the low-average range with Language and Motor lower. This is the developmental results section only, condensed but structurally complete. All details are fictional.
Child: M.R., chronological age 24 months (born at 27 weeks; corrected age 21 months) · Referral: NICU high-risk infant follow-up, developmental surveillance · Evaluator: J. Okafor, PT, DPT · Testing date: 07/09/2026 · Report date: 07/15/2026
Measures and administration: The Bayley Scales of Infant and Toddler Development, Fourth Edition (Bayley-4) was administered in paper format using United States norms (2023 revision for the Cognitive, Language, and Motor scales). The Cognitive, Language, and Motor scales were administered directly to the child through structured play; the Social-Emotional and Adaptive Behavior scales were completed by the child's mother on the caregiver questionnaire. All scores were derived using corrected age for prematurity per the clinic's convention, and both chronological and corrected age are reported. Composites are reported with 95% confidence intervals (CI), percentile ranks, and the descriptor set printed in the publisher's score reports, named here so the labels are interpretable.
Behavioral observations and validity: M.R. arrived rested and fed, separated easily onto his mother's lap, and warmed to the examiner within a few minutes. He engaged well with the play materials, reached and manipulated readily with either hand, and vocalized in single words and jargon. Attention faded near the end of the session; one break restored engagement, and standardized procedures were maintained throughout. Because an infant's observed state is what the validity of these results rests on, the results below are considered a valid estimate of current developmental functioning at his corrected age.
| Scale | Standard score | 95% CI | Percentile | Classification | Basis |
|---|---|---|---|---|---|
| Cognitive (COG) | 85 | 79-92 | 16 | Low Average | Directly administered |
| Language (LANG) | 74 | 68-82 | 4 | Below Average | Directly administered |
| Motor (MOT) | 76 | 70-84 | 5 | Below Average | Directly administered |
| Adaptive Behavior (ADAP) | 78 | 72-85 | 7 | Below Average | Caregiver report |
| Social-Emotional (SE) | 92 | 84-100 | 30 | Average | Caregiver report |
Domain results: M.R.'s Cognitive composite (85, Low Average, 16th percentile) places his play, problem-solving, and object exploration at the lower edge of the average range for his corrected age. His Language composite (74, Below Average, 4th percentile) is his lowest directly administered score, and within it receptive communication was stronger than expressive, consistent with his single-word output during testing. His Motor composite (76, Below Average, 5th percentile) reflects gross-motor skills weaker than fine-motor, a pattern common in children born very preterm. Standard scores, intervals, and percentiles carry these statements; the age-equivalent and growth-scale values from the score report are held for progress tracking rather than for ranking him against peers.
Caregiver-report domains: The Adaptive Behavior composite (78, Below Average, 7th percentile) and the Social-Emotional composite (92, Average, 30th percentile) come from his mother's ratings on the caregiver questionnaire, not from direct observation, and are labeled that way here. The below-average adaptive score is consistent with the directly observed language and motor findings; the average social-emotional score matches the warm, well-regulated engagement seen in the session. No large observed-versus-reported discrepancy was noted.
Corrected age and the reporting decision: Every score above uses corrected age, stated so a later evaluator can compare like with like. One caveat is documented rather than hidden: under the domain-differential rule Aylward (2020) proposes, Language and Motor composites would be corrected to 36 months at all degrees of prematurity, not stopped at 24, which would raise those two scores somewhat relative to a chronological-age reference. The clinic's convention is the 24-month default aligned with the American Academy of Pediatrics, so that is what the scores reflect; the alternative is named for the reader rather than applied silently. Delay severity is described here so the early-intervention team can map it onto this state's eligibility metric, which the results section does not itself decide.
Summary and linkage to recommendations: Testing at 21 months corrected shows low-average cognitive development (COG 85) with below-average language (LANG 74) and motor (MOT 76) composites and concordant below-average adaptive behavior by caregiver report (ADAP 78), alongside age-typical social-emotional functioning (SE 92). The language finding supports recommendation 1 (speech-language evaluation and early-intervention services). The motor finding supports recommendation 2 (physical and occupational therapy referral). The stability caveat supports recommendation 3 (re-evaluation at the next scheduled corrected-age visit, comparing same-norm scores). The full profile feeds recommendation 4: the early-intervention team should apply this state's developmental-delay criteria to the whole evaluation, not to any single score.
This sample is fictional and for educational purposes. It does not describe a real child or record, and the scores are invented for illustration and correspond to no real child or record.
Writing these after every session? BastionGPT drafts complete notes from bullets, dictation, or a transcript.
Generate a note from bulletsWrite the results section knowing which decision framework will read it, and label each requirement's strength honestly. For infants and toddlers the primary framework is IDEA Part C early intervention, and its structure is law but its threshold is not. Federal regulation defines an eligible child as one "experiencing a developmental delay...in one or more of the following areas" of cognitive, physical, communication, social or emotional, and adaptive development, or with a diagnosed condition carrying a high probability of delay (34 CFR 303.21). The level of delay, though, is state-defined: each state must adopt its own "rigorous definition of developmental delay" and "specify the level of developmental delay...or other comparable criteria" (34 CFR 303.111), so the ECTA Center's compilation documents criteria families as different as 2 standard deviations below the mean in one area, 1.5 standard deviations in two areas, and 25, 30 to 40, or 50 percent delay, and established-condition lists vary just as much (Barger and colleagues counted 620 unique qualifying conditions across 49 states, DC, and 4 territories, with 89.3% listed by fewer than 10 jurisdictions). The practical move: phrase delay severity in the exact metric the child's jurisdiction uses, because the same Bayley-4 profile can qualify a child in one state and not the next. NICU high-risk follow-up is convention and program policy rather than statute (California's CCS High Risk Infant Follow-Up program, for example, reimburses Bayley-or-equivalent developmental assessment to age 3). US testing time bills under the developmental-testing code family 96112 and 96113 as payer policy with plan-specific rules.
Norm currency is the Bayley-4's quiet defensibility question. As of July 2026 the Bayley-4 (2019) is the current edition and no Bayley-5 has been announced, but the 2023 revision of the Cognitive, Language, and Motor norms means a score's meaning depends on which norms produced it: name the norm version, and never compare a pre-revision score with a post-revision score as if they sat on the same metric. Editions are not interchangeable for longitudinal monitoring either. Because the Bayley-4 tracks close to the Bayley-III, in Pearson's technical report "0.1 scaled score points and 0.5 and 1.3 standard score points" lower across the three scales, the older edition's documented tendency to underestimate delay carries into interpretation, so switching editions mid-follow-up is a clinical-judgment call, not a clerical one. Access is the publisher's: qualification level B, with the Cognitive, Language, and Motor scales requiring in-person standardized administration, since per Pearson's own guidance those subtests "cannot be administered in a standardized format via" telepractice, and only the Social-Emotional and Adaptive Behavior questionnaires carry a validated remote path, so a hybrid workflow should document which components were in person. Outside the US, early intervention runs through provincial Infant and Child Development Programs in Canada (generally birth to age 3, with no single national cut score) and, in Australia, through the NDIS early childhood approach for children under 6 with developmental delay or developmental concerns, where "developmental delay" is defined in the NDIS Act 2013; clinicians outside the US should also confirm which normative version their Q-global account and print kit carry.
Bayley, Bayley-4, and Pearson are trademarks, in the US and other countries, of Pearson plc or its affiliates; Bayley-4 materials are copyrighted by NCS Pearson, Inc. BastionGPT is not affiliated with, or endorsed by, the publisher. This page reproduces no test items, stimuli, norms, or scoring materials.
There is no payer audit series for infant developmental write-ups; the accountability literature here is psychometric, and it cuts close to everyday practice. The Bayley-4 is reviewed in the Buros Center's Twenty-Second Mental Measurements Yearbook, and the sharper caution comes from the edition-comparison literature: an influential study of the Bayley-III found control-group composites running "between 0.55 and 1.23 SD above the normative mean" and concluded that proportions of children with delay "were grossly underestimated" (Anderson and colleagues, 2010), a concern that carries into the Bayley-4 because the two editions track closely. The mechanics of administration are the publisher's problem; the errors below are write-up errors, and they are yours. The BastionGPT Clinical Advisory Board sees the same ones most often in Bayley-4 report reviews:
BastionGPT is specifically trained, tuned, and clinically tested on psychological and psychoeducational evaluation reports.
See how clinicians use it day to day on the AI therapy notes page.
Many BastionGPT users report saving more than 90 minutes per day on documentation.
HIPAA-compliant with a signed BAA on every plan. Your data is never used to train models. BastionGPT drafts, you review and sign.
Two levels of scores carry the interpretation. At the subtest level the Bayley-4 uses scaled scores with a mean of 10 and a standard deviation of 3; at the scale level the Cognitive, Language, Motor, Social-Emotional, and Adaptive Behavior scales each yield a composite standard score with a mean of 100 and a standard deviation of 15, reported with a percentile rank and a 95% confidence interval. As a rough map, composites near 100 are average, and scores drop into monitoring and then significant-delay territory as they fall one and then two standard deviations below the mean, though the exact eligibility line is set by the jurisdiction, not the test. The reports also print developmental age equivalents and growth scale values, both of which carry published cautions and should never lead. Report the standard score, interval, and percentile together, name the descriptor set you use, and integrate the result into a whole evaluation.
Yes to the first, no to the second. In 2023 Pearson revised the normative, reliability, and validity information for the Cognitive, Language, and Motor scales after its postpublication research found 70 normative cases with performance concerns; those cases were replaced with demographically matched cases from the Bayley-4 UK validation study, and Q-global scoring was updated so it now applies the revised norms automatically. The change concentrated in the average-and-above score range. Any Cognitive, Language, or Motor score generated before the update derives from the original 2019 norms, so a report used for eligibility or longitudinal comparison should record which norm version produced it; the quickest check is the manual's copyright and normative-copyright dates. As of July 2026 the Bayley-4 (2019) is the current edition and no Bayley-5 has been announced.
State the rule you used on every score, because the field is genuinely unsettled. The dominant convention corrects to 24 months of chronological age across all domains and then stops, which aligns with the American Academy of Pediatrics, whose 2023 clinical report frames screening as "adjusted for a child's corrected age (if under 24 months)" (Davis and colleagues), and with Q-global's default. The Bayley-4 co-author Aylward (2020) argues a domain-differential rule instead: correct cognitive composite scores for the first 2 years, but correct "language and motor composite scores...to 3 years at all degrees of prematurity," concluding that not correcting at 3 years and younger places preterm infants at a distinct disadvantage. The practical write-up move is to use the 24-month default, and for language and motor composites in preterm children between 24 and 36 months, report both corrected and uncorrected scores and cite the rationale.
No. IDEA Part C is instrument-neutral: federal regulation requires a developmental delay measured by "appropriate diagnostic instruments and procedures" but names no specific test, and it leaves the level of delay to each state. So the requirement is structure, not a mandated tool: each state sets its own definition and criteria, and states approve lists of acceptable instruments on which the Bayley-4 commonly appears alongside measures such as the BDI-3. The result is real variation, since the same profile can qualify a child in one state and not the next. The honest sentence for a report is that the Bayley-4 is convention and, on many state instrument lists, accepted policy, not a legal mandate, and eligibility is a team determination under the jurisdiction's own delay criteria. Describe severity in that jurisdiction's metric, and let the team apply it.
Age and question drive the choice. The Bayley-4 covers 16 days to 42 months and gives a fine-grained infant and toddler developmental profile, so it carries most early-intervention and NICU follow-up referrals. The WPPSI-IV starts at 2:6, so at the 30-to-42-month overlap the convention is to move to it when a child can sustain a more cognitively demanding battery and an IQ estimate is the question. The BDI-3 spans birth to 7:11 and adds a personal-social domain, so it is often chosen when a child straddles the toddler-to-preschool boundary or when that broader profile is needed. Pearson's own comparison offers an anchor for the transition: at about 36 months a mean Bayley-4 Cognitive scaled score of 9.7 equated to a standard score of 98.5 while the mean WPPSI-IV Full Scale IQ was 103.3. Name the instrument, edition, and norms whichever you pick.
Only in part. Pearson supports remote administration of the Social-Emotional and Adaptive Behavior questionnaires through Q-global, either by a secure email link that does not require teleconferencing or by reading items aloud over video. The Cognitive, Language, and Motor scales are a different matter: per Pearson's own guidance those subtests "cannot be administered in a standardized format via" telepractice, because they require hands-on, in-person structured play with the child. So a valid Bayley-4 cognitive, language, or motor result comes from an in-person session. When a workflow is hybrid, document which components were in person and which questionnaires were completed remotely, and treat any departure from standardized administration as a stated limitation on the scores it touches.
Only with an explicit caution, and never as the headline number. Age equivalents are intuitive to parents, which is exactly why they mislead: they are not an equal-interval or ratio scale, so, in Pearson's own guidance, they "cannot be added, subtracted, or averaged," and their reliability is poorer near the top of the range. Growth scale values carry a separate caution, since Pearson states they "cannot be meaningfully compared between subtests or subdomains" and are meant to track one child's progress over time rather than to rank a child against peers. Both have a legitimate niche, describing function for a child who scores below the standard-score floor, but they belong after the standard scores, intervals, and percentiles, framed as supplementary. Lead with the standard scores; keep age equivalents in a clearly labeled, cautioned place.
Moderately, and asymmetrically, which is the part reports most often get wrong. In the Bayley-III literature, cognitive and language scores at age 2 correlated around .81 and .78 with preschool IQ and most children kept their broad classification from 2 to 4 years (Bode and colleagues, 2014), but a low score is far better at flagging risk than a passing score is at ruling it out: at a cutoff of 70, one study detected only 18% of children whose later IQ fell below 70 (Mansson and colleagues, 2021). Scores below 12 months are especially unstable. The write-up implication is to present an early Bayley-4 score as a current estimate rather than a prediction, to avoid trajectory language a single administration cannot support, and to state the re-evaluation plan, comparing same-norm scores at follow-up.
Yes. Paste your score summary (scales, standard scores, CIs, percentiles, descriptors, and the corrected or chronological age used) along with your behavioral observations, and it drafts the developmental results section for your review: the domains organized with standard scores leading, the caregiver-report scales labeled as such, and the corrected-age decision framed so you can confirm it. It can also cross-check a draft you wrote for the gaps reviewers flag, an unstated norm version, corrected age applied inconsistently, age equivalents reported without cautions, questionnaire scores presented as directly observed, or delay described outside the jurisdiction's metric, and it can produce a plain-language summary for parents and early-intervention teams. BastionGPT is HIPAA-compliant with a signed BAA on every plan, your data is never used to train models, and because you paste only your summary and observations, no test content leaves your records.
The instrument facts and compliance claims on this page trace to these sources, last verified July 2026:
Educational content, not legal or billing advice. Sample notes are fictional. Follow your organization's policies and your board, payer, and jurisdiction requirements.