The Bayley Scales of Infant and Toddler Development are the most widely used standardized tool for evaluating how young children are developing across cognitive, language, and motor abilities. The test is individually administered to children between 1 and 42 months of age, and it produces composite scores in five domains: cognitive, language, motor, adaptive behavior, and social-emotional development.1ScienceDirect. Theoretical Background and Structure of the Bayley Scales of Infant and Toddler Development, Third Edition Clinicians use the results to identify developmental delays, guide early intervention decisions, and track progress over time. But understanding what the scores actually mean, and what they cannot tell you, requires a closer look at how the test works and where its limits are.
What the Test Measures
The Bayley Scales assess a child’s developmental functioning through a combination of tasks the child performs during the session and questionnaires completed by a caregiver. A trained examiner sits with the child, typically for one to two hours, presenting age-appropriate activities: stacking blocks, responding to sounds, following instructions, grasping objects, crawling or walking across the room. The child’s performance is scored against norms derived from a large standardization sample of same-age peers.
The five domains break down into fairly intuitive categories. The cognitive scale measures things like attention, problem-solving, and how a child explores objects. The language scale splits into receptive communication (understanding words and commands) and expressive communication (producing words and gestures). The motor scale divides into fine motor skills (manipulating small objects, hand-eye coordination) and gross motor skills (sitting, standing, walking). The adaptive behavior and social-emotional scales rely on caregiver questionnaires rather than direct observation.2Academic Press. Bayley-III Clinical Use and Interpretation
Scores on each domain are reported as composite scores with a mean of 100 and a standard deviation of 15, similar to how IQ scores are structured. A composite score of 85 falls one standard deviation below the mean, and a score of 70 falls two standard deviations below. Historically, a score below 70 has been the conventional marker for significant developmental delay, though that threshold has become a subject of real debate.
How Reliable Are the Scores
A developmental assessment is only useful if different examiners get roughly the same result and the child scores consistently across sessions. On that front, the Bayley Scales perform well. In a study of children in Nepal, inter-rater reliability was excellent, with agreement coefficients reaching 0.95 to 0.99 across the standardization and quality-control samples. Internal consistency varied by subscale: the cognitive and gross motor subtests showed good reliability, while receptive communication and fine motor showed acceptable consistency, and expressive communication had the lowest internal consistency.3PubMed Central. Acceptability and Reliability of the Bayley Scales of Infant and Toddler Development-III Among Children in Bhaktapur, Nepal
Similarly, a psychometric study in Persian-speaking children found test-retest reliability above 0.98 across all five domains, and inter-rater reliability above 0.99 for cognitive scores.4PubMed Central. A Psychometric Study of the Bayley Scales of Infant and Toddler Development in Persian Language Children So when the same child is tested twice by the same or different examiners in a controlled setting, the numbers tend to hold steady. The challenge is not whether the test measures consistently, but whether what it measures at 18 months tells you something meaningful about the child years later.
What Early Scores Can and Cannot Predict
This is the question parents and clinicians care most about: if a toddler scores below average on the Bayley Scales, does that mean they will have lasting intellectual difficulties? The honest answer is that the test is better at identifying children who are clearly struggling than at making fine distinctions in the average range.
A study of extremely preterm children found that Bayley-III cognitive scores at 18 to 22 months had only a mild correlation with full-scale IQ at ages 6 to 7. Among children who scored below 70 on the Bayley-III cognitive composite, about two-thirds also had IQ scores below 70 at school age. That is meaningful, but it also means a third of children flagged as severely delayed in toddlerhood tested in the normal range later. More troubling, among children who scored in the 85 to 100 range on the Bayley-III (which looks reassuringly “average”), about 43% had IQ scores below 85 at school age, and one in ten had IQ below 70.5PubMed Central. Do Bayley-III Composite Scores at 18-22 Months Corrected Age Predict Full-Scale IQ at 6-7 Years in Children Born Extremely Preterm?
A birth cohort study from South India found that Bayley scores at 36 months predicted later IQ better than scores taken at 24 months, with correlations in the range of 0.40 to 0.49 between 36-month Bayley scores and IQ at age 5.6BMJ Open. Stability and predictability of Bayley Scales of Infant and Toddler Development: evidence from a south Indian birth cohort prospective study In another study of extremely preterm children, the Bayley-III cognitive index explained about 38% of the variance in later IQ, and a cut-off score of 70 had only 18% sensitivity for detecting children who would go on to have IQ below 70.7PubMed. The ability of Bayley-III scores to predict later intelligence in children born extremely preterm In plain terms, using the traditional “below 70” rule, the test missed the vast majority of children who ended up with intellectual difficulties.
The takeaway is not that the Bayley Scales are useless for prediction, but that they work best as a flag for children at highest risk. A very low score should be taken seriously. A score in the average range should not be treated as an all-clear, especially in children with medical risk factors like extreme prematurity.
The Cut-Off Score Problem
One of the most debated aspects of the Bayley Scales is where to draw the line between “delayed” and “not delayed.” When the third edition (Bayley-III) replaced the second edition, researchers quickly noticed that the newer version produced higher scores for the same children. On average, Bayley-III cognitive-language composite scores ran about 6 points higher than the equivalent score on the older Bayley-II, and motor composite scores ran about 8 points higher.8PubMed Central. Comparison of Second and Third Editions of the Bayley Scales in Children With Suspected Developmental Delay That shift meant that children who would have been classified as delayed on the old test now scored above the cut-off on the new one.
A widely cited Australian study drove the point home. Using the standard Bayley-III norms and a cut-off of 70, only about 13% of extremely preterm or extremely low birth weight children were classified as having cognitive delay. But when the researchers recalculated delay rates based on a control group of healthy full-term children tested on the same instrument, roughly a third of the preterm group showed cognitive delay, with even higher rates of language and motor delay.9Archives of Pediatrics & Adolescent Medicine. Underestimation of Developmental Delay by the New Bayley-III Scale In other words, the Bayley-III norms were painting an overly rosy picture.
Researchers have proposed various workarounds. One study found that combining cognitive and language scores below 85 on the Bayley-III produced 99% agreement with the older Bayley-II threshold of below 70 for identifying significant delay.10Pediatric Research. Using the Bayley-III to assess neurodevelopmental delay: which cut-off should be used? Another found that raising the Bayley-III cognitive composite cut-off to 80 optimized the ability to predict later IQ below 70.11Paediatrics & Child Health. Establishing Bayley-III cut-off scores at 21 months for predicting low IQ scores at 3 years of age in a preterm cohort A Turkish study proposed graduated cut-offs: roughly 93 for mild delay, 83 for moderate, and 71 for severe, rather than a single binary threshold.12PubMed Central. Which Bayley-III cut-off values should be used in different developmental levels? There is no single universally agreed-upon cut-off, and clinicians working with high-risk populations are generally advised to interpret scores with higher thresholds than the published norms suggest.
What Changed in the Bayley-4
The fourth edition of the Bayley Scales has been rolling out, and early comparisons to the third edition show a significant recalibration. In a study of very preterm children assessed at 24 months corrected age, Bayley-4 cognitive scores ran about 10 points lower than Bayley-III scores for the same population. Language scores dropped roughly 6 points and motor scores about 4 points. The proportion of children scoring more than two standard deviations below the mean on the cognitive scale jumped from 4% on the Bayley-III to about 13% on the Bayley-4.13Pediatrics. The Bayley-4 Versus the Bayley-III in Very Preterm Children at 24 Months’ Corrected Age
This is a major shift. It suggests the Bayley-4 has addressed the score inflation that plagued the Bayley-III, but it also means that follow-up programs comparing outcomes across eras need to be cautious. A child who scored 85 on the Bayley-III might score in the mid-70s on the Bayley-4, not because they declined, but because the measuring stick changed. Research using the Bayley-4 in children with autism spectrum disorder and developmental delay has also begun, with early findings suggesting that children identified with autism at very young ages tend to show global developmental deficits across all Bayley-4 subtests, which may help distinguish them from children with more targeted language delays.14Psychology in the Schools. Bayley‐4 performance of very young children with autism, developmental delay, and language impairment
Who Gets Tested and Why
The Bayley Scales are used most heavily in neonatal follow-up programs for preterm infants, where they have become essentially the standard outcome measure. A study of 70 preterm neonates found that children born before 30 weeks of gestation had mean cognitive composite scores around 80, compared to about 86 for those born after 30 weeks. Over 80% of the earliest-born group showed at least moderate cognitive delay.15Journal of Neonatology. Neurodevelopmental Outcomes of Preterm Infants Using Bayley Scale of Infant Development-III (BSID-III): A Tertiary Care Centre Study
The test is also used extensively in babies who experienced birth asphyxia or neonatal encephalopathy. In one study of infants who underwent cooling therapy for hypoxic-ischemic encephalopathy, over half showed cognitive delay and over half showed motor delay on the Bayley-III.16PubMed Central. Neurodevelopmental evaluation of newborns who underwent hypothermia with a diagnosis of hypoxic ischemic encephalopathy based on the Bayley-III scale Importantly, research has shown that infants who received therapeutic hypothermia scored significantly better on the Bayley-III than those who did not, which has helped validate cooling as a standard treatment.17Iranian Journal of Neonatology. Study of Neurodevelopmental Outcomes at 10-14 Months of Age Using Bayley Scale of Infant and Toddler Development in Asphyxiated Newborns with Hypoxic Ischemic Encephalopathy Treated with and without Therapeutic Hypothermia
Beyond prematurity and birth injury, the scales are used in children with genetic conditions like Down syndrome, where Bayley-III profiles have revealed distinct delay patterns. Researchers identified three clusters within a Down syndrome sample: mild, moderate, and pronounced delay profiles, with factors like heart surgery and receipt of occupational therapy influencing which profile a child fit into.18PubMed. Early developmental profiles among infants with Down syndrome
How the Home Environment Shapes Scores
Bayley scores do not reflect biology in a vacuum. A child’s home environment, family income, and parental education all influence the numbers. A cohort study measuring multiple environmental variables found that the degree to which parents promoted their child’s autonomy, such as letting the child explore and make choices, was independently linked to higher Bayley-III cognitive, language, and motor scores. That environmental factor also partially mediated the relationship between socioeconomic status and cognitive development, meaning some of the effect of family income on a child’s score operated through the quality of the home environment rather than income alone.19PLoS ONE. The Complex Interaction between Home Environment, Socioeconomic Status, Maternal IQ and Early Child Neurocognitive Development: A Multivariate Analysis of Data Collected in a Newborn Cohort Study
In a study of infants from a low-income area in São Paulo, Brazil, socioeconomic status was positively associated with language and motor performance, and fewer years of maternal education correlated with lower language and cognitive scores.20Trends in Psychiatry and Psychotherapy. Socioeconomic diversities and infant development at 6 to 9 months in a poverty area of São Paulo, Brazil These findings are a reminder that a low Bayley score does not automatically point to a neurological problem. It can reflect limited stimulation, poverty, or cultural differences in how children are raised.
Cross-Cultural Challenges
The Bayley Scales were normed on children in the United States, and applying them in other countries is not straightforward. Some test items rely on materials or activities that are culturally specific: a toy that is common in American homes may be unfamiliar to a child in rural Kenya or Ethiopia. When researchers culturally adapted the Bayley-III for Kenyan children aged 18 to 36 months, they found the adapted version performed reasonably well, but children scored lower than the normative population even after adaptation, likely because of cultural differences in how test items and scoring work, as well as differences in maternal education levels.21PubMed Central. Cultural Adaptation of the Bayley Scales of Infant and Toddler Development, 3rd Edition for use in Kenyan Children Aged 18–36 Months: A Psychometric Study
An Ethiopian adaptation study found that most items scaled onto the existing subscales with good internal consistency, but some items were simply not relevant in the local context and had to be removed or replaced.22PubMed Central. Adapting the Bayley Scales of infant and toddler development in Ethiopia: evaluation of reliability and validity These adaptation studies matter because the Bayley Scales are increasingly used in global health research and clinical trials in low- and middle-income countries. Without local norms, there is a real risk of labeling children as delayed when they are developing normally for their context.
The Bayley Scales Versus Screening Questionnaires
Parents sometimes encounter shorter tools like the Ages and Stages Questionnaire (ASQ), a parent-completed screener that takes minutes rather than the one to two hours the full Bayley assessment requires. These are different tools designed for different purposes. The ASQ is meant to flag children who need a closer look; the Bayley Scales are the closer look.
Studies comparing the two consistently show that the ASQ has high specificity, meaning it is good at correctly identifying children who do not have a delay, but lower sensitivity, meaning it misses a meaningful number of children who are delayed. In a study of extremely low birth weight infants, the ASQ would have missed about 27% of children who scored poorly on the Bayley-II.23PubMed Central. Use of the Ages and Stages Questionnaire and Bayley Scales of Infant Development-II in Neurodevelopmental Follow-up of Extremely Low Birth Weight Infants A Singapore cohort study comparing the ASQ-3 to the Bayley-III found strong specificity (often above 80%) but sensitivity as low as 19% in some motor domains.24Pediatrics & Neonatology. Concurrent validity of the ages and stages questionnaires with Bayley Scales of Infant Development-III at 2 years – Singapore cohort study In practice, this means the ASQ works well as a first pass: if the screener flags a concern, a full Bayley assessment is warranted. But a clean ASQ result in a high-risk child should not be treated as definitive.
When Scores Are Unstable
One issue that catches parents off guard is that a child’s Bayley classification can change substantially from one assessment to the next, especially in the first two years. A longitudinal study that tested low-risk and high-risk infants seven times between 3 and 24 months found highly unstable delay classifications. A child classified as delayed at one visit might test in the normal range at the next, or vice versa. The study’s sensitivity and positive predictive values for identifying persistent delay were poor across time points.25PubMed Central. Instability of delay classification and determination of early intervention eligibility in the first two years of life
This instability does not mean the test is broken. Infant development is genuinely variable. A baby might be slow to walk but then catch up rapidly; language may appear to lag at 12 months and then explode at 18 months. The Bayley Scales capture a snapshot, not a trajectory, and clinicians are generally advised to weigh the overall clinical picture, including medical history and risk factors, rather than relying on a single score at a single time point.
Brain Imaging and Bayley Scores
Researchers have started connecting Bayley scores to what they see on brain MRI, particularly in children who experienced birth complications. In infants with neonatal encephalopathy, several MRI scoring systems showed significant correlations with Bayley-III motor and cognitive composite scores, with the most comprehensive scoring system explaining about 30% of the variance in motor outcomes and about 26% in cognitive outcomes.26Pediatric Neurology. Relationship Between MRI Scoring Systems and Neurodevelopmental Outcome at Two Years in Infants With Neonatal Encephalopathy In healthy full-term newborns, the structural integrity of white matter tracts measured shortly after birth showed positive correlations with later Bayley cognitive, language, and motor scores.27PubMed Central. Diffusion Tensor MRI of White Matter of Healthy Full-term Newborns: Relationship to Neurodevelopmental Outcomes
These findings are still mostly in the research domain rather than guiding individual clinical decisions, but they reinforce that Bayley scores are measuring something real about brain development, not just a child’s mood on the day of testing.
Growth Scale Values for Tracking Change Over Time
Standard composite scores (the ones centered around 100) are designed to compare a child to same-age peers at a single point. They are less useful for tracking how an individual child’s abilities change over time, especially in children with progressive neurological conditions where development may plateau or regress. For this purpose, researchers increasingly turn to Growth Scale Values (GSVs), which are derived from the Bayley’s raw scores using a mathematical transformation that creates a true interval scale. A given difference in GSVs represents the same difference in ability no matter where on the scale you are, making it possible to measure genuine gains or losses rather than shifts in age-relative standing.28PubMed Central. Increasing precision in the measurement of change in pediatric neurodegenerative disease
GSVs have become recommended in natural history studies and clinical trials for children with neurological diseases, where the question is not “how does this child compare to peers” but “is this child gaining or losing skills.” For families and clinicians following a child with a neurodegenerative condition, GSVs offer a more precise way to measure whether an intervention is working.
Caregiver Report and Direct Assessment
The Bayley-4 continues to use a mix of direct observation and caregiver questionnaires, but how much do these two approaches agree? A study comparing Bayley-4 cognitive scores (from direct testing) with the Developmental Profile-4 cognitive scores (from parent report) found a moderately strong correlation between the two. The age of the child and the severity of autism-like features were also significant factors in how parent-reported and directly tested scores related to each other.29Journal of Autism and Developmental Disorders. Developmental Assessment in Children at Higher Likelihood for Developmental Delays – Comparison of Parent Report and Direct Assessment This is clinically useful because some children, particularly those with autism or severe anxiety, may not perform to their ability in a structured testing session. In those cases, caregiver input can provide a fuller picture of what the child can do at home.
The Bayley-4 also requires substantial examiner training. Assessors in one study completed a two-day accredited training program followed by supervised administrations before being allowed to score independently, with the authors noting that translation of these standards to routine clinical settings requires consideration of training costs and protected clinician time.30Pediatric Research. Inter-rater reliability and agreement of the Bayley-4 in a multidisciplinary team This is not a test that any practitioner can pick up and administer on a whim, and the quality of the results depends heavily on the examiner’s competence. For parents, this means it is reasonable to ask about the examiner’s training and experience, especially if the results will drive major decisions about intervention services or school placement.

