The Stanford-Binet Intelligence Scales are one of the oldest and most widely used individually administered IQ tests in the world, first developed in the early 1900s and now in its fifth edition (SB-5). The test measures cognitive ability across five broad domains in both verbal and nonverbal formats, producing a full-scale IQ score along with several composite scores that help clinicians understand a person’s cognitive strengths and weaknesses. It is used to evaluate children and adults for gifted programs, learning disabilities, intellectual disability, and developmental conditions like autism, and its structure has evolved considerably over more than a century of revisions.
What the Test Actually Measures
The fifth edition of the Stanford-Binet was designed to assess five factors considered important to intelligence, and it does so across both verbal and nonverbal channels.1Journal of Psychoeducational Assessment. Investigating the Theoretical Structure of the Stanford-Binet-Fifth Edition Those five factors are fluid reasoning, knowledge, quantitative reasoning, visual-spatial processing, and working memory. Each factor is tested with verbal tasks (where language is involved in asking and answering) and nonverbal tasks (where the test-taker responds by pointing, arranging objects, or completing visual patterns). This dual structure gives clinicians a richer picture than a single overall score would provide.
The design of the SB-5 draws heavily from the Cattell-Horn-Carroll (CHC) framework, a widely accepted model of cognitive abilities that organizes intelligence into broad and narrow factors. Research over two decades has confirmed that the factors produced by most major intelligence tests, including the Stanford-Binet, align well with this theoretical model.2Wiley Online Library. Cattell–Horn–Carroll abilities and cognitive tests: What we’ve learned from 20 years of research In practical terms, this means the test is not just measuring one undifferentiated “smartness.” It maps onto distinct cognitive abilities that people can have in different proportions, which matters when a clinician is trying to figure out why a child is struggling in school or excelling in certain areas but not others.
How a Testing Session Works
Unlike a group-administered standardized test you might take with a roomful of other people, the Stanford-Binet is given one-on-one by a trained examiner. The session typically takes between 45 minutes and about 90 minutes, depending on the person’s age and the number of subtests administered. Testing begins with two “routing” subtests that help the examiner determine the right difficulty level for the rest of the battery. This adaptive starting point is a key feature: a five-year-old and a twelve-year-old will not be given the same items from the outset, and within a subtest the difficulty ramps up until the test-taker reaches their ceiling.
The test can be used across a very wide age range, from about two years old through adulthood. For very young children, tasks lean heavily on nonverbal items like sorting objects, pointing at pictures, and manipulating blocks. Older children and adults encounter more abstract reasoning, vocabulary, and pattern-completion tasks. The SB-5 also offers an abbreviated battery (the ABIQ), which uses only the two routing subtests to produce a quick estimate of full-scale IQ. Research on nearly 1,700 youth with autism or ADHD found that this abbreviated score measures intelligence in the same way for both diagnostic groups and that scatter between the routing subtests did not differ between children with autism and those with ADHD.3PubMed. Measuring intelligence in Autism and ADHD: Measurement invariance of the-Binet 5th edition and impact of subtest scatter on abbreviated IQ accuracy That said, the abbreviated version trades precision for speed, and clinicians often opt for the full battery when high-stakes decisions are involved.
Clinical Uses in Autism and ADHD
The Stanford-Binet has become a go-to instrument in developmental assessment clinics, particularly for evaluating young children with autism spectrum disorder (ASD) or attention-deficit/hyperactivity disorder (ADHD). A large study using SB-5 data from over 1,600 youth aged two to sixteen identified 14 distinct cognitive profiles among children with ASD or ADHD, with those profiles appearing at roughly parallel levels of cognitive functioning across both conditions.4PubMed. Identifying Latent Cognitive Profiles in Autism Spectrum Disorder and Attention-Deficit/Hyperactivity Disorder Using the Stanford-Binet Intelligence Scales-5th Edition This kind of research helps clinicians move beyond a single IQ number and instead look at a child’s pattern of strengths and weaknesses.
One practical reason the SB-5 is favored for autism evaluations is its nonverbal domain. Children with autism who have limited spoken language can still demonstrate problem-solving ability through nonverbal subtests, giving a more complete picture than a heavily language-dependent test would. The test is also structured to engage younger children through hands-on materials, which matters when assessing a two- or three-year-old who cannot sit through a workbook-style evaluation.
Gifted Identification
At the other end of the ability spectrum, the Stanford-Binet has long been used to identify intellectually gifted children. Its ceiling extends high enough to differentiate among children performing well above average, which is important for school districts that need cutoff scores for gifted programs. Earlier editions were especially popular for this purpose, and research with the fourth edition found that correlations between the Stanford-Binet composite score and the Peabody Picture Vocabulary Test ranged from about .30 to .57 in a group of gifted students, suggesting that a quick vocabulary screening test has some value as a first pass but is not a replacement for a full cognitive assessment.5Diagnostique. Relationship between Scores of Gifted Children on Stanford-Binet IV and Peabody Picture Vocabulary Test — Revised
This finding reflects a broader principle: intelligence, as measured by tests like the Stanford-Binet, involves a range of abilities that a vocabulary test alone cannot capture. A child who is exceptionally strong in fluid reasoning or visual-spatial processing but average in verbal knowledge would look ordinary on a vocabulary screener and exceptional on a full Stanford-Binet battery. For gifted placement decisions, the multi-factor structure of the SB-5 gives evaluators considerably more information to work with than a single-score screener.
Challenges at the Lower End of the Scale
One of the more significant technical issues with the Stanford-Binet involves what happens when it is used to assess people with intellectual disability, particularly those with moderate-to-severe impairments. IQ tests in general have limited range and precision at the lower end of the scale.6PubMed Central. Improving IQ measurement in intellectual disabilities using true deviation from population norms The SB-5’s full-scale IQ score bottoms out at 40, and while that floor is lower than what some competing instruments offer, it still creates a problem for individuals whose true cognitive ability falls below that cutoff.
Research on Fragile X syndrome illustrates this vividly. In one study, half of the males and about 7% of the females with Fragile X had a full-scale IQ score right at the floor of 40 on the SB-5.7PubMed Central. Brief Report: Differences Between Stanford-Binet Abbreviated and Full-Scale Estimates of IQ in Fragile X Syndrome Vary Across Development When so many individuals are piled up at the minimum possible score, the test cannot distinguish between someone whose true ability is mildly below the floor and someone whose ability is far below it. This is called a floor effect, and it gets worse with age in Fragile X because IQ scores tend to decline not because adults are losing skills, but because their skill growth does not keep pace with same-age peers. The test treats standing still as falling behind.
Researchers have proposed workarounds, including statistical techniques that estimate “true deviation” scores below the published floor. These approaches can provide more meaningful differentiation among individuals with severe intellectual disability, but they are not part of the standard scoring procedures a clinician receives from the test publisher. For now, anyone reviewing a Stanford-Binet report showing an IQ of 40 should understand that the number may represent a floor rather than an actual measured ability level.
How the Stanford-Binet Compares to Wechsler Scales
The Stanford-Binet’s main competitor in clinical practice is the family of Wechsler intelligence tests, which includes the WAIS for adults, the WISC for school-age children, and the WPPSI for younger children. Both test families measure overlapping constructs and are normed on large representative samples, but they do not always produce the same scores for the same person, and the discrepancies can be clinically meaningful.
A study comparing Stanford-Binet and Wechsler Adult Intelligence Scale (WAIS) scores in adults with intellectual disability found that the WAIS full-scale IQ was higher than the Stanford-Binet composite IQ in every single case, with an average difference of nearly 17 points.8PubMed Central. Stanford-Binet & WAIS IQ Differences and Their Implications for Adults with Intellectual Disability (aka Mental Retardation) Additional comparisons with other measures suggested the WAIS may systematically underestimate the severity of intellectual impairment. That 17-point gap is not trivial. In contexts where an IQ score determines whether someone qualifies for disability services, housing support, or legal protections, the choice of test instrument can change the outcome.
This discrepancy does not appear to be caused simply by the Stanford-Binet’s lower floor. Even when the Stanford-Binet scores were well above their minimum, the WAIS still came out higher. The implication for clinicians is that scores from different IQ tests are not directly interchangeable, and a person assessed with one test cannot be assumed to have “the same” IQ on another. For families navigating disability evaluations, this is worth knowing: if a Wechsler test places someone just above a service-eligibility cutoff, a Stanford-Binet evaluation might place them below it.
The Role of IQ Testing in Special Education Law
Intelligence testing, including the Stanford-Binet, plays a complicated and sometimes contentious role in special education in the United States. Under the Individuals with Disabilities Education Act (IDEA), school districts use cognitive assessments to help determine whether a student qualifies for special education services. But due to a combination of legal precedent, state-level regulations, and high-profile court cases, intelligence tests are used inconsistently in special education decision-making across the country.9PubMed Central. Intelligence and the Individuals with Disabilities Education Act
Some states require IQ testing as part of the evaluation process for specific disability categories, while others allow alternative assessment approaches. In California, the use of standardized IQ tests for placing Black students in special education classes was restricted following the Larry P. v. Riles case in the 1970s, a ruling that reflected concerns about cultural bias in testing. This restriction was eventually relaxed, but the case illustrates how legal and political dynamics shape which tests get used and on whom. A parent seeking an evaluation for their child may find that the available testing options depend heavily on where they live and what disability category is being considered.
A Troubled History with Eugenics
The Stanford-Binet’s origins are inseparable from one of the darker chapters in American psychology. Lewis Terman, the Stanford psychologist who adapted Alfred Binet’s original French test for American use in 1916, was an active proponent of the eugenics movement. Under the testing frameworks of that era, people who scored poorly were labeled with terms like “moron,” “imbecile,” and “idiot,” classifications that carried legal and institutional consequences. Adults and children with low scores could be institutionalized, and in many states they could be involuntarily sterilized. Some of those individuals would today be recognized as having autism, learning disabilities, or other developmental differences rather than global intellectual deficits.
The Army Alpha and Beta tests used during World War I, which drew on the same intellectual tradition, were used to rank recruits and sort them into roles, but their results were also cited to argue for immigration restrictions targeting groups that scored lower on average. This legacy has left a lasting shadow over IQ testing broadly, and the Stanford-Binet specifically. Modern versions of the test have been substantially redesigned, normed on representative samples, and subjected to bias analyses, but critics rightly point out that the history of how intelligence testing was wielded as a tool of social control should inform how cautiously we interpret and apply scores today.
Socioeconomic Factors and Score Interpretation
Decades of research, stretching back well over a century, have documented a consistent positive correlation between socioeconomic status and IQ test performance. People from higher-income households with more education tend to score higher on tests like the Stanford-Binet, and this relationship has been replicated so many times that it is among the most robust findings in the field. The reasons are not simple and involve a tangle of factors including access to educational resources, nutrition, healthcare, exposure to environmental toxins, and the degree to which a child’s home environment stimulates the kinds of reasoning the test measures.
This matters for interpretation because an IQ score is not a pure readout of someone’s innate cognitive potential. A child who grew up in poverty, attended under-resourced schools, and experienced chronic stress may score lower on a Stanford-Binet than they would have under more supportive conditions. Clinicians are trained to consider these contextual factors when interpreting results, but how much weight they give them varies. For parents, the takeaway is that an IQ score is one piece of information, not a final verdict. It reflects performance on a given day under a given set of conditions, shaped partly by the child’s abilities and partly by everything that led up to that testing session.
What Brain Research Says About Intelligence
While IQ tests like the Stanford-Binet measure cognitive performance through behavioral tasks, a separate line of research has investigated what is happening in the brain that correlates with differences in intelligence. A large study using brain imaging data from the UK Biobank found that the correlation between total brain volume and a general intelligence factor was about 0.28, a statistically reliable relationship but a modest one, meaning brain size explains only a small fraction of the variation in intelligence across people.10Intelligence. Structural brain imaging correlates of general intelligence in UK Biobank
More telling than total volume is the distribution of gray matter. Research comparing brain scans with IQ scores has found that greater amounts of gray matter in specific regions of the frontal, temporal, parietal, and occipital lobes are associated with higher intelligence, underscoring that the brain basis of intelligence is distributed rather than localized to a single area.11PubMed. Structural brain variation and general intelligence Network-level analyses go further, showing that higher intelligence scores correspond to more efficient information transfer across the brain, as reflected by shorter path lengths and higher global efficiency in brain networks.12PLoS Computational Biology. Brain Anatomical Network and Intelligence
None of this brain research validates or invalidates the Stanford-Binet specifically. What it does suggest is that the “general intelligence” factor that tests like the SB-5 are designed to capture has genuine biological underpinnings rather than being a pure artifact of test design. At the same time, the modest size of the brain-volume correlation reminds us that intelligence is not reducible to any single anatomical measurement. The brain’s wiring patterns and efficiency seem to matter at least as much as its raw size.
Choosing Between the Stanford-Binet and Other Tests
If you or your child has been referred for cognitive testing, you may not have a choice about which test is used. The evaluating clinician typically selects the instrument based on the referral question, the person’s age, and sometimes institutional or legal requirements. But understanding the trade-offs can help you have a more informed conversation with the evaluator.
The Stanford-Binet tends to be preferred when the person being tested is very young (under about five), when a nonverbal pathway is important due to limited language, or when the referral question involves very high or very low ability levels where the test’s extended range is useful. Wechsler tests tend to dominate in school-age evaluations for learning disabilities and are more commonly used in neuropsychological batteries for adults. Neither is categorically “better.” They measure overlapping but non-identical constructs, and as the 17-point discrepancy study showed, the scores they produce for the same individual are not always comparable.13PubMed Central. Stanford-Binet & WAIS IQ Differences and Their Implications for Adults with Intellectual Disability (aka Mental Retardation)
If you are in a situation where a score difference between tests could affect eligibility for services, it is worth asking the evaluator why they chose the instrument they did and whether an alternative might produce a materially different result. This is not about gaming the system; it is about recognizing that test selection is itself a decision with consequences, and an informed consumer is better positioned to advocate for an accurate assessment.

