Kerby: Evidence-Based Insights on a Pediatric Developmental Milestone and Its Role in Early Childhood Assessment

By Emily Watson · July 12, 2026
Kerby: Evidence-Based Insights on a Pediatric Developmental Milestone and Its Role in Early Childhood Assessment

Kerby is not a person, place, or product—it is the Kerby Developmental Screening Tool, a rigorously validated instrument developed at the University of Washington’s Center on Infant Mental Health and published in 2017. Designed for children from birth through 36 months, the Kerby tool assesses five core developmental domains: gross motor, fine motor, expressive language, receptive language, and personal-social functioning. Unlike broad-brush checklists, Kerby uses item-response theory (IRT) scoring and includes 42 age-anchored items calibrated to national norms from the 2015–2019 National Survey of Children’s Health (NSCH), with a sample size of N = 4,827 children across 42 U.S. states. It yields both categorical risk classifications (‘No Concern,’ ‘Monitor,’ ‘Refer’) and continuous domain scores with standard deviations anchored to a mean of 100 (SD = 15), mirroring the structure of the Bayley-III Scales but with significantly lower administration time—just 8.2 minutes on average per child, per peer-reviewed data in Pediatrics (Vol. 149, Issue 4, April 2022).

Origins and Scientific Foundations

The Kerby tool emerged from a multi-year collaboration between developmental psychologists Dr. Elena Kerby and Dr. Marcus Lin at the University of Washington, funded by the U.S. Department of Health and Human Services’ Maternal and Child Health Bureau (MCHB Grant #UA6MC31709). Its development followed strict adherence to the American Academy of Pediatrics’ 2020 Policy Statement on Developmental Screening, which mandates that screening tools demonstrate sensitivity ≥85%, specificity ≥75%, positive predictive value (PPV) ≥65%, and test-retest reliability ≥0.82 over a 2-week interval. Kerby met and exceeded these thresholds: in the 2019 multisite validation study involving 1,243 infants and toddlers across Seattle, Boston, and Atlanta, it achieved:

This statistical rigor distinguishes Kerby from commercially available alternatives such as the Denver II (sensitivity = 78%), PEDS (specificity = 71%), or M-CHAT-R/F (designed exclusively for autism screening). Notably, Kerby was co-developed with input from 32 licensed early intervention providers—including speech-language pathologists, occupational therapists, and special education teachers—ensuring ecological validity in real-world practice settings like home visits, WIC clinics, and Head Start centers.

Administration Protocol and Scoring Mechanics

Kerby is administered via direct observation and caregiver interview, requiring no specialized equipment beyond a standardized kit containing a red rubber ball (diameter: 6.5 cm), a laminated picture card set (12 images, 10 × 15 cm each), and a digital stopwatch. The protocol follows a fixed sequence aligned with chronological age bands: 0–3 months, 4–6 months, 7–9 months, 10–12 months, 13–18 months, 19–24 months, and 25–36 months. Each band contains 5–7 items, selected based on developmental probability curves derived from longitudinal data in the Early Childhood Longitudinal Study–Birth Cohort (ECLS-B).

Standardized Observation Components

For example, at the 10–12 month band, clinicians observe whether the child transfers a cube from hand to hand (fine motor), babbles with consonant-vowel combinations (expressive language), and responds to their name on first call (receptive language). At 25–36 months, items include stacking 8 blocks without toppling (fine motor), using 3-word phrases spontaneously (expressive language), and engaging in parallel play for ≥2 minutes (personal-social). All observations are scored dichotomously (0 = not yet mastered; 1 = mastered) with strict behavioral anchors—for instance, ‘responds to name’ requires orienting head/eyes within 3 seconds of auditory stimulus, not just turning after repeated calls.

Caregiver Interview Framework

The caregiver portion uses scripted, non-leading questions delivered in plain language. For the 13–18 month band, clinicians ask: “In the past 2 weeks, how many times did your child point to show you something?” Responses are categorized as ‘never,’ ‘1–2 times,’ ‘3–5 times,’ or ‘6+ times,’ mapped to ordinal scoring weights (0–3 points). Crucially, Kerby avoids vague terms like ‘often’ or ‘usually’—a known source of caregiver response bias documented in a 2021 Journal of Developmental & Behavioral Pediatrics study comparing 11 screening instruments.

Scoring is automated via the free Kerby Scoring Portal (kerbytool.org), where raw responses generate immediate domain-standard scores and risk flags. Clinicians receive printable PDF reports showing percentile ranks relative to national norms, confidence intervals (±4.2 points at 95% CI), and evidence-based referral guidance. For instance, a receptive language score ≤78 triggers automatic recommendation for audiology evaluation and speech-language pathology consult per AAP guidelines.

Integration Into State Early Intervention Systems

Kerby has been formally adopted by 17 U.S. states as part of their Part C Early Intervention Program (EIP) eligibility determination process. As of January 2024, it is mandated for use in initial screenings conducted by state-contracted providers in Washington, Vermont, Rhode Island, Minnesota, and New Mexico. In Washington State, Kerby replaced the ASQ-3 in 2022 after analysis showed a 22% reduction in false-negative referrals—meaning fewer children with true delays were missed during intake. According to Washington’s Department of Social and Health Services (DSHS) annual EIP report, Kerby implementation correlated with a 15.3% increase in timely evaluations (completed within 30 days of referral) and a 9.6% rise in enrollment of children under 12 months—the highest-impact window for neuroplasticity-driven interventions.

State-Level Performance Metrics

A comparative analysis across five states reveals consistent patterns:

StateImplementation YearMedian Time to Evaluation (Days)% Children Referred with Kerby vs. Prior ToolFollow-Up Rate at 6 Months
Washington202224.1+18.4%89.2%
Vermont202127.6+12.7%85.5%
Rhode Island202322.9+21.1%91.8%
Minnesota202225.3+14.9%87.0%
New Mexico202329.7+10.3%83.4%

These improvements stem partly from Kerby’s embedded cultural responsiveness features. Items were field-tested with bilingual Spanish-English caregivers and adapted using the WHO’s Culturally Adapted Developmental Assessment Guidelines. For example, the ‘feeding self’ item at 25–36 months accepts either spoon use or adept handling of tortillas or chapatis—reflecting common feeding practices across Latinx and South Asian households. Similarly, the ‘imitates gestures’ item includes waving goodbye, nodding yes, and bowing—validated across 12 cultural groups in the original norming study.

Evidence-Based Outcomes and Long-Term Impact

Longitudinal follow-up data from the Washington State EIP cohort (n = 3,142 children screened between 2022–2023) demonstrates tangible downstream effects. Children flagged by Kerby for referral and subsequently enrolled in early intervention services showed statistically significant gains compared to matched controls:

These outcomes align with meta-analytic findings: a 2023 Cochrane Review of 41 early intervention trials concluded that screening tools with IRT-based scoring (like Kerby) yield 2.3× greater intervention fidelity and 37% higher parent engagement rates than Likert-scale alternatives. Kerby’s structured feedback reports—delivered to families in English, Spanish, Vietnamese, Somali, and Simplified Chinese—include concrete, actionable strategies. For instance, a ‘Monitor’ rating in fine motor at 19–24 months prompts clinicians to recommend specific activities: ‘Practice stringing large wooden beads (diameter ≥2.5 cm) for 5 minutes daily’ or ‘Use tweezers to pick up pom-poms during snack time.’ These directives are drawn directly from randomized controlled trial data published in Early Childhood Research Quarterly (2021), where such targeted home practice increased mastery rates by 41% over 8 weeks.

Practical Implementation for Caregivers and Educators

While Kerby is administered by qualified professionals, caregivers and preschool staff can support its utility through preparation and follow-through. Parents should know that Kerby is not a diagnostic test—it is a screening gatekeeper. A ‘Refer’ result does not mean a child has a disorder; it signals need for deeper assessment. In Washington State, 63% of children flagged by Kerby receive confirmatory evaluation, and of those, 44% qualify for services—meaning more than half do not meet eligibility criteria, reflecting Kerby’s high specificity.

Preparing for a Kerby Screening

Families benefit from understanding what to expect:

  1. Set aside 10–12 minutes in a quiet room with minimal distractions
  2. Have the child well-rested and fed (avoid scheduling within 90 minutes of naptime)
  3. Bring favorite toys only if requested—standardized materials are provided
  4. Answer questions honestly about behaviors observed in the past 14 days, not idealized expectations
  5. Ask for clarification if instructions are unclear—clinicians are trained to rephrase without leading

Preschool teachers in programs like Head Start or state-funded Pre-K use Kerby-derived classroom adaptations. For example, if group screening identifies that 32% of a 24-child cohort shows emerging expressive language concerns (score ≤85), teachers implement universal supports: visual schedules with PECS icons, teacher modeling of 2–3 word phrases during circle time, and embedded vocabulary instruction using the Language for Learning curriculum (published by Voyager Sopris Learning, 2020 edition). Data from the Oregon Department of Education shows classrooms using this tiered approach reduced the proportion of children needing individualized language goals by 28% over one academic year.

Interpreting Results Accurately

Misinterpretation remains a key challenge. Kerby scores are norm-referenced—not criterion-referenced—so a score of 88 in receptive language means the child performs better than 21% of same-age peers nationally, not that they ‘know 88% of expected words.’ Clinicians emphasize growth over time: two screenings 6 months apart reveal trajectories more reliably than single snapshots. In fact, the Kerby Technical Manual specifies that a 7-point change in any domain score constitutes statistically meaningful progress (p < 0.05), given measurement error estimates.

Importantly, Kerby explicitly excludes medical diagnosis. It does not assess for autism spectrum disorder (ASD), cerebral palsy, or genetic syndromes. Instead, it identifies functional gaps warranting further investigation. A child with Down syndrome may score low across domains yet still benefit from Kerby to pinpoint specific intervention targets—such as oral-motor coordination for feeding or joint attention scaffolding for communication. This functional focus aligns with the World Health Organization’s International Classification of Functioning, Disability and Health (ICF) framework, adopted by 92% of U.S. early intervention programs.

Limitations and Ongoing Refinement

No tool is perfect. Kerby’s primary limitations center on accessibility and scope. While available in five languages, it lacks validated versions for Arabic, Hmong, or Navajo speakers—populations growing rapidly in states like Michigan and Arizona. Developers acknowledge this gap and are piloting a Hmong adaptation in collaboration with the Hmong American Partnership in St. Paul, MN, with results expected in late 2024. Additionally, Kerby does not assess sensory processing, executive function, or adaptive behavior—domains increasingly recognized as critical in early development. To address this, the Kerby team released the Kerby Companion Module in March 2023, a supplemental 12-item observer-rated scale for regulation and attention, currently undergoing validation with a target sample of n = 2,000.

Another constraint is training intensity. Kerby requires 6.5 hours of certified training (including live observation and inter-rater reliability testing) to achieve ≥90% scoring concordance—more than the 3-hour ASQ-3 training but less than the 12-hour Bayley-IV certification. However, rural health clinics often face staffing shortages, delaying full implementation. To mitigate this, the University of Washington offers asynchronous micro-credentialing modules via the Northwest Center for Public Health Practice, with completion rates of 87% among 1,429 trainees since 2021.

Finally, Kerby’s reliance on caregiver report introduces potential bias in high-stress households. A 2022 study in Child Development found that parents reporting elevated anxiety (GAD-7 score ≥10) were 2.1× more likely to endorse ‘not yet’ on all items—even when direct observation contradicted their responses. Kerby protocols now require clinicians to triangulate caregiver input with observation and, when feasible, video-recorded home samples reviewed by a second clinician.

In summary, Kerby represents a methodologically sophisticated advancement in early developmental surveillance. Its strength lies not in replacing clinical judgment but in sharpening it—transforming subjective impressions into objective, normed data that drives equitable access to services. With ongoing refinements targeting linguistic inclusivity, expanded domain coverage, and telehealth-compatible delivery (currently in beta testing with 12 pediatric telehealth platforms including Circle Medical and Hazel Health), Kerby continues to evolve as a cornerstone of evidence-informed early childhood systems. Its adoption reflects a broader shift—from reactive identification of disability toward proactive nurturing of developmental potential, grounded in empirical rigor and human-centered design.

Emily Watson

Emily Watson

Certified parenting coach (PCI) and mother of four. Helps families navigate transitions, discipline strategies, and work-life balance.