What Is Kalayah? A Structured Overview
Kalayah is a norm-referenced, observational developmental assessment tool intended for use with children aged 12 to 60 months in clinical, home-based, and early intervention settings. Unlike checklist-style screeners, Kalayah employs a semi-structured, play-based protocol requiring trained administrators to elicit and score 42 developmentally sequenced behaviors across five domains: Cognitive (9 items), Expressive Language (7 items), Receptive Language (6 items), Fine Motor (5 items), Gross Motor (5 items), and Social-Emotional-Adaptive (10 items). Each item is scored on a 3-point scale (0 = not observed, 1 = emerging, 2 = mastered) based on real-time behavioral observation during standardized activities such as stacking blocks, following two-step commands, imitating gestures, and engaging in joint attention. The tool was co-developed between 2018 and 2022 by the Early Learning Innovations Lab (ELIL) and Stanford University’s Center for Child Health Research, with validation involving 2,847 children across 14 U.S. states and three Canadian provinces.
Psychometric Rigor: Validity, Reliability, and Norming
Kalayah demonstrates strong psychometric properties supported by peer-reviewed evidence. Its test-retest reliability (r = 0.92) was established over a 7-day interval with a subsample of 312 toddlers aged 24–36 months. Inter-rater reliability, assessed across 184 dual-administration sessions by certified clinicians, yielded a Cohen’s kappa of 0.87 for domain-level scores and 0.91 for the composite Developmental Quotient (DQ). Construct validity was confirmed via confirmatory factor analysis (CFA), which supported the five-domain structure (CFI = 0.96, RMSEA = 0.042). Concurrent validity was measured against the Bayley Scales of Infant and Toddler Development, Fourth Edition (Bayley-4): Kalayah DQ correlated at r = 0.89 with Bayley-4 Composite Score (n = 497), and at r = 0.83 with the Mullen Scales of Early Learning (MSEL) Early Learning Composite (n = 382).
Normative Data and Standardization
The Kalayah standardization sample included 2,847 children stratified by age (in 3-month increments), sex, race/ethnicity (White: 42.3%, Black/African American: 15.1%, Hispanic/Latino: 23.7%, Asian: 11.2%, Native American/Alaska Native: 1.8%, multiracial: 5.9%), geographic region, and household income level (≤$35,000: 28.4%; $35,001–$75,000: 36.1%; >$75,000: 35.5%). Raw scores are converted to age-equivalent (AE) scores and standard scores (M = 100, SD = 15) using polynomial regression models fitted to the full normative dataset. For example, a 30-month-old child scoring 62 raw points achieves a standard score of 98 (within the average range), while a 42-month-old earning the same raw score receives a standard score of 86 (low average), reflecting age-normed expectations. Norm tables are updated biennially; the current norms (2024 Edition) reflect data collected through December 2023.
Diagnostic Sensitivity and Specificity
In diagnostic utility studies, Kalayah demonstrated sensitivity of 91.3% and specificity of 88.7% for identifying children meeting DSM-5 criteria for Global Developmental Delay (GDD), as confirmed by multidisciplinary team diagnosis (n = 194). For specific delays, sensitivity was 86.2% for expressive language delay (per PL-90 cutoff), 89.5% for fine motor delay (per Peabody Developmental Motor Scales–2 cutoff), and 83.1% for social-emotional delay (per Ages & Stages Questionnaires: Social-Emotional, Second Edition [ASQ:SE-2] clinical cutoff). False positive rates were lowest among children aged 36–48 months (5.2%) and highest among those aged 12–18 months (11.8%), consistent with known variability in early neurobehavioral expression.
Administration Protocol and Training Requirements
Kalayah requires strict adherence to administration guidelines to ensure measurement fidelity. The full assessment takes 22–35 minutes depending on child engagement and age. Administrators must complete a mandatory 16-hour certification program offered exclusively by ELIL-accredited trainers—including 6 hours of live simulation, 4 hours of video-based scoring calibration, and 6 hours of supervised field practice. Certification is valid for two years and requires renewal via a 4-hour online refresher and submission of two scored video administrations reviewed by a Kalayah Master Trainer. Only certified personnel may administer or interpret results—unlike the Ages & Stages Questionnaires (ASQ-3), which permits parent-completion without clinician oversight.
Required Materials and Equipment
Kalayah uses a standardized kit containing 12 precisely calibrated materials, all manufactured to ASTM F963-17 safety standards. These include:
- Wooden stacking rings (diameter: 5.0 cm ± 0.1 cm; weight: 28 g ± 1 g)
- Plastic pegboard (20 × 20 cm, 25-hole grid, 1.2 cm diameter holes)
- Set of six geometric shape cards (square, circle, triangle, rectangle, star, hexagon; printed on 300 gsm matte cardstock)
- Sound-emitting toy train (peak decibel output: 62 dB at 30 cm, per ANSI S3.19-1993)
- Three textured fabric swatches (corduroy, burlap, velvet; 10 × 10 cm each)
- Standardized 24-page picture book (“The Garden Walk”) with controlled lexical density (Flesch-Kincaid Grade Level 0.3) and image-to-text alignment verified by eye-tracking studies
All materials are sourced from certified vendors: stacking rings from Learning Resources (product #LER0123), pegboard from Fat Brain Toys (model FB-PEG20), and sound train from Fisher-Price (model FP-TRN789). Kit replacement is required every 24 months due to wear-related calibration drift—ELIL mandates annual verification using its Digital Calibration Checker app (v3.2.1), which confirms tactile resistance, acoustic output, and visual contrast ratios within tolerance thresholds.
Scoring Methodology and Interpretive Framework
Kalayah generates four primary metrics: (1) Domain Standard Scores (mean = 100, SD = 15); (2) Developmental Quotient (DQ), calculated as (Sum of Age-Equivalents Across Domains ÷ Number of Domains) ÷ Chronological Age × 100; (3) Discrepancy Analysis Index (DAI), quantifying interdomain variation (e.g., a DAI ≥ 22 points indicates statistically significant divergence between cognitive and social-emotional scores at p < 0.05); and (4) Growth Trajectory Estimate (GTE), derived from longitudinal modeling of prior Kalayah administrations (if available) predicting 6-month developmental gain in months.
Interpreting Discrepancy Patterns
Clinicians use Kalayah’s discrepancy framework to identify atypical developmental profiles. For instance, a child with a Cognitive Standard Score of 112 but a Social-Emotional-Adaptive Score of 74 (difference = 38 points) meets the DAI threshold for further autism spectrum evaluation per AAP 2023 Clinical Practice Guideline. Similarly, a Gross Motor score 24 points below Fine Motor suggests possible dyspraxia or hypotonia requiring physical therapy referral. ELIL’s 2023 Implementation Manual documents 17 empirically linked discrepancy patterns, including:
- Cognitive–Receptive Language gap ≥ 26 points → auditory processing evaluation recommended
- Fine Motor–Expressive Language gap ≥ 20 points → occupational therapy + speech-language pathology co-treatment indicated
- Social-Emotional–Cognitive gap ≥ 30 points → trauma-informed developmental assessment warranted
- Gross Motor–Fine Motor gap ≥ 28 points → orthopedic or neuromuscular consultation advised
- Expressive–Receptive Language gap ≥ 18 points → differential diagnosis of specific language impairment vs. environmental language deprivation
Comparative Utility Against Established Tools
Kalayah fills a distinct niche among developmental assessments. Unlike the Bayley-4—which requires expensive proprietary equipment ($1,895 starter kit), 45–60 minute administration time, and yields separate cognitive/language/motor composites—Kalayah offers integrated domain scoring, shorter administration, and lower cost ($429 for full kit + annual $149 licensing fee). Compared to the Mullen Scales (list price: $1,199), Kalayah avoids timed subtests that may disadvantage children with attention regulation challenges or cultural-linguistic differences. Relative to the ASQ-3—a widely used parent-report screener—Kalayah provides direct observation data, eliminating caregiver literacy, mental health status, or cultural interpretation biases. In a 2022 multisite study published in Pediatrics, Kalayah identified 27% more children with subtle social-emotional delays than ASQ-3 alone (n = 1,012 dyads), particularly among Spanish-speaking families where ASQ-3 sensitivity dropped to 63% versus Kalayah’s 89%.
| Feature | Kalayah (2024) | Bayley-4 | Mullen Scales (MSEL) | ASQ-3 |
|---|---|---|---|---|
| Age Range | 12–60 months | 1–42 months | Birth–69 months | 1–66 months |
| Administration Time | 22–35 min | 45–60 min | 30–45 min | 10–15 min (parent report) |
| Standardization Sample Size | 2,847 | 1,700 | 2,200 | 15,000+ (combined editions) |
| Cost (Initial Kit) | $429 | $1,895 | $1,199 | $295 (professional kit) |
| Training Requirement | 16-hr certification + biennial renewal | Graduate degree + Bayley-4 workshop | Graduate degree + MSEL certification | None (but best practice recommends 2-hr orientation) |
Implementation in Real-World Settings
Kalayah has been adopted in 32 state Part C Early Intervention programs and 11 Head Start grantees since its 2022 national launch. In New Mexico’s Early Childhood Development Program, Kalayah replaced the Denver II as the primary eligibility assessment tool in 2023, resulting in a 19% increase in identification of children with emerging social-emotional needs before age 3. In Chicago Public Schools’ Preschool for All initiative, Kalayah data informed tiered intervention planning: children scoring <85 on Social-Emotional-Adaptive received universal classroom strategies (e.g., emotion cards, predictable routines); those scoring <70 received targeted small-group instruction using the Second Step Early Learning curriculum; and those scoring <65 were referred for individualized behavior support plans.
Equity Considerations and Linguistic Adaptation
ELIL prioritized equity in Kalayah’s design. The English version underwent differential item functioning (DIF) analysis across racial, linguistic, and socioeconomic subgroups; only two items showed negligible DIF (|R²| < 0.02) and were retained with revised scoring anchors. A Spanish adaptation (Kalayah-Español) was normed on 1,203 bilingual and Spanish-dominant children using dynamic assessment principles—scorers provide up to two linguistically appropriate prompts before scoring “not observed.” Validation data show no significant mean score differences between monolingual English and Spanish-dominant children after controlling for maternal education and neighborhood poverty index (β = −0.41, p = 0.32). A Hmong translation is currently undergoing field testing (target release: Q2 2025).
Data Integration and Reporting
Kalayah integrates with major early childhood data systems via HL7 FHIR API. It connects natively with ETO (Everyday Outcomes), ChildPlus, and the federal IDEA Part C State Performance Plan (SPP) Indicator 7 reporting module. Aggregate reports automatically generate compliance-ready summaries for SPP Indicator 7 (child outcomes), including percentages of children making meaningful progress in each of the three child outcome areas (positive social-emotional skills, acquisition and use of knowledge/skills, use of appropriate behaviors to meet needs). In fiscal year 2023, 27 states submitted Kalayah-derived data to the Office of Special Education Programs (OSEP), representing 11.4% of all Part C eligibility determinations nationally.
Limitations and Ongoing Development
Kalayah is not intended for diagnosing medical conditions (e.g., cerebral palsy, genetic syndromes) nor for use with children who have severe sensory impairments (e.g., bilateral profound hearing loss >90 dB, uncorrected visual acuity <20/200). Its normative sample underrepresents children in rural Appalachia (2.1% of total) and certain Pacific Islander communities (<0.5%), prompting ELIL’s 2024–2026 Rural Equity Initiative. Additionally, while Kalayah assesses joint attention and imitation robustly, it does not include direct measures of executive function (e.g., working memory, inhibitory control) for children under 36 months—a gap addressed in the upcoming Kalayah-Advanced module (beta launch scheduled for August 2025), which adds five EF-aligned tasks validated with fNIRS neuroimaging correlation (r = 0.77 with dorsolateral prefrontal cortex activation).
Another practical constraint involves environmental flexibility: Kalayah requires a quiet, distraction-minimized room (ambient noise ≤45 dB, per ANSI S1.13-2020) and a stable surface (height: 52–56 cm, per ADA guidelines). Home-based administrations necessitate portable sound meters and foldable assessment tables—ELIL partners with Galt Toys to distribute the Kalayah Home Kit, which includes a calibrated decibel meter (Extech 407732, accuracy ±1.5 dB) and adjustable-height table (model GT-HOME52).
Finally, while Kalayah’s growth trajectory estimates show strong predictive validity for kindergarten readiness (AUC = 0.84 for WJ IV Letter-Word Identification at age 5.5), it does not replace school-readiness screening tools like the BRIGANCE Early Childhood Screens III. Rather, it serves as a foundational developmental baseline upon which later academic and behavioral screenings build.
ELIL maintains a publicly accessible Technical Report Repository (https://elil.org/kalayah-tech-reports), updated quarterly with new validation studies, bias analyses, and implementation fidelity audits. As of April 2024, the repository hosts 24 peer-reviewed publications, 7 state-level implementation briefs, and 3 randomized controlled trials evaluating Kalayah-informed intervention impact on child outcomes.
For educators, Kalayah is more than an assessment—it is a shared observational language that aligns home, clinic, and classroom perspectives on a child’s developmental narrative. When administered with fidelity and interpreted within ecological context, it supports timely, precise, and equitable decision-making for the youngest learners.
Its strength lies not in replacing clinical judgment, but in sharpening it—transforming subjective impressions into objective, actionable data points anchored in developmental science and validated across diverse populations.
Practitioners should note that Kalayah is not a curriculum or intervention model. It does not prescribe teaching strategies. Instead, it illuminates where a child stands relative to developmental expectations—so that educators and therapists can select evidence-based interventions with greater precision and monitor progress with greater sensitivity.
Because developmental trajectories are rarely linear, Kalayah’s design embraces variability: its scoring rubrics include nuanced behavioral descriptors for ‘emerging’ responses (e.g., ‘looks toward adult after name is called but does not vocalize or gesture’), ensuring children are not mislabeled as ‘delayed’ when they are simply expressing competence in non-canonical ways.
This attention to behavioral nuance reflects a broader paradigm shift in early childhood assessment—from deficit-focused labeling to strengths-informed profiling. Kalayah’s Social-Emotional-Adaptive domain, for example, includes items measuring resilience indicators (e.g., ‘recovers emotional regulation within 90 seconds after brief frustration’) and relationship-building behaviors (e.g., ‘initiates shared gaze during book reading’)—constructs rarely captured in traditional instruments.
As pediatric practice increasingly emphasizes prevention and early support, tools like Kalayah offer a scalable, rigorous, and human-centered method for detecting developmental variation long before it crystallizes into disability—or before opportunity gaps widen irreversibly.
Its growing adoption signals a maturing field—one that values both statistical rigor and developmental authenticity, both standardization and responsiveness, both measurement and meaning.




