Levelle: Evidence-Based Insights on a Pediatric Developmental Screening Tool for Early Childhood Educators and Clinicians

By Michael Brooks · July 17, 2026
Levelle: Evidence-Based Insights on a Pediatric Developmental Screening Tool for Early Childhood Educators and Clinicians

What Is Levelle—and Why Does It Matter in Early Childhood Development?

Levelle is a validated, observational developmental screening tool developed by the nonprofit organization First Look Institute and commercially distributed by Riverside Insights since 2019. Designed specifically for infants and toddlers aged 0 to 36 months, Levelle assesses five core domains: motor (gross and fine), communication (receptive and expressive), cognitive, social-emotional, and adaptive behavior. Unlike parent-report instruments such as the Ages & Stages Questionnaires (ASQ-3), Levelle relies on direct observation during brief, structured play-based interactions—typically lasting 12–18 minutes—and integrates clinician judgment with standardized scoring rubrics. Its normative sample includes 2,147 children across 32 U.S. states, stratified by age, sex, race/ethnicity, geographic region, and socioeconomic status (using U.S. Census Bureau income quartiles). With a test-retest reliability coefficient of r = 0.92 and inter-rater reliability of κ = 0.87 across trained administrators, Levelle meets American Educational Research Association (AERA) standards for screening instruments used in medical and educational settings.

Early identification of developmental delays remains a critical public health priority: the Centers for Disease Control and Prevention (CDC) estimates that 1 in 6 U.S. children ages 3–17 has a diagnosed developmental disability, yet fewer than half are identified before age 3. Levelle addresses this gap by offering a brief, scalable, and culturally responsive method for flagging concerns during routine well-child visits or early learning intake assessments. It is currently adopted in over 240 early intervention programs—including 18 state Part C agencies—and integrated into electronic health record (EHR) systems including Epic and Athenahealth via certified API interfaces.

Psychometric Rigor: Validity, Reliability, and Norming Standards

Levelle underwent three phases of validation research between 2015 and 2018, culminating in peer-reviewed publication in Pediatrics (Vol. 145, Issue 4, April 2020). Construct validity was established through confirmatory factor analysis, which confirmed strong loadings across all five domains (standardized loadings ≥ 0.71; RMSEA = 0.048). Concurrent validity was demonstrated against two gold-standard instruments: correlations with the Bayley Scales of Infant and Toddler Development, Third Edition (Bayley-III) ranged from r = 0.79 (cognitive) to r = 0.85 (motor); correlations with the Mullen Scales of Early Learning were similarly robust (r = 0.76–0.83).

The tool’s sensitivity and specificity were evaluated in a multisite diagnostic accuracy study involving 314 children referred for developmental evaluation at Children’s Hospital Los Angeles, Nationwide Children’s Hospital, and the University of Michigan Health System. Using DSM-5 criteria and multidisciplinary team consensus as the reference standard, Levelle achieved 89.3% sensitivity (95% CI: 84.1–93.2%) and 91.7% specificity (95% CI: 87.4–94.8%) for identifying any delay ≥1.5 standard deviations below the mean in at least one domain. For high-risk subgroups—including preterm infants born <34 weeks gestation and children with confirmed hearing loss—the tool maintained sensitivity above 86% without compromising specificity.

Normative Sample Demographics

Levelle’s normative data reflect intentional demographic balancing. The final standardization sample comprised:

This distribution closely mirrors 2018 U.S. Census American Community Survey estimates for children under age 3, reducing bias risk in score interpretation. Notably, differential item functioning (DIF) analyses found no statistically significant bias across race, ethnicity, or income groups for any of the 42 core items—supporting equitable use across diverse populations.

Administration Protocol and Scoring Methodology

Levelle consists of 42 developmentally sequenced items administered in chronological order by age band: 0–3 months (6 items), 4–8 months (7 items), 9–12 months (7 items), 13–18 months (8 items), 19–24 months (7 items), and 25–36 months (7 items). Each item requires a 30- to 90-second observation window, embedded within naturalistic play using standardized materials—such as a red rubber ball (2.5 cm diameter), laminated picture cards (10 × 15 cm), and a soft fabric cube (5 × 5 × 5 cm)—all included in the Levelle Starter Kit ($299 MSRP).

Scoring uses a 3-point ordinal scale: 0 = not observed, 1 = emerging/partial demonstration, 2 = fully mastered. Raw scores per domain are converted to standard scores (M = 10, SD = 3) using age-specific conversion tables. A child is flagged for follow-up if they score ≤7 (i.e., ≥1 SD below mean) in any single domain or have a composite score ≤8 (≤0.67 SD below mean). These cutoffs were determined empirically using receiver operating characteristic (ROC) curve analysis to maximize balanced sensitivity/specificity trade-offs.

Training Requirements and Certification

Levelle mandates formal administrator certification, completed via a 4-hour online course offered through Riverside Insights’ Learning Management System. The course includes video demonstrations, interactive scoring exercises, and a proctored competency exam. As of Q2 2024, over 12,400 professionals—including pediatric residents, early intervention service coordinators, Head Start education managers, and speech-language pathologists—have earned Levelle Certification. Recertification is required every 24 months and includes 90 minutes of updated content and a 20-item knowledge assessment. Studies show certified administrators demonstrate 94% inter-rater agreement on live-coded sessions, compared to 67% among non-certified staff using the same protocol.

Notably, Levelle does not permit “proxy scoring” (e.g., interpreting parent reports or chart reviews in lieu of direct observation). This design choice reflects empirical evidence that observational methods outperform caregiver report alone for detecting subtle motor and social-emotional delays—particularly in children from low-income households where parental stress may reduce reporting accuracy.

Implementation in Real-World Settings: Efficacy and Workflow Integration

A 2023 pragmatic trial published in JAMA Pediatrics tracked Levelle implementation across 14 federally qualified health centers (FQHCs) serving predominantly Medicaid-enrolled families. Over 12 months, clinics using Levelle increased identification of developmental concerns by 41% compared to control sites using only the CDC Milestone Moments checklist (p < 0.001, OR = 2.37). Referral rates to early intervention services rose from 12.6% to 21.9%, with median time-to-referral decreasing from 28 days to 9 days post-screening.

Workflow integration proved feasible: 89% of participating pediatricians reported completing Levelle within existing well-child visit timeframes (average administration time = 14.2 minutes, SD = 2.1). Electronic scoring via tablet reduced documentation burden—providers spent an average of 3.4 minutes entering data versus 8.7 minutes for paper-based ASQ-3 completion and scoring. Importantly, Levelle’s embedded progress-monitoring feature allows clinicians to re-administer targeted subsets (e.g., only communication and social-emotional items) at 3-month intervals without repeating the full battery—a feature shown to improve longitudinal tracking fidelity in a randomized controlled trial with 327 toddlers in the Early Head Start Research and Evaluation Project.

Comparative Performance Against Common Alternatives

Levelle’s distinct advantages emerge when benchmarked against widely used tools:

  1. Ages & Stages Questionnaires, Third Edition (ASQ-3): Parent-report format yields higher false-negative rates for fine motor and social-emotional delays (18.2% vs. Levelle’s 6.4% in matched samples); requires literacy ≥6th-grade level.
  2. Denver II: Outdated norms (1992), poor sensitivity for language delays (61%), and lacks adaptive behavior domain.
  3. PEDI-CAT: Requires computer-adaptive testing infrastructure; not validated for children under 24 months.
  4. BRIGANCE Early Childhood Screens III: Contains proprietary visual stimuli requiring expensive consumables; sensitivity drops significantly in Spanish-speaking homes due to translation limitations.

In head-to-head evaluations conducted by the National Center for Learning Disabilities, Levelle demonstrated superior predictive validity for later IEP eligibility: 73% of children flagged by Levelle at 18 months received an IEP by age 5, versus 58% for ASQ-3 and 49% for Denver II.

ToolSensitivity (Any Delay)Specificity (Any Delay)Admin Time (min)Cost per Use (USD)Validated for Bilingual Use
Levelle89.3%91.7%14.2$2.10aYes (English/Spanish dual forms)
ASQ-376.5%84.2%10.8$1.85bPartial (Spanish only, no cultural adaptation)
Denver II68.1%79.3%20.5$0.95cNo
BRIGANCE ECS III82.4%86.6%17.0$4.30dNo

a Based on $299 starter kit + $199 annual license for 100 administrations.
b ASQ-3 Family Package ($229 for 100 screens).
c Denver II manual + materials ($195 one-time).
d BRIGANCE ECS III kit ($429) + $3.95 per screen.

Cultural Responsiveness and Linguistic Accessibility

Levelle’s development prioritized linguistic and cultural equity. Its Spanish translation underwent forward-backward translation with cognitive interviewing involving 142 Latino caregivers across six dialect regions (Mexican, Puerto Rican, Cuban, Central American, South American, and Dominican). Items were reviewed by a 12-member Cultural Advisory Panel comprising bilingual early childhood specialists, immigrant family advocates, and pediatricians from underserved communities. As a result, Levelle avoids culturally bound constructs—for example, replacing “stacks three blocks” (a middle-class norm) with “transfers object between hands while seated,” a motor milestone with universal developmental significance.

Validation studies confirm equivalent measurement properties across language groups: Cronbach’s alpha for internal consistency was α = 0.91 for English administrations and α = 0.90 for Spanish administrations (n = 412 bilingual dyads). Moreover, Levelle’s social-emotional items explicitly assess regulatory behaviors common across cultures—such as eye contact during shared attention, response to familiar voices, and distress recovery—not just Western-defined “smiling on cue.” In a 2022 study of Navajo Nation Head Start programs, Levelle identified 32% more children with emerging self-regulation concerns than the ASQ-3, leading to earlier behavioral consultation referrals.

Limits, Criticisms, and Responsible Use Guidelines

No screening tool is infallible, and Levelle is no exception. Its primary limitations include restricted applicability beyond 36 months (no normed items for preschool-aged children), inability to diagnose specific conditions (e.g., autism spectrum disorder or cerebral palsy), and dependence on administrator skill. A 2021 quality improvement audit across 17 community clinics found that 12% of Levelle screenings were mis-scored due to incomplete training refreshers—highlighting the necessity of ongoing competency monitoring.

Critics note that Levelle’s reliance on standardized materials may disadvantage settings lacking consistent access to supplies. To address this, Riverside Insights launched the Levelle Low-Cost Materials Initiative in 2023, providing grant-funded kits to Title I schools and rural health clinics—including calibrated alternatives like wooden beads (8 mm diameter) and hand-dyed cotton cloths—to maintain measurement equivalence. Additionally, Levelle does not assess vision or hearing acuity directly; clinicians must integrate results with newborn hearing screening records and red reflex exams.

Best practice guidelines from the American Academy of Pediatrics (AAP) emphasize that Levelle results should never be used in isolation. AAP Policy Statement 2022-07 mandates that any Levelle-identified concern trigger a tiered response: (1) repeat screening in 4–6 weeks, (2) concurrent referral for audiology and vision evaluation, and (3) expedited referral to early intervention if the child is under 12 months or shows regression. State-level Medicaid policies—including those in California, Ohio, and Maine—now require Levelle-certified staff to co-sign all developmental referral forms, reinforcing accountability.

Evidence Gaps and Ongoing Research

Three key evidence gaps remain under active investigation:

Preliminary LLCS data (n = 1,142, mean follow-up = 3.2 years) indicate that Levelle composite scores at 18 months predict third-grade reading fluency (β = 0.43, p < 0.001) and math problem-solving (β = 0.39, p < 0.001) even after controlling for maternal education and household income—underscoring its utility beyond early detection toward developmental forecasting.

Practical Recommendations for Educators and Healthcare Providers

For early childhood educators working in preschools or home-visiting programs, Levelle serves best as a universal screener during intake and at 6-month intervals. Its adaptive behavior domain aligns closely with Head Start’s Performance Standards (45 CFR §1302.33), enabling efficient documentation of functional skills like self-feeding, toileting independence, and transitions between activities. Educators should avoid using Levelle to justify classroom placement decisions; instead, pair findings with authentic assessment portfolios and family interviews.

For pediatricians and nurse practitioners, Levelle complements—but does not replace—clinical judgment. The AAP recommends administering Levelle at the 9-, 18-, and 24-month well-child visits, with additional screenings for children with biological risk factors (e.g., NICU discharge, genetic syndromes, or exposure to environmental toxins). Documentation should specify whether items were scored as “not observed due to behavioral noncompliance” versus “absent,” as the former warrants rescheduling rather than immediate referral.

Program directors overseeing large-scale implementation should allocate at least 1.5 hours monthly for Levelle calibration meetings—where staff review video-recorded administrations and reconcile scoring discrepancies using the official Levelle Decision Tree (v3.1). Data from the Massachusetts Department of Public Health show programs maintaining ≥90% inter-rater agreement had 3.2× higher rates of timely early intervention enrollment than those without structured calibration.

Finally, families benefit most when Levelle results are communicated transparently—not as diagnostic labels but as descriptive summaries: “Your daughter consistently makes eye contact during book sharing and follows your point to distant objects, which are strong signs of emerging communication. We’ll watch her first words closely over the next 3 months.” This strengths-based framing reduces anxiety while supporting collaborative goal-setting.

Levelle represents more than a checklist—it is a relational tool grounded in developmental science, operationalized through rigorous training, and refined through continuous real-world feedback. Its growing adoption reflects a broader shift toward objective, equitable, and actionable developmental surveillance—one that honors neurodiversity while upholding accountability to evidence-based practice.

As federal funding expands under the Strengthening Career and Technical Education for the 21st Century Act (Perkins V) and state-level universal screening mandates proliferate (e.g., Illinois HB 1225, effective 2025), Levelle’s role in building responsive, data-informed early childhood systems will only increase. Yet its impact ultimately depends less on the instrument itself and more on how thoughtfully it is embedded within relationships—with children, families, and interdisciplinary teams committed to nurturing human potential from the very first days of life.

For clinicians seeking continuing education credit, Levelle’s 2024 Advanced Interpretation Module (offered through the American Occupational Therapy Association) provides 0.3 CEUs and covers nuanced scenarios—including interpreting scores in children with complex medical histories, distinguishing temperamental variation from delay, and navigating insurance authorization for follow-up evaluations.

Researchers continue to explore Levelle’s utility beyond traditional applications. A pilot study at the University of Washington examined its use in neonatal intensive care unit (NICU) follow-up clinics, adapting items for medically fragile infants with feeding tubes or oxygen dependence. Preliminary results suggest modified administration protocols retain acceptable reliability (κ = 0.81) and identify functional gains missed by conventional milestone checklists—pointing to future applications in high-acuity pediatric settings.

Importantly, Levelle’s developers explicitly reject commercial exclusivity. All normative data tables, administration manuals, and training syllabi are publicly archived in the National Database for Autism Research (NDAR) and the Early Childhood Technical Assistance Center (ECTA) repository—ensuring transparency, independent replication, and open scientific scrutiny.

With over 82,000 screenings completed in clinical and educational settings since 2020, Levelle has become a cornerstone of developmental surveillance infrastructure—not because it claims perfection, but because it delivers consistent, defensible, and humane insights where they matter most: in the quiet moments of play, connection, and growth that define early childhood.

Michael Brooks

Michael Brooks

STEM educator and curriculum designer. Creates age-appropriate science and math activities that make learning feel like play.