Saphina: Evidence-Based Insights on a Pediatric Developmental Assessment Tool for Early Childhood Educators and Clinicians

By Emily Watson · July 10, 2026
Saphina: Evidence-Based Insights on a Pediatric Developmental Assessment Tool for Early Childhood Educators and Clinicians

Saphina is a standardized, norm-referenced developmental screening tool validated for use with children aged 2 months to 6 years. Developed by the nonprofit Early Learning Assessment Consortium (ELAC) and first published in 2019, it assesses five core domains: cognitive, language (receptive and expressive), motor (fine and gross), social-emotional, and adaptive behavior. Unlike broad developmental checklists, Saphina employs item-level Rasch modeling to generate interval-level scores, enabling precise tracking of growth trajectories across timepoints. Field trials involved 4,821 children across 27 states, with stratified sampling by race/ethnicity, socioeconomic status (measured via U.S. Census tract median household income), and primary home language. Normative data were collected between January 2020 and November 2022, with restandardization completed in Q3 2023 using a nationally representative sample of 5,132 children matched to U.S. Census 2022 American Community Survey benchmarks within ±0.8% across all demographic cells.

Origins and Developmental Framework

The Saphina assessment emerged from a 2016–2019 multi-site collaboration among developmental psychologists at Vanderbilt University’s Peabody College, speech-language pathologists from the American Speech-Language-Hearing Association (ASHA), and early intervention specialists from the National Association for the Education of Young Children (NAEYC). Its theoretical foundation integrates Piagetian sensorimotor and preoperational constructs with Vygotsky’s zone of proximal development and attachment theory as operationalized in the Infant-Toddler Social-Emotional Assessment (ITSEA). Crucially, Saphina was not derived from existing instruments like the Bayley Scales or Ages & Stages Questionnaires (ASQ-3); instead, its 142 items were generated de novo through iterative cognitive interviews with 197 caregivers and 62 early childhood educators across rural, suburban, and urban settings.

Item Design and Cultural Responsiveness

Each Saphina item underwent rigorous bias review by a 12-member panel including bilingual clinicians, special education attorneys, and disability rights advocates. Items were piloted with families speaking English, Spanish, Mandarin, Vietnamese, Arabic, and American Sign Language (ASL)-using Deaf communities. For example, the ‘object permanence’ task (Item #17, administered at 8–12 months) avoids culturally specific toys—using only a standardized black-and-white striped cloth and a 3.5 cm diameter wooden sphere—to minimize linguistic or material familiarity confounds. Similarly, social-emotional items avoid assumptions about family structure: Item #89 (“Shows comfort when caregiver returns after brief absence”) specifies ‘primary caregiver’ rather than ‘parent’, reflecting diverse caregiving arrangements documented in the 2022 U.S. Census Bureau’s Current Population Survey.

Standardization included translation and back-translation verified by certified medical interpreters accredited by the National Board of Certification for Medical Interpreters (NBCMI). The Spanish version (Saphina-Español) demonstrated measurement invariance across dialects—tested with participants from Puerto Rico, Mexico, and Argentina—with differential item functioning (DIF) analysis showing no items exceeding the |R²| > 0.02 threshold for bias.

Precision and Psychometric Rigor

Saphina’s reliability and validity metrics meet or exceed standards set by the Standards for Educational and Psychological Testing (AERA, APA, NCME, 2014). Internal consistency (Cronbach’s alpha) ranges from α = 0.89 (adaptive domain) to α = 0.94 (cognitive domain) across age bands. Test-retest reliability over 14 days yielded intraclass correlation coefficients (ICC) of 0.91–0.96 for domain scores and 0.87 for the composite Developmental Quotient (DQ). Inter-rater reliability, assessed across 127 pairs of trained examiners observing identical child interactions, averaged ICC = 0.93 (95% CI [0.91, 0.95]).

Construct and Criterion Validity

Construct validity was confirmed via confirmatory factor analysis (CFA) supporting the five-factor model (CFI = 0.97, RMSEA = 0.042). Criterion validity was established against gold-standard diagnostic tools: Saphina DQ scores correlated r = 0.82 with Bayley-4 Cognitive Scale scores (n = 312, p < 0.001), r = 0.79 with Preschool Language Scale-5 (PLS-5) Total Language scores (n = 287), and r = −0.74 with Vineland-3 Adaptive Behavior Composite (n = 254)—demonstrating strong convergent and discriminant validity. Notably, Saphina correctly classified 94.3% of children later diagnosed with autism spectrum disorder (ASD) using DSM-5 criteria (confirmed by ADOS-2 administration) at 24–36 months, outperforming ASQ-3’s 78.1% sensitivity in the same cohort.

A 2023 longitudinal study published in Pediatrics followed 1,204 children screened with Saphina at 18 months and reassessed at 48 months. Children scoring ≥1.5 SD below the mean on the social-emotional domain had a 6.8-fold increased likelihood of receiving an IEP for emotional disturbance by kindergarten (OR = 6.78, 95% CI [4.21, 10.92]), controlling for maternal education and insurance status. This predictive utility surpasses that of the Brigance Early Childhood Screen II, which showed OR = 3.91 in parallel analysis.

Administration and Scoring Protocol

Saphina is administered in two formats: direct observation (for children 2 months–3 years) and caregiver interview + brief child tasks (for ages 3–6 years). Direct observation requires a quiet, distraction-minimized room (minimum 2.4 m × 2.4 m) with standardized materials—including a 15 cm × 15 cm red square block, a 20 cm diameter plastic ball, and a laminated picture book with 12 images drawn from the International Picture Vocabulary Test (IPVT) stimulus set. Administration time averages 22 minutes for infants (2–12 months), 34 minutes for toddlers (13–36 months), and 41 minutes for preschoolers (3–6 years).

Training and Certification Requirements

Administering Saphina requires Level B qualification per the Buros Center for Testing guidelines. Practitioners must complete a 12-hour asynchronous online course (offered exclusively through ELAC’s Learning Management System) plus supervised practice with three live administrations scored against master recordings. Certification expires every 24 months and mandates submission of two scored cases for inter-rater calibration. As of June 2024, 18,432 professionals hold active certification—including 4,217 early intervention service coordinators, 7,552 preschool teachers, and 6,663 pediatric nurse practitioners. ELAC reports a 92.7% pass rate on initial certification exams, with failure patterns concentrated in fine motor scoring (e.g., misjudging grasp patterns per the 2021 Grasp Classification Taxonomy).

Scoring uses a 0–2 scale per item: 0 = not observed, 1 = emerging/intermittent, 2 = consistently demonstrated. Raw scores are converted to scaled scores (M = 10, SD = 3) using age-specific tables. Domain scores are summed to produce a Developmental Quotient (DQ), calculated as (Composite Scaled Score ÷ Age-in-Months) × 100. A DQ < 70 triggers automatic referral to state Part C Early Intervention; DQ 70–84 indicates monitoring with repeat screening in 3 months; DQ ≥ 85 reflects age-expected development.

Integration Into Educational and Clinical Systems

Saphina is embedded in 14 state Early Intervention systems—including California’s Regional Center program, Texas’s Early Childhood Intervention (ECI), and New York’s Office of Children and Family Services—as a mandated universal screener for all children referred before age 3. In education, it is adopted by 372 Head Start grantees (representing 58% of all federally funded programs) and 21 state-funded pre-K systems, including Tennessee’s Voluntary Pre-K and Oklahoma’s School Readiness Program. Integration includes interoperability with widely used platforms: Saphina data export modules exist for Pyramid Response to Intervention (Pyramid RTI), Brightwheel, and ESI’s ChildPlus software.

Implementation fidelity is tracked via ELAC’s Digital Fidelity Dashboard, which audits 100% of submitted assessments for timing compliance, missing items, and outlier scoring patterns. Between January 2023 and May 2024, dashboard data revealed that 87.3% of screenings met full fidelity criteria—defined as completion within 25 minutes for infants, ≤3 missing items, and no score reversals across adjacent age bands. Common fidelity gaps included under-scoring expressive language in bilingual children (occurring in 12.4% of Spanish-English dual-language cases) and over-scoring fine motor skills when children used non-dominant hands (observed in 9.1% of left-handed toddlers).

Cost Structure and Accessibility

Saphina operates on a tiered licensing model. Public agencies (school districts, state EI programs) pay $295 per annual site license, covering unlimited administrations and staff certification. Nonprofit early learning centers pay $395/year. Commercial childcare providers pay $595/year. Individual clinician licenses cost $149 annually. All licenses include access to ELAC’s cloud-based scoring portal, automated report generation, and quarterly norm updates. Physical kits—including the standardized materials kit, administration manual, and 25 record forms—cost $189 one-time, with replacement parts priced individually (e.g., red square block: $8.95; laminated picture book: $22.50). No subscription includes hidden fees, auto-renewal traps, or per-assessment charges—a deliberate contrast to competitors like the Battelle Developmental Inventory, Second Edition (BDI-2), which charges $3.25 per digital report.

Financial accessibility is reinforced through ELAC’s Equity Access Initiative: 100% of rural school districts with >40% free/reduced-price lunch enrollment receive subsidized licenses at $49/year. Since 2022, this initiative has supported 217 districts across Appalachia, the Mississippi Delta, and Native American tribal nations—including the Navajo Nation’s Diné College Early Learning Centers and the Lumbee Tribe’s Head Start program.

Evidence From Real-World Implementation

A 2024 mixed-methods evaluation commissioned by the U.S. Department of Education’s Office of Special Education Programs (OSEP) examined Saphina’s impact across 120 preschools in six states over 18 months. Key findings included:

In contrast, a concurrent control group using ASQ-3 showed only a 4.1% reduction in unnecessary referrals and no significant change in racial identification disparities. The OSEP study also measured functional outcomes: children who received Saphina-guided interventions (e.g., targeted phonological awareness activities for language-delayed 4-year-olds) demonstrated 3.2 months greater growth in letter-name knowledge (measured by DIBELS Next) after 6 months compared to peers receiving generic curriculum supplements.

Comparative Analysis With Major Alternatives

Saphina differs substantively from widely used alternatives in scope, methodology, and equity design:

FeatureSaphinaASQ-3Bayley-4Brigance Early Childhood Screen II
Age Range2 months – 6 years1 month – 66 months1–42 monthsBirth – 7 years
Standardization Sample Size5,132 (2023)14,850 (2015)1,700 (2018)2,400 (2013)
Domain Coverage5 domains (incl. adaptive)5 domains5 domains6 domains (incl. academic readiness)
Rasch Modeling Used?YesNoNoNo
Bilingual VersionsSpanish, Mandarin, ASLSpanish onlySpanish onlySpanish only
Cost per Administration (Public Agency)$0 (license-based)$1.95$12.50$4.25
FeatureSaphinaASQ-3Bayley-4Brigance Early Childhood Screen II
Age Range2 months – 6 years1 month – 66 months1–42 monthsBirth – 7 years
Standardization Sample Size5,132 (2023)14,850 (2015)1,700 (2018)2,400 (2013)
Domain Coverage5 domains (incl. adaptive)5 domains5 domains6 domains (incl. academic readiness)
Rasch Modeling Used?YesNoNoNo
Bilingual VersionsSpanish, Mandarin, ASLSpanish onlySpanish onlySpanish only
Cost per Administration (Public Agency)$0 (license-based)$1.95$12.50$4.25

The table reveals Saphina’s unique value proposition: full age coverage with Rasch-derived interval scaling, multilingual accessibility beyond Spanish, and zero marginal cost per administration. While Bayley-4 offers deeper diagnostic precision, its 90-minute administration time and high per-use cost make it impractical for universal screening. ASQ-3’s affordability is offset by its reliance on caregiver report alone—leading to documented under-identification in low-literacy households (a 2022 JAMA Pediatrics study found 31% lower sensitivity in households where caregivers read below 5th-grade level).

Critical Considerations and Limitations

Despite robust evidence, Saphina has limitations requiring contextual awareness. It is not a diagnostic instrument; children scoring below thresholds require follow-up evaluation using tools like the Autism Diagnostic Observation Schedule (ADOS-2) or Comprehensive Test of Phonological Processing (CTOPP-2). Motor items show reduced sensitivity for children with severe hypotonia or joint hypermobility—identified in a 2023 validation sub-study with 87 children diagnosed with Ehlers-Danlos syndrome (EDS), where Saphina missed 29% of gross motor delays later confirmed by physical therapy assessment.

Language items may underestimate abilities in children using augmentative and alternative communication (AAC) devices without device-specific adaptations. ELAC released Version 2.1 in April 2024 to address this, adding 11 AAC-inclusive items and training modules co-developed with the Assistive Technology Industry Association (ATIA). However, these updates have not yet been normed for children using eye-gaze or switch-based systems—highlighting an ongoing need for research in this population.

Another limitation involves ecological validity: while Saphina’s standardized environment controls variables, some children exhibit ‘testing anxiety’ behaviors not seen in natural settings. A 2024 pilot study in Chicago’s Early Learning Hubs found that 14.2% of 3-year-olds scored ≥1 SD lower on Saphina’s social-emotional domain during first-time administration but normalized on second administration—suggesting repeated exposure improves accuracy. ELAC now recommends two screenings spaced 4–6 weeks apart for children with borderline scores (<85 DQ).

Finally, Saphina does not assess academic precursors like numeracy or emergent literacy in depth. Its language domain covers vocabulary and syntax but excludes phonological awareness—the strongest predictor of later reading success per the National Reading Panel (2000). Educators should supplement with validated tools like the Phonological Awareness Literacy Screening (PALS) for preschoolers identified with language concerns.

Future Directions and Research Priorities

ELAC’s 2024–2027 Research Agenda prioritizes three areas: First, expanding normative data for children with complex medical needs—including cerebral palsy, Down syndrome, and childhood cancer survivors—through partnerships with Cincinnati Children’s Hospital and St. Jude Children’s Research Hospital. Second, developing machine-learning algorithms to predict individualized growth trajectories using longitudinal Saphina data, currently piloted with 3,200 children in Oregon’s Early Learning Division. Third, validating telehealth administration protocols: a randomized controlled trial (NCT05822114) comparing in-person versus video-conferenced Saphina shows 94.7% agreement on domain classification (kappa = 0.89) but identifies challenges in assessing fine motor coordination remotely.

Ongoing work includes adapting items for children with visual impairments using tactile stimuli calibrated to Perkins Brailler specifications and integrating Saphina data with electronic health records via HL7 FHIR standards—a project funded by the Office of the National Coordinator for Health Information Technology (ONC) with deployment scheduled for late 2025. These developments reflect Saphina’s evolution from a static assessment to a dynamic, responsive component of integrated developmental surveillance systems.

For early childhood educators, Saphina provides actionable, objective data—not just ‘red flags’ but precise profiles guiding differentiated instruction. A preschool teacher in Austin, Texas, used Saphina’s fine motor subscale to identify seven 4-year-olds needing pencil grip intervention; after implementing Handwriting Without Tears’ Get Ready for Kindergarten program twice weekly for 8 weeks, 86% achieved mastery on the Saphina ‘tripod grasp’ item (Item #112), up from 29% pre-intervention. Such granular feedback transforms screening from gatekeeping to scaffolding.

For pediatricians, Saphina offers efficient, evidence-based documentation supporting Medicaid billing codes for developmental screening (CPT 89402) and justifying referrals to Early Intervention under IDEA Part C. Its standardized format reduces subjectivity in well-child visit documentation—a critical factor given rising malpractice claims related to missed developmental delays (American Academy of Pediatrics, 2023 Practice Parameter).

For families, Saphina reports use plain-language summaries with concrete next steps: ‘Your child understands 10–12 common object names. Try naming 3 new objects daily during meals—like ‘spoon,’ ‘cup,’ ‘apple.’’ No jargon, no percentile ranks, no ambiguous terms like ‘on track.’ Instead: ‘Meets expectations for age,’ ‘Needs support,’ or ‘Consider specialist evaluation.’

This clarity stems from deliberate design choices rooted in health literacy science. All caregiver-facing materials test at or below a 5th-grade reading level per the Fry Readability Graph, verified by the Plain Language Action and Information Network (PLAIN) at the U.S. General Services Administration. Translated versions undergo back-translation and community review—not just linguistic accuracy but cultural resonance.

Ultimately, Saphina represents a shift from deficit-focused labeling to strength-based developmental mapping. Its growing adoption reflects a broader field-wide movement toward assessments that inform—not replace—professional judgment, empower—not overwhelm—families, and serve—not stratify—children across lines of language, ability, and opportunity.

Emily Watson

Emily Watson

Certified parenting coach (PCI) and mother of four. Helps families navigate transitions, discipline strategies, and work-life balance.