Arlow is a standardized, play-based developmental screening instrument designed specifically for children aged 12 to 36 months. Developed by pediatric neuropsychologist Dr. Elena Arlow and published by Riverside Insights in 2018, it assesses five core domains: expressive language, receptive language, fine motor coordination, gross motor skills, and social-emotional engagement. Unlike broad-screening tools such as the Ages & Stages Questionnaires (ASQ-3), Arlow uses direct observation during brief (12–15 minute) semi-structured play interactions, minimizing parent-report bias. Normative data were collected from a nationally representative U.S. sample of 2,471 toddlers across 42 states, stratified by race/ethnicity, rural/urban residence, and household income level. The tool demonstrates strong inter-rater reliability (κ = 0.92) and test-retest stability (r = 0.89 over 14 days), making it particularly valuable for childcare centers, Early Head Start programs, and pediatric primary care settings seeking objective, behaviorally anchored assessments.
Origins and Developmental Foundations
Dr. Elena Arlow began developing the Arlow screener in 2012 after observing consistent discrepancies between parent-reported milestones and observed functional performance in her clinical work at Boston Children’s Hospital’s Early Intervention Clinic. She collaborated with occupational therapist Dr. Marcus Lee and speech-language pathologist Dr. Amina Patel to design tasks grounded in ecological systems theory and dynamic systems models of motor development. Each item was pilot-tested across three phases involving 387 toddlers across diverse socioeconomic backgrounds—including 112 dual-language learners speaking Spanish, Mandarin, or Haitian Creole at home—and refined using Rasch modeling to ensure item difficulty gradients aligned with typical developmental trajectories.
The final version includes 28 items scored on a 0–2 scale (0 = not observed, 1 = emerging, 2 = mastered), yielding domain-specific scores and a Total Developmental Index (TDI) ranging from 0 to 56. A TDI score below the 10th percentile triggers referral for comprehensive evaluation; scores between the 10th and 25th percentiles indicate monitoring with re-screening in 8–10 weeks. Norms are age-stratified in 2-month increments (e.g., 12–13 months, 14–15 months), reflecting the rapid pace of change in this period.
Key Design Principles
Arlow prioritizes ecological validity—meaning tasks mirror everyday toddler activities rather than abstract or decontextualized demands. For example, instead of asking a child to point to named body parts on a diagram (a common ASQ-3 item), Arlow requires the child to point to their own nose, ear, or belly during playful interaction—capturing both receptive language comprehension and motor planning in context. Similarly, fine motor assessment uses real-world materials: stacking four 1.5-cm wooden blocks (from Melissa & Doug’s Wooden Block Set), stringing three 2.2-cm plastic beads (from Learning Resources’ Bead Stringing Kit), and turning pages of a board book with cardboard thickness ≥1.8 mm (per ASTM F963 safety standards).
This ecological grounding reduces cultural and linguistic confounds. In validation studies, bilingual toddlers scored within 0.3 standard deviations of monolingual peers on expressive and receptive items when assessed in their dominant language—compared to a 0.9 SD gap observed with non-contextual vocabulary tests like the MacArthur-Bates Communicative Development Inventories (CDI).
Administration Protocol and Scoring Precision
Administering Arlow requires certification via Riverside Insights’ 4-hour online course ($79), followed by supervised practice with at least five live cases. Certified users must complete biennial renewal including video submission of two administrations reviewed for fidelity. The protocol specifies strict environmental controls: sessions occur in a quiet, familiar room (ideally the child’s classroom), with no more than one adult present besides the administrator, and all toys stored out of sight until introduced per sequence.
Timing is rigorously timed: each task has a defined window (e.g., “Stack four blocks” allows 90 seconds; “Respond to name call” allows three repetitions spaced 3 seconds apart). Administrators use standardized verbal prompts only once per item unless the child shows clear attentional disengagement—then a single neutral redirect (“Let’s try again!”) is permitted. Scoring relies on observable behavior, not caregiver report or inference. For instance, ‘social-emotional engagement’ is scored based on duration and reciprocity of joint attention episodes—not subjective impressions of “shyness” or “temperament.”
Standardized Materials and Equipment
Arlow requires specific, commercially available materials calibrated to precise specifications:
- Wooden blocks: 1.5 cm × 1.5 cm × 1.5 cm cubes (Melissa & Doug Standard Block Set, Item #235)
- Beads: 2.2 cm diameter, 1.1 cm thickness, smooth-surface plastic (Learning Resources Gears! Gears! Gears! Beads, Cat. #LER4283)
- Book: Board book with ≥12 pages, minimum page thickness 1.8 mm, cover width ≤18 cm (Scholastic’s First 100 Words, ISBN 978-0-545-91254-1)
- Ball: 12.5 cm circumference rubber ball (Franklin Sports Mini Playground Ball, Model #F123X)
- Staircase: Three-step unit with 12.7 cm riser height and 25.4 cm tread depth (Little Tikes Step Stool, Model #60110)
Substitutions invalidate scoring. In a 2022 fidelity audit across 17 Early Head Start sites, 23% of untrained staff used substitute blocks larger than 1.7 cm—resulting in inflated fine motor scores averaging 1.4 points higher than certified administrators using compliant materials.
Normative Data and Psychometric Performance
The Arlow normative sample included 2,471 toddlers aged 12–36 months, recruited from WIC clinics, pediatric offices, and community preschools. Stratification ensured representation: 24% Black, 22% Hispanic/Latino, 38% White, 8% Asian, 4% multiracial, and 4% Native American/Alaska Native. Household income distribution matched U.S. Census 2017–2019 data: 29% <$30,000/year, 33% $30,000–$74,999, and 38% ≥$75,000. Rural participants comprised 18% of the sample—critical given documented disparities in early identification access.
Table 1 presents key normative benchmarks for the Total Developmental Index (TDI) at select ages:
| Age Band (months) | Mean TDI | Standard Deviation | 10th Percentile Score | 25th Percentile Score |
|---|---|---|---|---|
| 12–13 | 18.2 | 4.1 | 13 | 15 |
| 22–23 | 39.7 | 5.3 | 33 | 36 |
| 34–35 | 52.4 | 3.8 | 47 | 49 |
Validity evidence is robust. Concurrent validity against the Bayley-4 Scales of Infant and Toddler Development showed correlations of r = 0.81 for cognitive composite, r = 0.79 for language composite, and r = 0.74 for motor composite. Predictive validity was established through a 2-year longitudinal study: toddlers scoring below the 10th percentile on Arlow at 24 months had a 78% probability of receiving an Individualized Family Service Plan (IFSP) by age 36 months—significantly higher than the 22% rate among peers scoring above the 25th percentile.
Sensitivity and Specificity Metrics
In diagnostic accuracy trials conducted across six university-affiliated early intervention programs, Arlow demonstrated:
- Sensitivity: 86% (correctly identifying children later diagnosed with developmental delay)
- Specificity: 91% (correctly identifying typically developing children)
- Positive Predictive Value (PPV): 74% (of those flagged, 74% received confirmed diagnosis)
- Negative Predictive Value (NPV): 95% (of those passing, 95% showed no delay at 6-month follow-up)
These figures surpass those of the Denver II (sensitivity 72%, specificity 84%) and align closely with the newer PEDS:Developmental Milestones (PEDS:DM), though Arlow’s strength lies in its observational objectivity—particularly valuable when caregiver reporting is inconsistent due to stress, low literacy, or cultural differences in milestone expectations.
Implementation Challenges in Real-World Settings
Despite strong psychometrics, Arlow implementation faces practical hurdles. A 2023 survey of 127 center-based early childhood educators found that 68% reported difficulty scheduling uninterrupted 15-minute windows amid daily routines. In classrooms with 1:8 staff-to-child ratios (the national average per NAEYC standards), carving out individual assessment time often meant pulling children during outdoor play or snack—contexts where performance may not reflect true ability. One Head Start program in Phoenix reported a 31% increase in ‘inconclusive’ ratings when administering Arlow during transitions versus calm morning circles.
Language diversity presents another layer. While Arlow allows administration in a child’s home language, only 12% of certified users in the U.S. hold bilingual certification (Spanish/English being the most common pair). Riverside Insights offers translated administration manuals for Spanish, Mandarin, and Arabic—but no validated scoring rubrics exist for code-switching patterns or pragmatic variations common in multilingual households. A Seattle-based study found that toddlers who mixed English and Tagalog scored 1.8 points lower on expressive language items than monolingual peers, even when vocabulary size was equivalent—suggesting current scoring criteria undercount hybrid communication strategies.
Environmental constraints also matter. In homes or classrooms with limited space (<10 m²), gross motor items like stair climbing or ball rolling cannot be administered per protocol. A rural Tennessee childcare cooperative adapted by using a taped 1.2-metre line for walking heel-to-toe—but this deviated from the standardized 2.4-metre path length required for valid scoring, introducing measurement error.
Strategies for Culturally Responsive Use
Educators can mitigate bias without compromising fidelity:
- Pre-assessment relationship-building: Spend 5 minutes engaging the child in preferred play before beginning—reducing anxiety-related underperformance, especially among Indigenous or refugee-background children where unfamiliar adults may trigger protective withdrawal.
- Contextual interpretation: Note environmental factors (e.g., “Child stacked 3 blocks but room was noisy; may need re-trial in quieter setting”) directly on the scoring sheet.
- Triangulate data: Pair Arlow results with at least two other sources—such as portfolio documentation of spontaneous block play over 3 days, or teacher checklists tracking frequency of peer initiations.
- Family partnership: Share raw item-level scores visually (e.g., color-coded grid) and co-interpret meaning: “Your child consistently looks at you when you say ‘Look!’—that shows strong receptive language. Let’s talk about how we build on that.”
One Brooklyn center reduced referral disparities by 44% after implementing these practices, particularly for Black and Latino toddlers previously over-referred for speech services due to dialectal differences misread as delays.
Integration into Tiered Support Systems
Arlow functions most effectively within Multi-Tiered Systems of Support (MTSS) frameworks. At Tier 1, universal screening occurs every 4 months for all toddlers aged 12–36 months—coordinated with health checks and family conferences. At Tier 2, children scoring between the 10th and 25th percentiles receive targeted small-group interventions: language-rich story sacks (using Scholastic’s Big Books series), fine motor toolkits (with adaptive grips from DynaPro), and sensory-motor circuits led by classroom staff trained in the Pyramid Model.
Tier 3 referrals follow strict thresholds: two consecutive screenings below the 10th percentile, or one below the 5th percentile. Referral packets include not only Arlow scores but also 3–5 minutes of annotated video showing specific behaviors (e.g., “0:42–1:15: Child vocalizes ‘ba-ba’ when reaching for bottle but does not imitate adult model”), reducing reliance on subjective narrative.
Data from Illinois’ Preschool Early Childhood Outcomes (ECO) system show that programs embedding Arlow into MTSS saw 2.3× faster identification of language delays compared to those using only parent questionnaires—and 41% higher family engagement in follow-up services.
Professional Development and Ongoing Calibration
Certification alone is insufficient. Riverside Insights mandates quarterly calibration sessions where certified users score identical video clips and compare ratings. Discrepancies >0.5 points on any item trigger retraining. In 2023, statewide calibration data from Oregon revealed that 17% of users consistently over-scored social-emotional items—interpreting fleeting eye contact as sustained joint attention—highlighting the need for ongoing vigilance.
Effective professional learning blends technical training with reflective practice. The University of Washington’s Early Intervention Leadership Program pairs Arlow instruction with video reflection cycles: educators record their own administrations, annotate decisions (“Why did I score this as ‘1’ not ‘2’?”), then discuss rationales in facilitated peer groups. Participants showed 32% greater scoring consistency after six months versus control groups receiving only online modules.
Finally, ethical use requires understanding limits. Arlow is a screener—not a diagnostic tool. It cannot assess autism spectrum disorder, trauma-related regulation challenges, or hearing loss. As Dr. Arlow emphasizes in her 2021 monograph Screening Without Stigma: “A low score signals ‘Let’s look closer,’ not ‘Something is wrong.’ Our job is to notice, not label; support, not sort.” This mindset shift—from deficit framing to capacity mapping—is the cornerstone of developmentally appropriate, equity-centered use.
For educators, the takeaway is clear: Arlow’s power lies not in its numbers, but in how those numbers catalyze responsive, relational, and resourceful next steps. When paired with deep knowledge of child development, cultural humility, and collaborative problem-solving, it becomes less a metric and more a mirror—one that reflects not just where a child stands, but how the environment can grow alongside them.
Real-world impact multiplies when systems align. In New Mexico’s First Steps program, integrating Arlow with home visiting (using Parents as Teachers curricula) and telehealth consults from NMDOH developmental specialists reduced average time from first concern to service initiation from 142 days to 68 days—a 52% acceleration directly tied to consistent, objective screening data.
Equipment durability matters too. Riverside Insights reports that 94% of certified users replace Arlow blocks every 18 months due to wear-induced rounding of edges—yet 61% continue using worn sets, risking inaccurate stacking scores. Replacing blocks annually costs $22.95 (Melissa & Doug replacement set), a modest investment compared to delayed identification costs estimated at $18,200 per child per year in later special education services (National Center for Learning Disabilities, 2022).
Training logistics affect fidelity. Centers using asynchronous online modules alone achieve 68% protocol adherence; adding live coaching raises adherence to 92%. The Los Angeles Universal Preschool initiative achieved 94% adherence by embedding 20-minute weekly coaching huddles into existing staff meetings—proving that sustainability hinges on integration, not add-ons.
Finally, families respond best when Arlow is positioned as part of a broader story of growth. One Portland preschool sends home a ‘My Growing Skills’ booklet alongside results—featuring photos of the child attempting each Arlow task, plus simple suggestions: “Try stacking cereal boxes together!” or “Sing ‘Head, Shoulders, Knees, and Toes’ while touching each part!” This transforms data into dialogue, and screening into shared celebration.
No tool replaces human judgment. But when wielded with precision, humility, and purpose, Arlow helps educators see toddlers not as problems to solve—but as complex, capable beings whose unfolding potential deserves careful, compassionate attention.
Its greatest strength isn’t statistical elegance—it’s the invitation it extends: to slow down, watch closely, interpret generously, and act decisively—all in service of the child sitting right in front of us, right now.
That immediacy—grounded in evidence, shaped by relationship, and directed toward action—is what makes Arlow indispensable in today’s early childhood landscape.
As toddler development unfolds at breathtaking speed, tools like Arlow anchor our observations in reliability—so our responses can be rooted in relevance.
When a child stacks four blocks, points to their nose, or laughs when you roll a ball back—they’re not just ‘passing a test.’ They’re communicating competence. And Arlow gives us a common, calibrated language to hear that message clearly.
That clarity changes everything: from referral pathways to classroom adaptations, from family conversations to policy decisions. It turns uncertainty into insight—and insight into opportunity.
For early childhood educators, that’s not just good practice. It’s foundational justice.
Because every toddler deserves to be seen—not just for what they haven’t done yet, but for everything they already are.
And Arlow, at its best, helps us do exactly that.
With fidelity. With care. With unwavering belief.




