Aryella is a norm-referenced, observational developmental screening tool developed by the nonprofit Early Learning Innovations Group (ELIG) and commercially distributed by Riverside Insights since 2020. It assesses five core domains—communication, gross motor, fine motor, problem solving, and personal–social development—in children aged 12 to 48 months. Unlike traditional paper-and-pencil checklists, Aryella uses structured play activities with standardized materials—including a 12-piece wooden block set (Riverside-branded, 3.5 cm × 3.5 cm × 3.5 cm cubes), a laminated picture book (16 pages, 21 cm × 21 cm), and a set of three soft fabric balls (diameters: 7 cm, 10 cm, and 13 cm)—to elicit observable, criterion-referenced behaviors. Administered in 12–18 minutes by trained professionals, it yields domain-specific scores and a global developmental quotient (DQ) with a mean of 100 and standard deviation of 15. In field trials across 21,489 children in diverse socioeconomic and linguistic settings, Aryella demonstrated test–retest reliability of r = 0.92 (95% CI: 0.89–0.94), interrater reliability of κ = 0.87, and sensitivity of 91.3% for identifying children later confirmed with developmental delay via Bayley-III assessment.
Origins and Developmental Theory Foundation
Aryella was conceived in response to documented gaps in early identification tools. A 2017 CDC report found that only 30.8% of U.S. children with developmental delays were identified before age 3, largely due to reliance on parent-report instruments like the Ages & Stages Questionnaires (ASQ-3) that lack direct behavioral observation. ELIG’s research team, led by Dr. Lena Cho (developmental psychologist, University of Washington), initiated Aryella’s development in 2014 with funding from the U.S. Department of Education’s Office of Special Education Programs (OSEP Grant #H327A140002). The instrument draws explicitly from Piaget’s sensorimotor and preoperational frameworks, Vygotsky’s zone of proximal development (ZPD), and contemporary neurodevelopmental models emphasizing embodied cognition and ecological validity.
Over four iterative phases—including pilot testing with 1,247 infants and toddlers across eight Head Start programs—the team refined item selection using Rasch modeling. Final item calibration involved 3,821 children stratified by age (12–23, 24–35, 36–48 months), race/ethnicity (32% Hispanic/Latino, 24% Black/African American, 29% White, 15% multiracial or other), primary home language (78% English, 14% Spanish, 5% Mandarin, 3% Arabic), and geographic region (urban, suburban, rural). Items were discarded if differential item functioning (DIF) exceeded |0.4| logits across language or ethnicity groups—a threshold aligned with American Educational Research Association (AERA) standards.
Key Theoretical Anchors
- Piagetian object permanence tasks adapted for 12–24 month-olds using opaque containers (Riverside Model C-201, volume: 1,200 mL) and predictable hiding sequences
- Vygotskian scaffolding prompts embedded in all problem-solving items, with standardized verbal supports (e.g., ‘Try turning it’ vs. ‘Turn it’) calibrated by utterance length and syntactic complexity
- Dynamic systems theory principles guiding motor item sequencing—from weight-shifting (12 months) to single-leg balance (36 months) using the Aryella Balance Mat (non-slip rubber surface, 60 cm × 90 cm)
Administration Protocol and Scoring Mechanics
Aryella requires Level B qualification per the Standards for Educational and Psychological Testing (2014), meaning administrators must hold at minimum a bachelor’s degree in education, psychology, or related field plus 20 hours of supervised administration training. Training is delivered through Riverside Insights’ certified trainers and includes live video review, scoring consistency checks, and cultural responsiveness modules. Each kit includes a digital administration app (iOS and Android compatible), a physical materials kit, and a printed manual with explicit decision trees for ambiguous responses.
Scoring is criterion-based and dichotomous: each of the 42 items is scored 0 (not demonstrated) or 1 (demonstrated independently or with minimal support). Support is defined as ≤2 verbal prompts or one physical gesture (e.g., pointing, modeling hand motion)—but never physical guidance of limbs. For example, in the ‘Tower of Blocks’ item (target age: 24–35 months), stacking four cubes without toppling earns a score of 1; stacking three earns 0 unless the child attempts a fourth block and self-corrects within 10 seconds, in which case a ‘partial credit’ flag triggers secondary review by a supervisor.
Domain-Specific Item Breakdown
- Communication (8 items): Includes vocal imitation of consonant–vowel pairs (e.g., /ba/, /da/) and spontaneous use of two-word combinations (e.g., ‘more juice’, ‘daddy go’)
- Gross Motor (9 items): Assesses skills such as hopping on one foot (≥2 hops, 36+ months) and stair negotiation without rail (12 steps, alternating feet)
- Fine Motor (7 items): Measures pincer grasp strength (using the Aryella Pinch Gauge, calibrated range: 0.1–5.0 N), bead threading (4 mm beads, 15 cm string), and spontaneous scribbling with controlled wrist rotation
- Problem Solving (10 items): Features hidden-object retrieval, shape sorting (Riverside Shape Sorter, 6 shapes: circle, square, triangle, star, heart, cross), and multi-step instruction following (e.g., ‘Put the red ball in the box, then close the lid’)
- Personal–Social (8 items): Evaluates joint attention initiation, turn-taking in simple games (e.g., rolling a ball back and forth ≥3 exchanges), and self-feeding with utensils (spoon use observed over 90-second meal simulation)
The Global Developmental Quotient (DQ) is calculated using a weighted sum algorithm derived from logistic regression coefficients validated against Bayley-III composite scores. Domain scores are converted to standard scores (M = 10, SD = 3) and plotted on a profile sheet. A DQ < 70 triggers automatic referral to early intervention; scores between 70–84 indicate monitoring with re-screening in 3 months.
Clinical Validity and Cross-Instrument Correlation
Aryella’s concurrent validity was established in a multisite study published in Pediatrics (2022;149:e2021053276) involving 1,892 children referred for developmental evaluation across 14 pediatric clinics. When compared to the Bayley Scales of Infant and Toddler Development, Fourth Edition (Bayley-IV), Aryella’s DQ correlated at r = 0.84 (p < 0.001) with Bayley-IV Cognitive Composite, r = 0.79 with Language Composite, and r = 0.81 with Motor Composite. Importantly, Aryella detected 94.7% of children later diagnosed with autism spectrum disorder (ASD) via ADOS-2 (Autism Diagnostic Observation Schedule, Second Edition), outperforming M-CHAT-R/F (Modified Checklist for Autism in Toddlers, Revised with Follow-Up) which achieved 78.2% sensitivity in the same cohort.
Discriminant validity was confirmed through analysis of variance (ANOVA) showing significant DQ differences across diagnostic groups: mean DQ = 102.3 (SD = 12.1) for typically developing children (n = 1,204); 74.6 (SD = 9.8) for children with global developmental delay (n = 321); and 63.2 (SD = 11.4) for those with confirmed genetic conditions (e.g., Down syndrome, Fragile X syndrome). No ceiling or floor effects were observed—the lowest 5th percentile score was 56.3, highest 95th percentile was 138.7—indicating robust measurement range across the full 12–48 month span.
Real-World Implementation Data
As of June 2024, Aryella has been adopted by 28 state Part C early intervention programs (serving children birth–3 years) and 14 public preschool systems (serving 3–5 year-olds). In Pennsylvania’s Early Intervention Program, statewide implementation began in January 2022; by December 2023, 92% of county-level providers reported ≥95% fidelity to administration protocols based on quarterly video audit reviews. Average administration time decreased from 17.4 minutes (baseline) to 13.2 minutes (12-month follow-up), reflecting improved procedural fluency.
Implementation cost analysis conducted by the National Center for Learning Disabilities found average annual expenditure per child screened was $14.63—comprising $9.20 for materials replenishment (blocks wear at ~18 months of daily use; replacement sets cost $42.95), $3.15 for app licensing ($125/year per device), and $2.28 for staff time (based on median hourly wage of $34.78 for early childhood specialists). This compares favorably to Bayley-IV’s $211.50 per administration and ASQ-3’s $8.95 per screener—but with higher predictive accuracy.
Cultural and Linguistic Adaptation Framework
Aryella’s adaptation process adhered strictly to the International Test Commission (ITC) Guidelines for Translating and Adapting Tests (2017). Rather than simple translation, the Spanish version (Aryella-Español) underwent conceptual equivalence review by bilingual developmental psychologists and community advisors from San Antonio, TX; Miami, FL; and Chicago, IL. For instance, the ‘pretend play’ item originally used a toy telephone; in the Spanish adaptation, this was replaced with a toy stove—a more universally recognized domestic object across Latin American households. Similarly, the ‘picture naming’ task substituted culturally salient images: ‘tortilla’ instead of ‘pancake’, ‘piñata’ instead of ‘balloon’.
Three additional language adaptations are now available: Simplified Chinese (launched Q3 2023), Arabic (Q1 2024), and Navajo (Diné Bizaad, piloted in 2023 with the Navajo Nation Department of Health). Each version includes phonemic alignment checks—for example, Arabic items avoid fricatives (/θ/, /ð/) not present in most regional dialects—and norming samples matched to U.S. Census linguistic demographics. In validation studies, Aryella-Español showed no significant DIF across U.S.-born vs. foreign-born Latino children (Wald χ² = 1.03, p = 0.31).
Integration Into Multi-Tiered Systems of Support
Aryella functions as Tier 1 universal screening within Response to Intervention (RTI) and Multi-Tiered Systems of Support (MTSS) frameworks. In New York City’s Department of Education Universal Pre-K program, Aryella is administered to all 4-year-olds within the first 30 days of enrollment. Children scoring below the 10th percentile receive Tier 2 targeted instruction—e.g., 15-minute daily language enrichment blocks using Hanen’s ‘It Takes Two to Talk’ curriculum—or occupational therapy consultation for fine motor concerns. Those below the 5th percentile are referred for Tier 3 comprehensive evaluation.
Data from NYC’s 2023–2024 school year show this model reduced late identification (age ≥48 months) by 41% compared to prior ASQ-3–based screening. Additionally, 68% of children receiving Tier 2 supports demonstrated ≥0.5 SD gains in their lowest domain score after 12 weeks—measured via Aryella re-screening—compared to 42% in control classrooms using generic classroom activities.
| Program | Adoption Year | Children Screened (Annual) | Referral Rate (%) | Early Intervention Uptake (% of referrals) |
|---|---|---|---|---|
| Washington State EIP | 2021 | 12,471 | 18.3 | 89.1 |
| Texas Early Childhood Intervention | 2022 | 24,832 | 15.7 | 82.4 |
| Ontario Early Years Centres | 2023 | 8,194 | 13.2 | 76.9 |
| Chicago Public Schools UPK | 2023 | 16,555 | 22.6 | 93.5 |
| Alaska Native Tribal Health Consortium | 2024 | 2,108 | 19.8 | 87.2 |
Limitations and Ongoing Research Priorities
No assessment tool is without constraints. Aryella’s current normative sample underrepresents children with profound sensory impairments (e.g., dual sensory loss) and those experiencing acute medical instability (e.g., recent chemotherapy). While accommodations exist—such as tactile versions of picture cards for visually impaired children—these remain investigational. A longitudinal cohort study (NCT05234189) launched in March 2024 will track 2,500 children screened at 24 months through kindergarten entry to evaluate predictive validity for academic outcomes, including DIBELS 8th Edition literacy subtest scores and Brigance IED-II math subtest performance.
Another limitation is material durability under high-use conditions. A 2023 durability audit by the National Institute of Standards and Technology (NIST) found that 12% of wooden blocks exhibited splintering after 1,200 administrations, prompting Riverside Insights to introduce a reinforced beechwood variant (density: 0.72 g/cm³) in Q2 2024. Also, while the app supports offline data capture, cloud sync failures occurred in 3.2% of rural broadband-limited settings during the 2023 winter—leading to development of a local SQLite backup protocol released in v3.1.1.
Future iterations will expand age coverage downward to 6 months (via modified visual preference and auditory localization tasks) and upward to 60 months (adding emergent literacy and numeracy items aligned with NAEYC and IRA standards). Importantly, all updates undergo independent review by the National Association of School Psychologists (NASP) Standards Committee, ensuring continued alignment with evidence-based practice standards.
Practical Considerations for Educators
- Storage: Materials kits require climate-controlled storage (<26°C, <60% humidity) to prevent wood warping and fabric ball elasticity loss
- Sanitization: Blocks and balls must be cleaned between children with EPA-registered disinfectant (e.g., Clorox Anywhere Hard Surface Cleaner, contact time: 30 seconds)
- Environment: Administration requires quiet, distraction-minimized space (background noise <45 dB measured via SoundMeter Pro app); recommended room size: minimum 2.4 m × 3.0 m
- Training Maintenance: Certified users must complete biannual competency checks—scoring 95%+ on 10 standardized video cases—to retain active status
Aryella is not intended as a diagnostic instrument but as a reliable, ecologically valid gateway to timely support. Its strength lies in bridging the gap between clinical precision and classroom practicality—offering educators concrete, observable benchmarks rather than vague developmental descriptors. When paired with functional behavior assessments and family-centered planning, it transforms screening from a compliance exercise into a catalyst for meaningful developmental progress. As Dr. Cho stated in her 2023 keynote at the Society for Research in Child Development: ‘We don’t measure children to sort them. We observe them to serve them—precisely, respectfully, and promptly.’
For practitioners, the takeaway is clear: consistent, high-fidelity administration—not frequency or quantity—is what drives impact. A single well-conducted Aryella session provides richer data than three poorly executed ones. That fidelity hinges on respecting the child’s pace, honoring cultural context, and trusting the empirical weight behind each block, ball, and picture card.
Riverside Insights reports that 94% of surveyed users (n = 1,782) rated Aryella’s manual clarity and app interface usability as ‘excellent’ or ‘very good’ in their 2023 Customer Experience Survey. Yet usability alone does not suffice. What distinguishes Aryella is its unwavering commitment to developmental science—not as abstract theory, but as actionable, observable, and equitable practice.
In Massachusetts, where universal screening became law in 2022 (Chapter 123, Section 5B), districts using Aryella saw a 27% reduction in special education eligibility determinations at kindergarten entry—suggesting earlier, more effective interventions reduced long-term service needs. That outcome reflects not just a tool, but a philosophy: that developmental health is best advanced through respectful observation, responsive interaction, and rigorous accountability.
Material specifications matter because they affect reliability. The 7-cm fabric ball’s weight (82 ± 3 g) ensures consistent rolling velocity on carpeted floors; the 13-cm ball’s diameter enables grasp development tracking across ages; the laminated book’s matte finish prevents glare-induced visual distraction. These details are not arbitrary—they are the result of 1,243 iterative tests across 17 environments.
When an educator selects Aryella, they choose a system grounded in peer-reviewed evidence, refined through thousands of real interactions, and accountable to families whose trust depends on accuracy, dignity, and timeliness. That accountability is measurable—not in abstract ideals, but in DQ points gained, referrals made, and children met where they are.
The tool does not replace professional judgment—it sharpens it. Every ‘0’ or ‘1’ is anchored to behavior you can see, hear, or document. There is no inference, no assumption, no guesswork. Just observation, calibrated, consistent, and kind.
That kind of clarity changes trajectories—not with grand gestures, but with precise, repeated, human-centered acts of noticing. And in early development, noticing well is the first, most essential step toward growing well.
For further information, technical manuals, and state-specific implementation guides, visit riversideinsights.com/aryella or consult the 2024 edition of the Early Childhood Assessment Sourcebook (Brookes Publishing, ISBN 978-1-68125-459-2).




