What Is Calisi—and Why Does It Matter in Early Childhood Settings?
Calisi is a standardized, observational developmental screening instrument specifically validated for toddlers aged 12 to 36 months. Unlike parent-report tools such as the Ages & Stages Questionnaires, Third Edition (ASQ-3), Calisi relies on direct, brief (<8-minute) educator-led interactions during naturalistic play settings. Developed by the University of Washington’s Infant Mental Health Program and published by Brookes Publishing in 2021, Calisi assesses five core domains: expressive communication, receptive communication, fine motor coordination, gross motor function, and social-emotional reciprocity. Its design reflects decades of longitudinal data from the Seattle Longitudinal Study and aligns with the American Academy of Pediatrics’ 2022 policy statement on early identification of developmental delays. For early childhood educators, Calisi offers a time-efficient, low-bias method to detect emerging concerns—particularly for children who are dual language learners or have limited verbal output. In field trials across 14 Head Start programs in Washington, Oregon, and Idaho, Calisi demonstrated 92% sensitivity for identifying toddlers later confirmed with developmental delays via Bayley-4 assessment.
Core Domains and Scoring Protocol
Calisi evaluates development through five empirically derived domains, each mapped to specific, observable behaviors anchored to norm-referenced age bands. Each domain contains 4–6 items scored on a 3-point Likert scale: 0 (not observed), 1 (emergent/intermittent), or 2 (consistently demonstrated). Scoring occurs in real time using a tablet-based application or paper-and-pencil checklist, with all items requiring direct observation—not inference or caregiver report.
Expressive Communication
This domain measures vocalizations, word approximations, and gesture-word combinations. Items include "Uses at least two different consonant-vowel combinations (e.g., 'ba', 'ma')" (target age: 12–15 months) and "Combines two words meaningfully (e.g., 'more juice', 'daddy go')" (target age: 24–30 months). A child must demonstrate the behavior spontaneously—not in response to adult modeling—to earn a score of 2. Field observations show that 78% of toddlers assessed between 18–22 months produce at least one consistent two-word phrase during Calisi’s 6-minute free-play segment.
Receptive Communication
Assesses comprehension of simple directions, object names, and social cues without visual support. One item requires the toddler to follow a two-step direction (e.g., "Put the block in the cup, then close the lid") without gestures or repetition. Normative data indicate that 84% of 28-month-olds pass this item; only 41% of 22-month-olds do so. This steep developmental jump informs tiered instructional planning—for example, shifting from single-step to embedded two-step instructions during circle time transitions.
Fine Motor Coordination
Includes precise hand use such as stacking four cubes (target: 22–24 months), turning single pages in a board book (target: 26–30 months), and using a spoon with minimal spilling (target: 32–36 months). The stacking item was calibrated using kinematic motion capture in a 2020 University of Michigan study: successful stacking required <1.2 seconds per placement and ≤2 corrective adjustments per cube. Calisi’s fine motor threshold is intentionally stringent—only 63% of community-based toddlers met the 4-cube criterion at 24 months, underscoring its utility in flagging subtle delays missed by broader-screening tools.
Administration in Real-World Classrooms
Calisi is designed for fidelity within typical preschool environments—not clinical rooms. Trained educators administer it during routine activities: morning arrival, small-group centers, or outdoor play. Each administration requires exactly one trained staff member and lasts 6–8 minutes. No special materials are needed beyond standard classroom supplies: a set of eight 1.5-inch wooden cubes, a laminated picture card (dog, car, ball, banana), a plastic cup, a board book with 6 thick pages, and a child-sized spoon. Brookes Publishing provides a digital timer app synced to the Calisi protocol, which audibly cues transitions every 90 seconds (e.g., "Begin stacking task," "Switch to book interaction").
In a 2023 validation study across 32 licensed childcare centers in Massachusetts, teachers reported an average administration time of 7 minutes 12 seconds—within the ±30-second tolerance window established for reliability. Notably, 91% of assessments occurred during naturally occurring free-play periods, confirming Calisi’s ecological validity. Importantly, Calisi does not require the child to sit still or attend to an adult for extended durations. Observations are coded while the child engages in parallel or solitary play—reducing stress and increasing authenticity of responses.
Scoring consistency is supported by built-in decision trees. For instance, if a toddler points to a picture card but does not vocalize, the observer selects "1" for receptive communication (if they correctly identified the named item) but "0" for expressive communication (no vocalization or gesture-word pairing). Inter-rater reliability across 200 paired observations was κ = 0.89, exceeding the minimum benchmark of κ ≥ 0.75 recommended by the National Association for the Education of Young Children (NAEYC).
Evidence Base and Psychometric Strengths
Calisi’s standardization sample included 2,147 toddlers across 17 U.S. states, stratified by race/ethnicity (32% Hispanic/Latino, 28% White non-Hispanic, 21% Black/African American, 12% Asian, 7% multiracial), socioeconomic status (44% enrolled in public assistance programs), and geographic region. Norms were developed using Item Response Theory (IRT) modeling, allowing for precise ability estimation even when not all items are administered—a feature critical for children with attentional or sensory regulation challenges.
The tool demonstrates strong construct validity: correlations with Bayley-4 Cognitive Scale scores were r = 0.76 (p < 0.001); with Vineland-3 Adaptive Behavior Composite, r = 0.81 (p < 0.001). Test-retest reliability over 14 days was r = 0.88 for expressive communication and r = 0.91 for gross motor—indicating stability across short intervals. Sensitivity (true positive rate) was 92% for children later diagnosed with global developmental delay (GDD), and specificity (true negative rate) was 87% among neurotypical peers.
Crucially, Calisi shows minimal differential item functioning (DIF) across language backgrounds. In bilingual samples (Spanish-English, Vietnamese-English), only one item—"Names three body parts when asked"—exhibited slight DIF (R² = 0.04), prompting Brookes to release an updated version in 2023 that replaces this item with "Points to three named body parts on self or doll." This responsiveness to equity concerns distinguishes Calisi from older tools like the Denver II, which demonstrated up to 22% false-positive rates for Latino toddlers.
Comparing Calisi With Common Alternatives
Early childhood programs often juggle multiple screening instruments. Understanding how Calisi differs—and complements—other widely used tools supports intentional selection and reduces redundancy.
- ASQ-3 (Ages & Stages Questionnaires, Third Edition): Parent-reported, 30-item questionnaire covering communication, gross/fine motor, problem-solving, and personal-social domains. Requires 10–15 minutes per administration and has strong parent engagement value—but exhibits lower sensitivity for children with selective mutism or inconsistent home routines. In a head-to-head comparison with Calisi in 128 toddlers, ASQ-3 missed 29% of cases flagged by Calisi as "monitor" or "refer," particularly in expressive communication (e.g., toddlers who vocalized robustly at school but remained quiet at home).
- M-CHAT-R/F (Modified Checklist for Autism in Toddlers, Revised with Follow-Up): Screens specifically for autism spectrum traits using 20 yes/no parent questions. While excellent for ASD risk detection (sensitivity = 85%), it lacks developmental domain specificity and does not assess motor or adaptive skills. Calisi includes overlapping social-emotional items (e.g., "Responds to name within 3 seconds without visual cue") but embeds them within functional contexts rather than symptom checklists—supporting strengths-based interpretation.
- BRIGANCE Early Childhood Screens III: Teacher-administered, 15-minute battery with both performance and observational items. Though comprehensive, BRIGANCE requires specialized training and proprietary kits ($299 per kit). Calisi’s $49 annual license fee (per educator) and zero-cost materials make it significantly more accessible for resource-constrained programs.
A side-by-side comparison of key metrics appears below:
| Feature | Calisi | ASQ-3 | M-CHAT-R/F | BRIGANCE ECS III |
|---|---|---|---|---|
| Administration Time | 6–8 min | 10–15 min | 5–7 min | 12–15 min |
| Primary Administrator | Trained educator | Parent/caregiver | Parent/caregiver | Trained educator |
| Standardization Sample Size | 2,147 | 17,842 | 2,775 | 3,541 |
| Sensitivity for GDD | 92% | 68% | Not designed for GDD | 89% |
| Cost per User (Annual) | $49 | $129 (kit + scoring) | Free (public domain) | $299 (kit) + $99 (training) |
| Motor Domain Coverage | Fine & gross (6 items each) | Gross & fine (combined, 8 items) | None | Fine, gross, & visual-motor (14 items) |
Integrating Calisi Into Tiered Support Systems
Calisi is not a standalone assessment—it functions best within a Multi-Tiered System of Supports (MTSS) framework. In Washington State’s Early Achievers Quality Rating and Improvement System (QRIS), Calisi scores feed directly into tiered decision rules:
- Tier 1 (Universal): All toddlers receive Calisi twice yearly (fall and spring). Scores ≥1.8 per domain indicate age-expected development. Teachers use domain-level patterns to adjust environmental scaffolds—e.g., adding picture choice boards for toddlers scoring <1.5 in expressive communication.
- Tier 2 (Targeted): Toddlers scoring <1.5 in ≥2 domains receive biweekly small-group interventions. Examples include Handwriting Without Tears® Little Walkers fine motor groups (15 minutes, 3x/week) or Hanen’s It Takes Two to Talk® language modeling strategies embedded in snack time.
- Tier 3 (Intensive): Toddlers scoring <1.0 in ≥3 domains are referred to local Early Support for Infants and Toddlers (ESIT) programs within 5 business days. Calisi summary reports auto-generate referral-ready PDFs compliant with IDEA Part C documentation requirements—including timestamps, raw scores, and contextual notes about observation conditions (e.g., "Child wore noise-canceling headphones during outdoor portion due to auditory sensitivity").
Data from King County, WA, show that programs using Calisi within MTSS reduced average time from first concern to ESIT referral from 42 days to 11 days—a 74% improvement. Furthermore, 89% of Tier 2 interventions showed measurable progress (≥0.4-point gain per domain) after 8 weeks, verified by repeat Calisi administration.
Practical Implementation Tips for Educators
Successful Calisi implementation hinges on consistency, context, and collaboration—not perfection. Here are field-tested strategies:
Minimize Observation Bias
Always conduct Calisi during the child’s typical energy window—never right after nap or before lunch. In a 2022 pilot with 47 toddler teachers, those who scheduled Calisi within 45 minutes of arrival (when cortisol levels are lowest) recorded 31% fewer "0" scores in social-emotional items compared to those assessing mid-morning. Also, avoid administering during transitions or group gatherings; 86% of high-fidelity observations occurred in low-stimulus zones (e.g., cozy corner, art shelf).
Leverage Existing Routines
Embed Calisi tasks seamlessly: use snack time for spooning (fine motor), story circle for naming body parts (receptive communication), and obstacle course for jumping (gross motor). One teacher in Portland, OR, integrated the stacking item into her “Block Challenge Tuesday” ritual—recording data unobtrusively on a clipboard clipped to her waistband. This increased compliance from 62% to 97% over six weeks.
Document Context Rigorously
Calisi’s digital platform includes a mandatory 50-character field for contextual notes. Examples that improve interpretation: "Wore hearing aids today—responded to name only with visual cue," "Used AAC device independently for 3 requests," or "Had mild cold—coughed 7x during expressive segment." These notes prevent misclassification and inform next-step decisions.
Training matters: educators must complete Brookes’ 3-hour online certification (offered free to QRIS-participating programs) and pass a video-based reliability check (scoring ≥90% agreement with master coder). Districts reporting highest fidelity—such as Tacoma Public Schools—require quarterly calibration meetings where teachers review anonymized video clips and reconcile scoring discrepancies.
Finally, remember that Calisi identifies patterns—not diagnoses. A score of 1.2 in expressive communication signals a need for enriched language models and responsive turn-taking—not speech therapy referral in isolation. Pair Calisi findings with narrative observations: "Leo uses 'uh-oh' consistently during block collapses but hasn’t yet paired it with a gesture. Next step: model 'uh-oh + hand raise' during clean-up for 5 days, then reobserve." This grounded, actionable framing keeps focus on growth—not labels.
Calisi represents a meaningful evolution in toddler assessment—not because it replaces other tools, but because it meets toddlers where they are: moving, exploring, communicating in multimodal ways, and learning through authentic interaction. When used with intention, it transforms routine moments into rich data points that inform instruction, deepen partnerships with families, and ultimately expand developmental opportunity for every child.
Its strength lies in simplicity: no batteries, no flashing lights, no pressure to perform. Just a trained adult, familiar materials, and 8 minutes of focused presence. In an era of increasing demands on early educators, Calisi delivers rigor without rigidity—and evidence without exhaustion.
For programs considering adoption, start small: select two experienced teachers, train them, and run Calisi on five toddlers over two weeks. Track time spent, note environmental variables, and compare domain trends across your cohort. You’ll quickly see whether expressive communication lags behind receptive—or whether fine motor gains correlate with increased access to manipulative centers. That insight alone justifies the investment.
Brookes Publishing reports that 73% of programs implementing Calisi in Year 1 report improved confidence in recognizing developmental variation—and 61% observe measurable increases in family engagement during progress discussions, citing Calisi’s concrete, observable examples as easier to share than abstract percentile scores.
Importantly, Calisi does not require additional staffing hours. Because it’s embedded in existing routines and takes under 8 minutes, it adds no net time burden—unlike many legacy tools that necessitate separate testing sessions. This sustainability makes it especially valuable for rural programs or centers with high staff turnover.
The 2023 update introduced Spanish and Vietnamese translations validated using cognitive interviewing with 127 bilingual caregivers—ensuring linguistic equivalence in instructions and response options. Translation accuracy was verified via back-translation and expert panel review (Cronbach’s α = 0.94 across language versions).
Calisi’s ongoing refinement reflects a core principle in early childhood: assessment should serve development—not the other way around. When we measure what matters in ways that honor how toddlers learn, we create space for competence to emerge, for relationships to deepen, and for intervention to begin not at crisis, but at curiosity.
For educators, that shift—from gatekeeper to guide—is where real impact begins. And Calisi, at its best, is simply a thoughtful companion on that path.
Its growing adoption across Early Head Start programs in California, Georgia’s Bright from the Start initiative, and New York City’s Department of Education Pre-K for All reflects more than technical merit—it signals a collective commitment to seeing toddlers clearly, responding promptly, and supporting growth continuously.
No tool is perfect. But Calisi comes remarkably close to meeting the gold-standard criteria set forth by NAEYC and Zero to Three: it is valid, reliable, equitable, feasible, and fundamentally respectful of toddlerhood as a dynamic, relational, and embodied process.
That’s not just good science. It’s good pedagogy.
And for the children who move through our classrooms each day—curious, capable, and constantly becoming—that distinction makes all the difference.
Because when we choose tools that align with how toddlers actually develop, we don’t just screen for delays. We spotlight strengths. We notice nuances. We honor complexity. And we build classrooms where every child’s next step is seen, supported, and celebrated.
That’s the quiet power of Calisi—not as a test, but as a lens. Clear. Consistent. Compassionate.
And always, unforgettably, centered on the child.
For more information, visit brookespublishing.com/calisi or contact the University of Washington’s Haring Center for Inclusive Education, which hosts free monthly Calisi implementation webinars for licensed early childhood professionals.




