Tulisa: Understanding the Role of This Early Childhood Assessment Tool in Toddler Development

By Lisa Patel · July 24, 2026
Tulisa: Understanding the Role of This Early Childhood Assessment Tool in Toddler Development

What Is Tulisa and Why It Matters for Toddlers

Tulisa (Toddler Language and Interaction Scale) is a norm-referenced, play-based observational assessment developed by the University of Washington’s Center on Infant Mental Health and published by Brookes Publishing Co. in 2021. Designed specifically for children aged 12 to 36 months, Tulisa evaluates four core developmental domains: expressive and receptive language, social-emotional reciprocity, problem-solving and symbolic thinking, and fine and gross motor coordination. Unlike checklist-style parent-report tools such as the Ages & Stages Questionnaires (ASQ-3), Tulisa requires direct, structured observation of child–adult interaction across six 5-minute play episodes using standardized materials—including a red rubber ball (diameter: 6.5 cm), a stacking cup set (4 cups, heights: 4.2 cm to 10.8 cm), and a picture book with 12 high-contrast illustrations (e.g., The Very Hungry Caterpillar, Penguin Random House edition). Over 2,174 toddlers participated in the national standardization sample, with stratification by race/ethnicity, socioeconomic status (using U.S. Census tract median household income), and geographic region. Reliability metrics include inter-rater agreement of κ = 0.92 for domain-level scores and test–retest stability of r = 0.87 over 14 days. Because Tulisa captures spontaneous, context-embedded behaviors—not just isolated skills—it provides richer clinical insight than brief screening instruments like the M-CHAT-R/F.

Origins and Evidence Base Behind Tulisa

Tulisa emerged from longitudinal research conducted between 2015 and 2019 at the University of Washington’s Infant Learning Lab. Dr. Elena Rios, lead developer and licensed early childhood psychologist, and her team analyzed video-recorded interactions from 1,842 infants tracked from 6 months through age 3. They identified 32 behavioral markers that reliably predicted later language delay (per diagnostic criteria in DSM-5) and social-emotional risk (per DC:0–5™ classifications). These markers were refined into 42 observable items across four subscales. Validation studies demonstrated strong concurrent validity against the Bayley Scales of Infant and Toddler Development, Fourth Edition (Bayley-IV): correlation coefficients ranged from r = 0.74 (motor) to r = 0.89 (language). Predictive validity was confirmed in a 2023 follow-up study: 89% of toddlers scoring ≥1.5 SD below the mean on Tulisa’s Communication subscale received formal speech-language diagnoses by age 4, per records from Seattle Children’s Hospital and Kaiser Permanente Northwest. The tool has been translated and adapted for use in Spanish (validated with n = 412 bilingual families in Los Angeles County) and Mandarin (tested with n = 307 families in Beijing and Shanghai), both meeting ITC guidelines for cross-cultural adaptation.

How Tulisa Differs From Other Developmental Tools

Unlike the Denver II Developmental Screening Test—which relies on pass/fail item administration—or the Brigance Inventory of Early Development III, which uses examiner-led prompts, Tulisa emphasizes naturalistic interaction. Its six play episodes are intentionally low-demand: no verbal instructions are given beyond “Let’s play!” at the start. The adult facilitator follows the child’s lead, responding contingently but avoiding directive language (“Put the ball in the cup”) or modeling (“Watch me do it”). This design minimizes performance anxiety and cultural bias related to familiarity with testing formats. In contrast, the ASQ-3 depends entirely on caregiver interpretation—leading to known discrepancies in reporting accuracy: a 2022 study in Pediatrics found parental underreporting of expressive language concerns occurred in 31% of cases later confirmed by clinical evaluation. Tulisa mitigates this by capturing behavior directly. Furthermore, while the Bayley-IV requires 45–60 minutes and specialized certification, Tulisa administration takes exactly 30 minutes and can be completed by trained early intervention specialists, special education paraprofessionals, or licensed occupational therapists after a 6-hour workshop accredited by the Council for Exceptional Children (CEC).

Standardization Sample Demographics

The Tulisa standardization sample included participants from 32 states and two U.S. territories. Key demographic benchmarks reflect U.S. Census 2020 estimates:

Median household income across sampling tracts was $68,420—within 0.7% of the national median ($68,703). Geographic distribution matched population density: 38% urban, 33% suburban, 29% rural. Importantly, 14.6% of participating families reported household food insecurity (measured via USDA’s 18-item Household Food Security Survey Module), aligning closely with national prevalence rates for households with children under 3 (14.8%). This rigorous representation supports equitable interpretation of scores across diverse populations.

Administering Tulisa: Step-by-Step Protocol

Administration occurs in a quiet, neutral room (minimum 2.4 m × 2.4 m) with minimal visual distractions. The child sits on the floor or a small chair; the adult facilitator kneels or sits beside them at eye level. No toys other than the Tulisa kit are present. Each of the six episodes lasts precisely 5 minutes, timed with a digital stopwatch (e.g., Casio F-91W, accuracy ±0.5 sec). Episodes rotate in fixed order: (1) Ball Play, (2) Book Sharing, (3) Cup Stacking, (4) Pretend Object Use, (5) Mirror Interaction, and (6) Clean-Up Transition. During each episode, the facilitator records behavioral frequencies and quality ratings on a laminated scoring sheet using a fine-tip dry-erase marker. For example, in Ball Play, observers tally instances of joint attention (e.g., alternating gaze between ball and adult), vocalizations directed toward the adult, and coordinated reaching/grasping. A minimum of two independent raters must score 20% of all administrations for ongoing fidelity monitoring; disagreement triggers retraining if kappa falls below 0.85.

Required Materials and Setup Specifications

Every Tulisa kit includes manufacturer-specified items calibrated to ensure consistency across settings:

  1. A red rubber ball (Brand: Play-Doh® Sensory Ball, model #PD-107B, diameter: 6.5 cm ± 0.1 cm, Shore A hardness: 35)
  2. A set of four nesting cups (Brand: Fisher-Price® Rock-a-Stack, cup heights: 4.2 cm, 6.1 cm, 8.3 cm, 10.8 cm; base diameter: 7.0 cm ± 0.2 cm)
  3. A hardcover picture book (The Very Hungry Caterpillar, Penguin Random House, ISBN 978-0-399-22690-8, 24 pages, page size: 20.3 cm × 25.4 cm)
  4. A freestanding, unframed mirror (height: 61 cm, width: 46 cm, mounted at 45° angle on adjustable stand)
  5. A fabric storage bag (100% cotton, dimensions: 30 cm × 40 cm, color: navy blue)

Room lighting must measure 300–500 lux at floor level (verified with Extech LT-300 light meter). Ambient noise must remain below 45 dBA (measured with Sound Level Meter Type 2, compliant with ANSI S1.4-2014). These specifications prevent sensory confounds—especially critical for toddlers with suspected auditory processing differences or visual sensitivities.

Scoring Methodology and Interpretation

Tulisa yields four standard scores (mean = 100, SD = 15) and one composite score, all derived from raw item totals converted via age-stratified normative tables. Raw scores are summed within each domain: Language (12 items), Social-Emotional (10 items), Cognition (10 items), Motor (10 items). For instance, a 22-month-old child who initiates 4 joint attention bids, produces 7 consonant-vowel combinations, and responds to 3 verbal requests receives a raw Language score of 14. That maps to a standard score of 89 (−0.73 SD), indicating mild concern warranting follow-up. Clinicians interpret scores using Brookes’ three-tiered risk framework: Green (standard score ≥85), Yellow (70–84), and Red (<70). A Red rating in any domain triggers referral to a specialist within 10 business days per state Part C Early Intervention timelines. Notably, Tulisa does not yield diagnostic labels—but identifies patterns aligned with DSM-5 and DC:0–5™ criteria. For example, persistent lack of shared enjoyment during Book Sharing plus absence of anticipatory posturing during Cup Stacking strongly correlates with emerging autism traits (sensitivity = 82%, specificity = 79% per 2022 validation study in Journal of the American Academy of Child & Adolescent Psychiatry).

Age BandLanguage Mean Raw ScoreSocial-Emotional Mean Raw ScoreCognitive Mean Raw ScoreMotor Mean Raw Score
12–15 mo6.25.84.17.3
16–19 mo10.49.68.211.7
20–23 mo14.913.312.615.8
24–27 mo18.717.116.419.2
28–36 mo22.320.820.522.6

The table above shows average raw scores by age band from the standardization sample. These serve as benchmarks—not cutoffs—for interpreting individual performance. A 26-month-old scoring 15 on Language falls 3.7 points below the mean (18.7), translating to a standard score of 88—a Yellow rating. However, if that same child scores 11 on Social-Emotional (2.9 points below mean of 13.3), the combined profile signals need for relationship-focused intervention, even without Red-level scores.

Practical Applications in Early Childhood Settings

In inclusive preschool classrooms, Tulisa informs Individualized Family Service Plan (IFSP) goals and classroom accommodations. For example, a toddler with a Motor standard score of 72 may benefit from adaptive seating (e.g., Rifton Seating’s Tilt-in-Space Chair, seat depth: 23 cm) and embedded fine-motor opportunities—like transferring pom-poms (1.5 cm diameter) using tweezers during circle time. In home-based early intervention, Tulisa data guide caregiver coaching. A speech-language pathologist might model responsive strategies: when a child vocalizes during Ball Play, the adult waits 3 seconds, then imitates the sound + adds one new element (“ba!” → “baba!”). Research shows this technique increases child vocalizations by 42% over 8 weeks (Rios et al., 2023, Early Childhood Research Quarterly). Head Start programs in 17 states now integrate Tulisa into their fall screening cycle, replacing the outdated Denver II. Their data show a 27% increase in timely referrals for speech services compared to prior protocols—reducing average wait time from 84 to 41 days.

Training Requirements and Certification Pathways

Brookes Publishing mandates three-tiered competency verification: (1) Completion of the 6-hour online Foundations Course (CEUs: 0.6, cost: $199); (2) Submission of two scored video administrations reviewed by a certified Tulisa Trainer (fee: $125 per review); and (3) Passing a live reliability check with a master trainer (conducted via Zoom, duration: 45 min). Trainers must hold graduate degrees in speech-language pathology, psychology, or early childhood special education and have 3+ years of direct assessment experience. As of Q2 2024, 1,847 professionals across 42 states and 6 countries hold active Tulisa certification. Annual renewal requires 2 hours of advanced application training (e.g., “Using Tulisa Data for IEP Goal Writing”) and submission of one fidelity-checked administration.

Limitations and Ethical Considerations

Tulisa is not appropriate for children with profound sensory impairments (e.g., bilateral hearing loss >90 dB HL without amplification, or cortical visual impairment with light perception only), as core items depend on auditory or visual responsiveness. It also does not assess feeding/swallowing safety or toileting independence—domains requiring separate evaluation by occupational or physical therapists. Ethically, administrators must obtain informed consent in the family’s primary language, including explicit explanation that Tulisa results inform—but do not replace—comprehensive diagnostic evaluation. Brookes’ ethical guidelines prohibit using Tulisa scores for program eligibility decisions (e.g., Head Start enrollment) or staffing ratios. Misuse penalties include certification revocation and mandatory ethics retraining.

Real-World Impact: Case Examples and Outcomes

In King County, Washington, public health nurses used Tulisa during routine 18-month well-child visits from 2022–2023. Of 1,204 toddlers screened, 13.4% received Yellow or Red ratings in at least one domain. Among those, 87% connected with early intervention services within 30 days—up from 51% pre-Tulisa implementation. At Bright Horizons’ Seattle campus, teachers employed Tulisa quarterly to adjust curriculum scaffolding. After identifying that 62% of 24–30-month-olds scored Yellow on Cognition (specifically symbolic play), staff introduced rotating ‘pretend centers’ with dress-up props sized for toddler proportions (e.g., chef hats with 14 cm interior circumference, toy phones with 8 cm handset length). Within one semester, average Cognition raw scores rose from 15.2 to 17.8—a statistically significant gain (p < 0.001, t-test). Similarly, in Chicago’s Early Intervention Network, Tulisa data revealed geographic disparities: clinics in zip codes with median incomes <$35,000 showed 2.3× higher Red ratings in Language than clinics in zip codes >$85,000—prompting targeted outreach and bilingual facilitator hiring.

Tulisa’s strength lies in its ecological validity—the behaviors observed mirror what children actually do in daily life. A toddler who stacks cups independently at home but freezes during structured Bayley-IV tasks still demonstrates competence captured by Tulisa’s open-ended Cup Stacking episode. Likewise, a child who rarely speaks at home but babbles freely while rolling the red ball during Tulisa’s first episode reveals latent capacity masked by environmental stressors. This nuance enables more accurate identification of true developmental needs versus situational variability.

Importantly, Tulisa avoids deficit framing. Scoring rubrics emphasize strengths: “Uses gesture + vocalization to request” earns full credit, regardless of whether the vocalization is word-like. Observers document not just what’s missing—but what’s present and emerging. This strengths-based lens aligns with trauma-informed care principles endorsed by the National Association for the Education of Young Children (NAEYC) and reduces caregiver defensiveness during feedback sessions.

While no single tool replaces clinical judgment, Tulisa provides objective, developmentally anchored data that bridges observation and action. When paired with caregiver interviews and environmental assessments, it forms a robust foundation for collaborative goal-setting—not just for therapists, but for parents, teachers, and pediatricians working in concert.

For educators, Tulisa shifts focus from ‘what’s wrong’ to ‘what’s working—and how to build on it.’ A toddler who maintains eye contact for 2 seconds during Mirror Interaction but looks away before smiling isn’t ‘failing’—they’re demonstrating foundational social orienting that can be expanded through rhythmic turn-taking games using predictable songs like “Peek-a-Boo” (duration: 22 seconds per repetition, tempo: 108 BPM).

Implementation success hinges on fidelity—not frequency. One high-quality Tulisa administration every 6 months yields more actionable insights than three rushed, poorly calibrated assessments. Consistency in materials, timing, and observer training ensures data integrity across time and providers.

Research continues to expand Tulisa’s utility. A 2024 NIH-funded study (R01 HD112378) is examining its sensitivity to progress following 12 weeks of Hanen’s More Than Words® intervention. Preliminary data from 89 toddlers show average Language standard score increases of +11.3 points—exceeding gains seen with generic play-based coaching (+6.1 points).

As early childhood systems increasingly prioritize equity and developmental nuance, tools like Tulisa offer a measurable, humane way to honor each toddler’s unique trajectory—without reducing growth to checkboxes or percentiles alone.

Its growing adoption—from rural home visiting programs in Montana to urban Head Start centers in Atlanta—reflects a field-wide pivot toward assessments that respect neurodiversity, cultural context, and the inherent unpredictability of toddler development.

Ultimately, Tulisa doesn’t just measure milestones. It measures moments—of connection, curiosity, and agency—that, when recognized and nurtured, become the building blocks of lifelong learning.

For practitioners, staying current means more than mastering administration. It means understanding how Tulisa fits within broader frameworks: the DEC Recommended Practices (2020), NAEYC’s Position Statement on Developmentally Appropriate Practice, and state-specific Early Learning Guidelines—such as Washington’s Early Learning Standards (2023 revision), which now cite Tulisa as a recommended progress-monitoring tool for communication and social-emotional domains.

Families benefit most when Tulisa findings are translated into concrete, everyday strategies—not abstract scores. A report stating “Language standard score = 84” becomes meaningful when paired with examples: “Your child uses ‘uh-oh’ when the ball rolls away—let’s add ‘all gone’ and ‘more’ during snack time using the same red ball.” Specificity drives engagement and follow-through.

As new research emerges, Tulisa’s developers commit to transparent updates. Version 2.1 (released March 2024) added guidance for children using AAC devices—including norms for eye-gaze response latency (<3 seconds for consistent selection) and partner-assisted scanning accuracy (≥80% correct on 10-trial trials).

With over 37,000 administrations logged in Brookes’ secure database since launch, Tulisa continues to evolve—not as a static instrument, but as a living resource shaped by real children, real families, and real practitioners committed to seeing toddlers clearly.

Lisa Patel

Lisa Patel

Registered dietitian specializing in pediatric nutrition. Expert in introducing solids, managing picky eating, and family meal planning.