Portal & CoCoComelon: Sing & Play with JJ2024×CoComelon: Sing & Play with JJ.
Designing CoComelon's First Voice-Interactive Game Suite for Preschoolers.
The brief.
CoComelon: Sing & Play with JJ is a voice-interactive game suite for connected TV and mobile. Kids ages two to four sing, answer, move, and play along with JJ across seven mini-games. I led narrative and UX design alongside a Volley PM, taking the product from early concepts through launch with Volley Games and Moonbug's CoComelon IP. The design problem was timing. Play sessions showed that preschoolers lose attention within roughly 2 to 4 seconds if the system does not respond, but they also need that same 2 to 4 second pause to form a response. The experience had to wait long enough for a child to participate without letting the game feel stalled. That timing model shaped the prompt structure, animation cues, recovery behavior, and pacing across the full suite.

Designing CoComelon's First Voice-Interactive Game Suite for Preschoolers.
Wheels on the Bus
Action verse participation with two-pass loop
Peek A Boo
Hide and seek with JJ using visual cues
Baa Baa Black Sheep
Color & wool collection with micro-rewards
Bingo
Spelling & rhythm game accepting any sound
Look & Learn Bus
An I-spy type educational game
Yummy Rainbow Popsicles
Giving children autonomy and a role in the game
Twinkle Twinkle Little Star
Emotional regulation and winding-down
CoComelon: Sing & Play with JJ
CoComelon: Sing & Play with JJ
The challenge.
The game depended on living-room microphones, preschool speech, and a two-screen TV-plus-phone setup. That introduced noisy input, delayed responses, parental handoffs, and children who might speak softly, freeze, or answer with movement instead of words. The experience could not punish any of that. CoComelon's brand requires positive reinforcement, and a child should never feel corrected because hardware failed to hear them. The design challenge was to build a voice-first system with no failure state: clear enough for parents, forgiving enough for toddlers, and responsive enough to feel alive.
Research.
Research reframed the FTUE. Privacy still had to be handled clearly, but caregivers were more immediately worried about their child feeling put on the spot. The risk was not only "is my child being recorded?" It was "what happens if my child does not know what to do?" That changed the onboarding strategy. The first-time experience needed to establish co-viewing, model safe participation, and give parents permission to help before the child encountered a prompt.
“Is my child being recorded?”
“What if my child doesn't know what to do?”
“Is this actually helping my child learn?”
Process.
The design process centered on voice-first interaction loops: listen, invite, wait, affirm, and model. I sketched flows that used visual prompts and animation timing as the primary interface, with narration supporting context instead of carrying every instruction.
- Early FTUE scribbles exploring mic permissions, calibration, and parental hand-off moments
- Low-fidelity Wheels on the Bus flow implementing the listen-first, interact-second structure (listen verse → action verse → location transition → confetti)
- Rough Playroom reward layout showing micro-rewards (toys unlocking) at the end of games
- Mini-game loops showing how kids move from "verse listening" → "participation moment" → "reward" based on the two-pass structure
Building Parental Trust Through Transparent FTUE.
The FTUE was designed as reassurance, not a legal hurdle. Parents needed to understand recording, but they also needed to know they could participate. Co-viewing gave caregivers a clear role: if a child froze, whispered, or looked away, the parent could step in without the moment reading as failure. Calibration doubled as interaction education. It modeled acceptable volume, tested the input path, and showed that quieter households or less verbal children could still complete the experience.
Visual Cues & Timing Over Verbal Prompts.
Each mini-game used a consistent interaction pattern. Narration introduced the subject, animation created the invitation, the system waited for participation, and JJ affirmed the child before moving forward. The prompt lived in the visual layer because preschoolers can follow gesture, rhythm, and character attention faster than spoken instructions. This also protected the recovery model. If a child did not respond, JJ could demonstrate the action again through animation without saying the child was wrong. There was no buzzer, timeout scold, or corrective "try again." The system taught by modeling and kept the emotional state positive.
Choosing 3D Over 2D for Brand Consistency.
3D was both a production decision and a trust decision. Moonbug already had recognizable 3D assets, and reusable environments such as the bus, playroom, and backyard could support multiple mini-games. A 2D approach would have required new scene art for every flow and introduced a style shift from the show. For preschoolers, character recognition is part of usability. JJ had to look like the JJ they already knew before they would confidently respond. Reusing the established 3D language let the team spend more effort on pacing, prompts, and interaction design instead of rebuilding brand familiarity.
Diegetic Navigation System in Unity.
The hub used diegetic navigation instead of a conventional menu. A list of game titles is efficient for adults, but preschoolers respond better to places, objects, and make-believe. JJ's playroom made the game suite feel like entering a familiar world rather than operating an app. Toys functioned as navigation objects. As children played, more toys could appear, giving progression a spatial form and leaving room for future character rooms or themed hubs. The IA was designed to scale without losing the child-facing metaphor.
Iteration.
Early builds prompted too often and too directly. What sounded encouraging in a script could feel like pressure to a shy child in front of the TV. We slowed the response window to a full 2 to 4 seconds, delayed reveals, softened prompt language, and leaned harder on visual cues. Participation improved because the interaction stopped treating hesitation as an error.
Collaboration.
I created storyboards, UX flows, naming conventions, and shot maps in Figma, Notion, and Milanote. PMs used the documentation for planning, and animators used the shot maps to align character motion with voice-input moments. I also built reusable templates for scene structure, interaction points, rewards, and recovery behavior so engineering and animation could ship multiple titles from the same interaction model.
The outcome.
Product Launch
- Multiple interactive titles shipped using the same research-backed UX framework
- The FTUE and interaction loop became a template adopted across new games
- Moonbug approved and praised the educational pacing and tone
Qualitative Wins
- Caregivers found it easy to guide children through the songs
- Kids replayed favorite songs multiple times per session (validating the engagement patterns)
- Internal teams reported improved speed and clarity thanks to reusable templates
Platform Impact
This project established a voice-first, education-aligned UX framework for CoComelon's interactive content on connected TV.
AI workflows.
Voice Design: Casting AI by Ear
The game teaches the child, so the narrator needed one consistent voice across a licensed game. I bypassed standard rules engines to manually generate, curate, and cast the AI voice like a traditional role.
- Generated a wide range of AI voice candidates and line readings.
- Curated takes by ear to isolate authentic, on-brand inflections.
- Stitched the best readings into a "canonical voice" clip.
- Applied the canonical clip universally to prevent voice drift.
AI Storyboarding & Previz
To unblock the animation team and accelerate production, I leveraged AI to generate storyboards and establish visual context earlier in the pipeline.
- Built a shared Milanote board with AI-generated previz to align design and animation.
- Collaborated directly with animators to map out narrative beats, branching logic, and dialogue in one hub.
- Allowed animation to begin pulling previz while game mechanics were still being refined.
- Translated concepts into production-ready docs, ensuring strict adherence to the CoComelon IP.
Impact: Moving the storyboarding phase earlier with AI meant the animation team never had to wait on design bottlenecks.
Reflection.
What I Learned
One extra second of silence changed whether a two-year-old felt heard or talked over. The most important design unit was not the screen; it was the wait state.
What I'm Most Proud Of
Creating a reusable voice-first interaction model grounded in research findings that became the backbone for multiple CoComelon games.
Biggest Challenge
Balancing caregiver trust, preschool cognitive load, and unreliable voice input while keeping every response positive.
What I Would Do Differently
Establish success metrics earlier and implement a lightweight analytics layer to measure attention and participation patterns—quantifying what we observed qualitatively.