Ruby Wu

RUBY WU ©
2020—2026.

LooperRoom AI

An AI companion designed to know its limits

An AI companion designed to know its limits

Client

LooperRoom

Year

2025

Team

Ruby Wu / Sole Product UI/UX Designer

Akash Patel / Developer

Aniket Sinha / Developer

Angelica Toledo / Developer

Chris Crawford / Developer

Stella Sakhon / Project Manager

OVERVIEW

An AI mental health companion that knows exactly where its job ends.

LooperRoom connects patients and clinicians between therapy sessions, the gap where 35-50% of patients drop out. I joined as the sole product designer when the product was a working but broken beta. The team had a project manager coordinating delivery, but no one owning product strategy, so I stepped into that gap. I defined what this AI is allowed to do, designed its conversational behavior, rebuilt the clinician workspace, and validated every major call through a structured pilot study before shipping. This case is about owning an AI product end to end: from deciding what an autonomous system should never do, to proving with evidence that users actually trust and adopt it. The product is now used by real patients and clinicians, won the NY Product Design Silver Award, and became a Fast Company Innovation by Design 2025 Finalist.

SCOPE OF OWNERSHIP

What I owned solo, and what I led with alignment.

I owned solo

AI personas and behavior rules · Conversational design principles · Information architecture · UI & interaction redesign

I led with alignment

Engineering implementation specs · Business positioning as clinician extension · Validation methodology

PROBLEM

Patients don't quit therapy in the therapist's office. They quit in the weeks between visits.

Between sessions, symptoms worsen in silence and 35-50% of patients drop out. LooperRoom was built to close that gap, but the beta I inherited was designed by engineers for engineers. Long fragmented forms, unreadable chat, interfaces that looked like internal tools.

But our users are not typical users. One side is people at their most vulnerable, and the other is clinicians with no time. Here, a design failure means someone closes the app on the day they most needed it.

Between sessions, symptoms worsen in silence and 35-50% of patients drop out. LooperRoom was built to close that gap, but the beta I inherited was designed by engineers for engineers. Long fragmented forms, unreadable chat, interfaces that looked like internal tools.

These are not typical users. One side is people at their most vulnerable. The other is clinicians with no time. Here, a design failure means someone closes the app on the day they most needed it. Since no one owned product thinking, so I did.

RESEARCH & PROBLEM BREAKDOWN

Dropout looks like one problem, but it's actually four.

I broke dropout down by the two sides of the product and kept asking why until real problems surfaced:

01

Patient side

the AI companion felt cold and mechanical

02

Patient side

crisis support was a buried phone number

03

Clinician side

UI forced clicking through dropdowns when the real workflow is scanning

04

Clinician side

nothing patients shared ever reached clinicians in useful form for diagnosis

STRATEGY: DEFINING WHAT THE AI SHOULD BE

Before designing a single screen, we had to decide what this AI is allowed to be.

Too passive and it helps no one. Too capable sounding and people treat it like a therapist, which is dangerous and a regulatory nightmare, so we drew the lines below:

LooperRoom is a bridge, not a destination

LooperRoom is a bridge, not a destination

It keeps people company between sessions and hands everything useful back to the humans who do treat.

It keeps people company between sessions and hands everything useful back to the humans who do treat.

Business strategy of clinician extension

Business strategy of clinician extension

An AI therapist positioning makes clinicians opponents. A clinician extension makes them our distribution channel, and every point of dropout prevented is direct revenue retention for clinics.

An AI therapist positioning makes clinicians opponents. A clinician extension makes them our distribution channel, and every point of dropout prevented is direct revenue retention for clinics.

Two AI personas for two user groups

Two AI personas for two user groups

Lulu for patients, warm and never judging ; Melo for clinicians, organized and data driven

Lulu for patients, warm and never judging ; Melo for clinicians, organized and data driven

The persona split was challenged, but the transcripts settled it.

The PM's position was reasonable: one capable AI should serve both sides, and one persona is cheaper to build. But in early pilot sessions, a single AI kept confusing its own role, slipping into analyst mode mid conversation and summarizing symptoms back at patients, exactly the diagnostic tone we had banned. Companionship and analysis are different jobs with different rules. That is when the Lulu and Melo split stopped being a design preference and became a safety requirement.

Skeptical clinicians became our referral channel.

Clinicians came in skeptical, reasonably so, since an AI in mental health reads as a replacement threat. The extension framing plus Melo's human in the loop design moved them along an observable path: skeptical at intro, willing to pilot with a few patients, then actively referring patients to use Lulu between sessions.

PRIORITIZATION: DESIGNING UNDER UNCERTAINTY

You can't prioritize an AI product the normal way. Half of what you need to know can't be known upfront.

The core of an AI product is probabilistic, so instead of a standard impact vs effort matrix, I mapped every design challenge onto the Uncertainty Matrix, sorting them by what kind of knowledge, or lack of it. Each quadrant calls for a different design method.

JUDGMENT FRAMEWORK

One question sat under every call: how reversible is the harm, and how vulnerable is the user right now?

Looking back at the decisions in this project, they all ran through the same two axes. High vulnerability with irreversible harm gets deterministic design and zero improvisation, which is the crisis switch. High vulnerability with reversible harm gets tested iteration, which is Lulu's persona work. Low vulnerability with reversible harm ships directly, which is latency honesty. This framework was not on my wall at the time. It is what my decisions turn out to have in common, and it is how I would brief the next designer on day one.

DECISION 1: TEACHING LULU TO LISTEN

Our AI companion had one job, and she was failing at it… she sounded like a robot.

Early testing showed Lulu repeating the same patterns, never reflecting what users said. Talking to her felt like filling out a form. For a person at their lowest point, that means closing the app. My diagnosis is: Lulu was collecting data, not keeping company. With the PM, I rebuilt her conversational rules on two principles, drawn from motivational interviewing, a clinical framework built on one insight: people keep talking when they feel heard.

01

Reflect before you ask

Reflect what the user said, confirm the understanding, then ask an open question.

02

Hard language boundaries

No diagnostic phrasing. No advice. Lulu only reflects, confirms, and asks questions.

Then we put the redesign in front of real users instead of assuming it worked. The full method and results are in the Impact section, but the short version: people stopped feeling interviewed and started feeling heard.

Before

After

DECISION 2: AI INTERACTION PATTERNS

Designing an AI product means designing behavior, not just UIs.

Four moments show what that means:

01

The thinking state

Models have latency. Lulu floats up and speaks the wait in her own voice: "remembering our conversations," never "processing data." Latency honesty is trust design.

02

The crisis switch

When high risk signals appear, generation stops and a deterministic flow takes over, so resources are one tap away, each explained in plain language.

03

Consent as conversation

Who sees your data and why, explained inside onboarding in the same warm voice, not in a legal page.

04

The human in the loop

Melo summarizes trends and shifts, but never concludes. The clinician makes the judgment based on AI insights.

DECISION 3: THE CLINICIAN WORKSPACE

Not every problem in an AI product is an AI problem.

The original interface made clinicians pick patients from a dropdown, one at a time. But their real workflow is scanning: who messaged, who is trending down, who needs me today.

I rebuilt it as a message list, the pattern every clinician knows from their phone. All patients at a glance, latest activity surfaced, instant search, so the UI is shaped around a real clinician workflow.

Before

After

DECISION 4: CARRYING EACH CLINICIAN'S VOICE

One Lulu was never going to sound like every clinician.

Clinician interviews showed that therapeutic style is a structured variable, not personal magic: CBT oriented clinicians move fast to coping actions, person centered ones validate emotion first, some have rituals like always asking about sleep. A patient who hears one approach in session and another from Lulu starts wondering which to trust.

So I designed Clinician Voice Calibration: clinicians mark which sample replies sound like them, the system builds a style profile applied to all their patients, and Melo periodically checks whether Lulu still sounds right. Carrying each clinician's standard across every touchpoint is not personalization. It is clinical consistency.

01

Voice calibration & clinician profile

Clinicians mark which sample replies sound like them, no rule writing needed. The system turns those choices into a style profile applied to every patient under their name.

02

Melo check-in

Calibration is never one and done. Melo periodically shows recent Lulu replies and asks "do these still sound like you," keeping each clinician in charge of their own voice.

IMPACT

I didn't assume the redesign worked. I measured it.

We ran a structured pilot study with target users before shipping the conversational redesign. I used three methods that each covers the others’ blind spots:

Think aloud sessions

Think aloud sessions

Users talked with Lulu in simulated conversations while narrating their thoughts and feelings out loud, so we could catch the exact moments a reply felt mechanical.

Users talked with Lulu in simulated conversations while narrating their thoughts and feelings out loud, so we could catch the exact moments a reply felt mechanical.

Pre/post self assessment

Pre/post self assessment

Participants rated their emotional state, sense of being understood, and trust toward Lulu before and after each conversation.

Participants rated their emotional state, sense of being understood, and trust toward Lulu before and after each conversation.

Participants rated their emotional state, sense of being understood, and trust toward Lulu before and after each conversation.

Conversation log analysis

Conversation log analysis

We reviewed transcripts against the reflect, confirm, ask principle to verify the dialogue rules actually held in real conversations, not just in the spec.

We reviewed transcripts against the reflect, confirm, ask principle to verify the dialogue rules actually held in real conversations, not just in the spec.

And below are the results we got:

What we found

What we found

Participants reported a clear rise in feeling understood rather than interviewed, and in their willingness to continue talking with Lulu. The redesigned crisis flow also tested well: participants could articulate which resource fit which situation, something the old phone number list never achieved.

Participants reported a clear rise in feeling understood rather than interviewed, and in their willingness to continue talking with Lulu. The redesigned crisis flow also tested well: participants could articulate which resource fit which situation, something the old phone number list never achieved.

What adoption confirmed

What adoption confirmed

The clearest evidence was behavioral. Of the 12 pilot clinicians, 9 moved from observing to actively referring patients to use Lulu between sessions within the first 4 weeks. Patients holding at least 3 conversations a week rose from 42% to 63% over the same pilot period. The skeptic to referrer shift is the adoption story: the group most likely to oppose AI mental health tools became the ones recommending it.

The clearest evidence was behavioral. Of the 12 pilot clinicians, 9 moved from observing to actively referring patients to use Lulu between sessions within the first 4 weeks. Patients holding at least 3 conversations a week rose from 42% to 63% over the same pilot period. The skeptic to referrer shift is the adoption story: the group most likely to oppose AI mental health tools became the ones recommending it.

Recognition

Recognition

LooperRoom AI won the NY Product Design Silver Award and Fast Company Innovation by Design 2025 Finalist!

LooperRoom AI won the NY Product Design Silver Award and Fast Company Innovation by Design 2025 Finalist!

One honest note on metrics: dropout reduction itself is a long horizon metric that requires longitudinal tracking beyond the pilot stage. What we established first were its leading indicators, such as patients willing to keep talking, and clinicians willing to keep recommending.

REFLECTION

The most important design decisions in AI products don't always happen on the UI.

Next step if resources allowed, I would redesign Melo from presenting data to surfacing insight, flagging when a patient's pattern shifts. But the deeper lesson is, the real design work is the lines you draw for the AI: what it can say, when it stays silent, and when it hands off to a human. Draw them right and AI becomes an assistant people trust, but draw them wrong it becomes a risk people avoid. And this logic is not specific to mental health. Any AI trusted to act on someone's behalf, in their voice, at their standard, needs the same lines drawn with the same care. That is the design work I want to keep doing.