Vitals · Confidence through Transparency

Vitals · Confidence through Transparency

Trust in Healthcare is not an add-on, it's a necessity. But nobody ever designed it. Until now

Trust in Healthcare is not an add-on, it's a necessity. But nobody ever designed it. Until now

Trust in Healthcare is not an add-on, it's a necessity. But nobody ever designed it. Until now

Healthcare

Healthcare

XAI

XAI

B2C

B2C

Mobile

Mobile

Project

Project

Vitals

Vitals

Type

Type

Independent Case Study

Independent Case Study

Role

Role

Product Designer

Product Designer

Platform

Platform

Mobile

Mobile

Problem

Problem

Health AI tools give confident answers. None of them tell you how certain they actually are.

Health AI tools give confident answers. None of them tell you how certain they actually are.

Gap

Gap

90% of FDA-approved medical AI devices show no information on how recommendations are generated

90% of FDA-approved medical AI devices show no information on how recommendations are generated

User

User

Wearable device owners who ask AI health questions but can't evaluate the answers with confidence

Wearable device owners who ask AI health questions but can't evaluate the answers with confidence

Solution

Solution

A visual confidence layer that appears before every AI response

A visual confidence layer that appears before every AI response

Problem backed by research

Problem backed by research

Too tired to read? Listen instead.

Too tired to read? Listen instead.

0:00/1:34

AI gives you answers. It never tells you how sure it is.

You’ve probably done it. Asked AI a health question at 11pm. Got an answer that sounded certain. Then wondered if you should trust it. That doubt is rational. The system was never designed to be questioned.

It Started with a Real Consequence

In 2024, a patient used ChatGPT to evaluate symptoms of a mini-stroke. ChatGPT gave an incorrect diagnosis. The patient delayed seeking care. The consequences were life-threatening. This is not an edge case. It is the logical end result of a design failure happening at scale right now.

The Research Confirmed What the Story Suggested

A 2025 MIT study published in NEJM AI found that people cannot distinguish between low-accuracy AI medical responses and responses from actual doctors. They trust both equally. 90% of FDA-approved medical AI devices provide zero confidence information to end users. The problem is systemic, not isolated.

The Research Had a Name for What Was Missing

The field calls it trust calibration. A 2025 longitudinal review found that visual explanations alone do not improve appropriate trust in AI. They only work when accompanied by uncertainty signals and intended-use boundaries. Without them, even well-designed explanations push users toward overtrust. That finding reframed the entire design problem.

The question was never

The question was never

How do we make AI feel more trustworthy?

It was always

It was always

How do we give users the tools to calibrate their own trust in real time?

The Biggest Names in Health AI Are Not Exempt

WHOOP

Describes its Recovery Score internally as "a directional indicator, not a diagnostic tool." Their interface shows a single confident green number. Zero qualifiers. A 2023 peer-reviewed study found zero correlation between that score and any measured physiological variable.

Google Fitbit

Gemini-powered AI coach. Positioned as personalized health guidance, yet early reviews describe confusing advice, weak coaching, and occasional inaccuracies. Even Google notes outputs may be wrong and not for medical reliance. The interface shows little of that uncertainty.

Apple Health+

Apple’s reported AI health coach was scaled back before launch, with features expected to roll into the Health app instead. Even Apple appears cautious about turning health guidance into a standalone AI product. In health, trust and clarity are harder problems than intelligence.

The User

Not a demographic. Not a persona. A psychological state. Someone with a wearable tracking their body around the clock, asking AI questions about training, recovery, sleep, and stress. Caught between the only two options current products give them: trust the AI blindly, or dismiss it entirely.

"I want to trust it. But I have been wrong before following its suggestion. So now I just never know."

— Jatin, 28.

Therefore, VITALS.

Therefore, VITALS.

A confidence layer before every AI health response, showing certainty, reasoning, and what to do next.

A confidence layer before every AI health response, showing certainty, reasoning, and what to do next.

The AI knows. Here is the proof.

The AI knows. Here is the proof.

Color signals certainty. Percentage earns it. Bars explain it. The frame shapes how content is received.

Green means go. 97% means something was actually calculated, not just claimed. The bars don't show status. They show the reasoning behind it.

A user can look at those three signals and independently agree or disagree. That's a fundamentally different experience than being told "You're ready" with nothing behind it. The recommendation sits inside the card, not below it. Attached to the confidence that earned it. For the first time, the user isn't being asked to trust the AI. They're being given the tools to agree with it.

Uncertainty as collaboration, not failure.

Uncertainty as collaboration, not failure.

Most AI products fail here. They either fake confidence or hide behind "consult a doctor." VITALS does neither. "I see something. I am not certain what it means. And I think you deserve to know that."

The amber card signals uncertainty without triggering panic. Two buttons. Two directions. Both empowering. Verify the data the AI is working from. Or give it more context so it can recalculate.

The Data Origin screen is the strongest moment in the product. Not because it shows the heart rate chart or the 98.7% sensor accuracy. Because it makes one distinction no health AI currently makes: The data is not bad. The interpretation is uncertain. Those are two completely different problems. And only one of them is worth worrying about.

Some questions should never be answered by a machine.

Some questions should never be answered by a machine.

The color tells you before anything is read, this is not a failure state. The AI isn't hitting a wall. It's making an intentional, responsible choice. Presence first. Disclaimer second. Presence before disclaimer. Acknowledgment before boundary.

The copy opens with "I hear you", not "I cannot help with that." That sequence matters more than any other decision on this screen.

The crisis number sits alone in its own box. Not a design preference. A safety decision. When someone is in distress, friction kills follow-through. At the bottom, a small note explains why the AI stepped aside instead of answering.

The reasoning behind the decision, made visible. The confidence card concept, applied to knowing your own limits.

Built on a System, Not Just Screens

Built on a System, Not Just Screens

A unified UI guide ensures every screen speaks the same visual language, across both light and dark environments.

Designed decisions I made and rejected

Designed decisions I made and rejected

01

Why a card and not inline text?

Why a card and not inline text?

Why a card and not inline text?

Inline disclosures get skimmed. A visually distinct card breaks the reading pattern and forces a moment of evaluation before the answer lands.

Behavioral insight:

Inline disclosures follow the same visual pattern as the answer itself. The brain skips them. A visually distinct container forces a mandatory pause before the content lands.

Behavioral insight:

Inline disclosures follow the same visual pattern as the answer itself. The brain skips them. A visually distinct container forces a mandatory pause before the content lands.

02

Why show a percentage?

Why show a percentage?

Why show a percentage?

"High confidence" is what every AI already claims. "97%" implies something was actually calculated. It invites scrutiny rather than demanding acceptance.

Behavioral insight:

Vague labels like "high confidence" trigger social proof bias. People accept them without question. A specific number like 97% activates analytical thinking and invites the user to actually evaluate the claim.

03

Why color-coded and not icon-coded?

Why color-coded and not icon-coded?

Why color-coded and not icon-coded?

Color is processed faster than icon meaning. Green, amber, pink communicates trust level in under one second. Icons require learned association. Color is instinctive.

Behavioral insight:

Color communicates state pre-attentively, before conscious reading begins. Icons require learned association. In a health context where a user may already be anxious, instinctive processing beats learned processing every time.

04

Why the card above the answer, not below?

Why the card above the answer, not below?

Why the card above the answer, not below?

Because the frame shapes how content is received. A 97% card before the recommendation makes it feel earned. The same recommendation without the card feels like generic AI output. Sequence is meaning.

Behavioral insight:

Anchoring bias. The first piece of information a person receives frames everything that follows. Confidence before the answer makes the answer feel earned. Confidence after the answer is a footnote nobody reads.

Designed with Doubt. Shipped with Conviction.

Designed with Doubt. Shipped with Conviction.

Two bets shaped this entire project :

  1. A confidence percentage only builds trust if consistently accurate. A single wrong 97% recommendation could destroy the system's credibility faster than no score at all.

  1. 41% certainty is a risky product decision. The safer choice is always to hide uncertainty. VITALS makes the opposite bet, that users shown honest uncertainty will trust the product more over time, not less.

Designed with Doubt. Shipped with Conviction.

Two bets shaped this entire project :

  1. A confidence percentage only builds trust if consistently accurate. A single wrong 97% recommendation could destroy the system's credibility faster than no score at all.

  1. 41% certainty is a risky product decision. The safer choice is always to hide uncertainty. VITALS makes the opposite bet, if users shown honest uncertainty will trust the product more over time, not less.

What Users Said

What Users Said

Qualitative testing across four distinct profiles. Goal: measure cognitive load, emotional response, and interaction accuracy.

Qualitative testing across four distinct profiles. Goal: measure cognitive load, emotional response, and interaction accuracy.

High Confidence · 22, Male

High Confidence · 22, Male

Seeing the certainty score before the text shifted the mental model from passive consumption to informed agreement.

Seeing the certainty score before the text shifted the mental model from passive consumption to informed agreement.

"Seeing the exact confidence score before the answer actually gives me the permission to trust it."

"Seeing the exact confidence score before the answer actually gives me the permission to trust."

Low Confidence · 28, Male

Low Confidence · 28, Male

The amber card did not erode trust. It triggered a collaborative instinct to provide better data rather than abandon the product.

The amber card did not erode trust. It triggered a collaborative instinct to provide better data rather than abandon the product.

"If I see that AI is unsure, my immediate instinct is to clarify my situation to help it get the right answer instead of just abandoning it completely."

"If I see that AI is unsure, my immediate instinct is to clarify my situation to help it get the right answer instead of just abandoning it completely."

Mental Wellness · 16, Female

Mental Wellness · 16, Female

Users in emotional distress skip long AI responses. The pink card broke that pattern and secured attention.

Users in emotional distress skip long AI responses. The pink card broke that pattern and secured attention.

"When I am anxious and venting to an AI, I usually skip reading the long responses. But this card stands out. I would actually read this and take action."

"When I am anxious and venting to an AI, I usually skip reading the long responses. But this card stands out. I would actually read this and take action."

Concept Validation · 52, Female

Concept Validation · 52, Female

A non-technical user with no AI literacy understood the system immediately through the color alone.

A non-technical user with no AI literacy understood the system immediately through the color alone.

"The visual card makes it instantly clear how much I should rely on the answer, which feels much safer than reading a standard plain text response."

"The visual card makes it instantly clear how much I should rely on the answer, which feels much safer than reading a standard plain text response."

How would I measure success

How would I measure success

Trust calibration rate

Trust calibration rate

Do users make different decisions based on confidence level? Do high-confidence responses lead to more action? Do low-confidence responses lead to more verification behaviour? This is the primary signal.

Context update rate

Context update rate

In low-confidence scenarios, what percentage engage with the Update Context flow? A high rate means uncertainty lands as an invitation, not a rejection.

Data Origin engagement

Data Origin engagement

Do users who view that screen report higher trust in the AI response afterwards? This tests whether transparency increases engagement rather than undermining it.

Mental wellness follow-through

Mental wellness follow-through

What percentage of users in scenario 3 tap the crisis line, browse therapists, or open self-help resources? This measures whether the card is a bridge or a visual disclaimer.

Return rate after low confidence

Return rate after low confidence

Do users who see a 41% card come back with follow-up questions more than users who see a 97% card? A higher return rate signals durable trust, not just momentary reassurance.

Future Scope

Future Scope

Confidence history

Confidence history

A log of past AI decisions with their confidence levels and real-world outcomes. That is the deepest form of trust-building possible in any product.

Shared confidence view

Shared confidence view

The ability to share a confidence card directly with a health coach or physician. This bridges consumer health AI and professional healthcare in a way that does not currently exist anywhere.

What I learned

What I learned

When something is opaque, people either trust it blindly or reject it entirely. Good design lives in the middle ground. VITALS is a design argument that AI transparency can be beautiful, scannable, and emotionally intelligent at the same time.

If one person building a health AI product looks at this and thinks "Why aren’t we doing this already?", the design worked.