AI gives you answers. It never tells you how sure it is.
You’ve probably done it. Asked AI a health question at 11pm. Got an answer that sounded certain. Then wondered if you should trust it. That doubt is rational. The system was never designed to be questioned.
It Started with a Real Consequence
In 2024, a patient used ChatGPT to evaluate symptoms of a mini-stroke. ChatGPT gave an incorrect diagnosis. The patient delayed seeking care. The consequences were life-threatening. This is not an edge case. It is the logical end result of a design failure happening at scale right now.
The Research Confirmed What the Story Suggested
A 2025 MIT study published in NEJM AI found that people cannot distinguish between low-accuracy AI medical responses and responses from actual doctors. They trust both equally. 90% of FDA-approved medical AI devices provide zero confidence information to end users. The problem is systemic, not isolated.
The Research Had a Name for What Was Missing
The field calls it trust calibration. A 2025 longitudinal review found that visual explanations alone do not improve appropriate trust in AI. They only work when accompanied by uncertainty signals and intended-use boundaries. Without them, even well-designed explanations push users toward overtrust. That finding reframed the entire design problem.
How do we make AI feel more trustworthy?
How do we give users the tools to calibrate their own trust in real time?
The Biggest Names in Health AI Are Not Exempt

WHOOP
Describes its Recovery Score internally as "a directional indicator, not a diagnostic tool." Their interface shows a single confident green number. Zero qualifiers. A 2023 peer-reviewed study found zero correlation between that score and any measured physiological variable.
Google Fitbit
Gemini-powered AI coach. Positioned as personalized health guidance, yet early reviews describe confusing advice, weak coaching, and occasional inaccuracies. Even Google notes outputs may be wrong and not for medical reliance. The interface shows little of that uncertainty.


Apple Health+
Apple’s reported AI health coach was scaled back before launch, with features expected to roll into the Health app instead. Even Apple appears cautious about turning health guidance into a standalone AI product. In health, trust and clarity are harder problems than intelligence.
The User
Not a demographic. Not a persona. A psychological state. Someone with a wearable tracking their body around the clock, asking AI questions about training, recovery, sleep, and stress. Caught between the only two options current products give them: trust the AI blindly, or dismiss it entirely.
"I want to trust it. But I have been wrong before following its suggestion. So now I just never know."
— Jatin, 28.


Color signals certainty. Percentage earns it. Bars explain it. The frame shapes how content is received.
Green means go. 97% means something was actually calculated, not just claimed. The bars don't show status. They show the reasoning behind it.
A user can look at those three signals and independently agree or disagree. That's a fundamentally different experience than being told "You're ready" with nothing behind it. The recommendation sits inside the card, not below it. Attached to the confidence that earned it. For the first time, the user isn't being asked to trust the AI. They're being given the tools to agree with it.

Most AI products fail here. They either fake confidence or hide behind "consult a doctor." VITALS does neither. "I see something. I am not certain what it means. And I think you deserve to know that."
The amber card signals uncertainty without triggering panic. Two buttons. Two directions. Both empowering. Verify the data the AI is working from. Or give it more context so it can recalculate.
The Data Origin screen is the strongest moment in the product. Not because it shows the heart rate chart or the 98.7% sensor accuracy. Because it makes one distinction no health AI currently makes: The data is not bad. The interpretation is uncertain. Those are two completely different problems. And only one of them is worth worrying about.

The color tells you before anything is read, this is not a failure state. The AI isn't hitting a wall. It's making an intentional, responsible choice. Presence first. Disclaimer second. Presence before disclaimer. Acknowledgment before boundary.
The copy opens with "I hear you", not "I cannot help with that." That sequence matters more than any other decision on this screen.
The crisis number sits alone in its own box. Not a design preference. A safety decision. When someone is in distress, friction kills follow-through. At the bottom, a small note explains why the AI stepped aside instead of answering.
The reasoning behind the decision, made visible. The confidence card concept, applied to knowing your own limits.
A unified UI guide ensures every screen speaks the same visual language, across both light and dark environments.

Inline disclosures get skimmed. A visually distinct card breaks the reading pattern and forces a moment of evaluation before the answer lands.
"High confidence" is what every AI already claims. "97%" implies something was actually calculated. It invites scrutiny rather than demanding acceptance.
Behavioral insight:
Vague labels like "high confidence" trigger social proof bias. People accept them without question. A specific number like 97% activates analytical thinking and invites the user to actually evaluate the claim.


Color is processed faster than icon meaning. Green, amber, pink communicates trust level in under one second. Icons require learned association. Color is instinctive.
Behavioral insight:
Color communicates state pre-attentively, before conscious reading begins. Icons require learned association. In a health context where a user may already be anxious, instinctive processing beats learned processing every time.
Because the frame shapes how content is received. A 97% card before the recommendation makes it feel earned. The same recommendation without the card feels like generic AI output. Sequence is meaning.
Behavioral insight:
Anchoring bias. The first piece of information a person receives frames everything that follows. Confidence before the answer makes the answer feel earned. Confidence after the answer is a footnote nobody reads.

Do users make different decisions based on confidence level? Do high-confidence responses lead to more action? Do low-confidence responses lead to more verification behaviour? This is the primary signal.
In low-confidence scenarios, what percentage engage with the Update Context flow? A high rate means uncertainty lands as an invitation, not a rejection.
Do users who view that screen report higher trust in the AI response afterwards? This tests whether transparency increases engagement rather than undermining it.
What percentage of users in scenario 3 tap the crisis line, browse therapists, or open self-help resources? This measures whether the card is a bridge or a visual disclaimer.
Do users who see a 41% card come back with follow-up questions more than users who see a 97% card? A higher return rate signals durable trust, not just momentary reassurance.
A log of past AI decisions with their confidence levels and real-world outcomes. That is the deepest form of trust-building possible in any product.
The ability to share a confidence card directly with a health coach or physician. This bridges consumer health AI and professional healthcare in a way that does not currently exist anywhere.
When something is opaque, people either trust it blindly or reject it entirely. Good design lives in the middle ground. VITALS is a design argument that AI transparency can be beautiful, scannable, and emotionally intelligent at the same time.
If one person building a health AI product looks at this and thinks "Why aren’t we doing this already?", the design worked.

