What users of health chatbots struggle with — three recurring breakdowns across 59 apps and over 15,000 reviews
A study analysing over 15,000 user reviews across 59 AI health chatbot apps. Breakdowns recur in three forms, and the most negative experiences were associated with concerns about privacy and security.
Paper overview (our summary)
- Field (arXiv category)cs.HC(+1)
- AuthorsMuhammad Hassan, Ramazan Yener, Ece Gumusel, et al. (4)
- Submitted2026-06-25
- arXiv ID2606.27302v1
Key points
- Over 15,000 user reviews across 59 AI healthcare chatbot apps were analysed.
- Three recurring breakdowns were identified: access barriers and service unreliability, user experience and interaction quality, and billing and customer support.
- None of the three bears directly on the medical correctness of answers.
- Privacy and security concerns were associated with the most negative experiences.
- The authors treat these chatbots as information infrastructure rather than individual products, drawing attention to failures of access, usability and trust.
1Looking at use rather than performance
AI in health is usually assessed on whether answers are correct. What users actually struggle with, though, may not be the content of answers. This study analyses over 15,000 reviews across 59 apps — a record written from the user side.
Reviews are not answers to questions a researcher set. They contain what users chose to write. What gets articulated as a problem therefore yields a different picture than a survey built on prepared questions.
2The three breakdowns
| Form of breakdown | Content |
|---|---|
| Access barriers and service unreliability | It cannot be used at all, or does not run steadily |
| User experience and interaction quality | It is hard to use, or the exchange does not connect |
| Billing and customer support | Problems around charges and handling of enquiries |
None of the three bears directly on the medical correctness of answers. Whether it can be used, whether it is usable, and whether it works as a commercial dealing. Breakdown occurs in a layer distinct from technical performance.
3What accompanied the worst experiences
The authors report that privacy and security concerns were associated with the most negative experiences. Health information presumes trust in how it is handled. Where that trust is damaged, capability however good goes unused.
4Framing them as infrastructure
The framing of the paper treats these chatbots not as individual products but as information infrastructures. For infrastructure, not stopping, being reachable by anyone and being trustworthy come before sophistication of function. This site also covers many medical device recall records, where a software defect is likewise remedied only once an update actually reaches people.
Arriving, and being usable, are conditions separate from performance. This article is our own summary of public research information and is neither a warranty of its contents nor medical advice.
Why it matters
The layer where users actually struggle can sit apart from technical performance. In domains resting on trust, concerns about privacy shape the assessment of the whole experience.
FAQ
Why does answer correctness not appear among the breakdowns?
What does framing them as infrastructure mean?
Sources (primary)
Source: arXiv (descriptive metadata is CC0 public domain). Summaries are our own; see arXiv for the original text and PDF.
- arXiv abstract page (original, official)
- PDF (arXiv)
- arXiv ID: 2606.27302