A health AI met people in their own words
Medical AI is usually tested on neat case notes. Real people are not neat case notes.
They forget dates, use ordinary language and start with the detail that worries them most. Google researchers wanted to see what happened when an AI had to work through that mess in conversation.
Their experimental system, SymptomAI, ran inside Fitbit Labs from June 2025 to April 2026. A total of 13,917 consenting participants across the United States described symptoms to one of five agents built on Gemini 2.0 Flash.
Each agent could ask questions and then produce a differential diagnosis: not one final answer, but a list of possible causes and suggested next steps. The five versions differed in how tightly their interviews were structured.
This was a research study. The AI's output was not sent to a participant's doctor and had no role in their treatment. Google has not announced SymptomAI as a consumer diagnostic product.
The headline result needs its full comparison
The study did not have a confirmed medical diagnosis for every conversation. Of the full group, 1,228 participants later reported a diagnosis they said they had received from a healthcare provider.
Researchers selected 517 of those cases for a detailed clinical review. Board-certified clinicians read the transcript collected by SymptomAI, wrote their own lists of possible diagnoses and then rated anonymised lists from the AI and other clinicians.
On that comparison, the current paper reports that SymptomAI was more likely to include the participant-reported diagnosis in its top five. It gives a median odds ratio of 2.56 and a statistical result below the paper's significance threshold.
The clinicians also ranked the AI list first in 53.3% of comparisons. A clinician list was ranked first in 23.5%. Those numbers are striking, but they do not mean an AI independently diagnosed patients better than doctors in a clinic.
The doctors were working from an interview run by the AI. They could not see the patient, check records, examine them or ask different follow-up questions. The study compares diagnosis lists built from the same limited transcript.
Asking better questions did real work
The five study arms make the experiment more useful than a simple model leaderboard.
One version behaved like a basic chatbot and relied on the participant to supply the useful details. Others followed standard history-taking questions or were allowed to choose new questions as the conversation developed.
Every agent-led interview strategy significantly outperformed the basic user-led condition, according to the paper. In plain terms, the AI did better when it took responsibility for finding missing information before answering.
That matters beyond health. A model can know a great deal and still fail if the interface lets a vague first message become the whole brief. Good agent design includes knowing when the input is not enough.
It also creates a harder safety question. The system choosing the next question is shaping the evidence it will later use. An omitted question can narrow the diagnosis before anyone notices.
Wearable data added a second, softer signal
The researchers also used SymptomAI's top diagnoses as labels for a much larger analysis of Fitbit measurements.
They examined more than 500,000 days of wearable data across nearly 400 conditions. Conversations labelled as acute infections often lined up with changes in heart rate, breathing, skin temperature and sleep around the time symptoms were reported.
That pattern is interesting because it links what a person said with a separate physiological signal. It could help researchers study the early shape of an illness at a scale that manual clinical labelling would make difficult.
It is not independent proof that every AI diagnosis was correct. The labels for the full cohort came from SymptomAI itself, and an association between a wearable signal and an AI label can carry the model's mistakes forward.
The paper presents this as a research route: symptom conversations, possible diagnoses and passive measurements studied together. It does not establish a clinical test.
Promising evidence, not a medical handover
Several facts are solid. The study was large, participants used natural language, the comparison was blinded and active questioning beat a basic chat format. The paper also reports strong performance in a separate panel of 1,509 people beyond the Fitbit-user cohort.
The authors' larger claim is that a dedicated conversational agent can gather useful clinical context and support symptom assessment at population scale. The experiment gives that claim real weight.
The uncertainties are just as important. The paper is a preprint. Most reference diagnoses came through participants rather than medical records. Only a fraction of conversations entered the clinician comparison, and those clinicians inherited the AI's interview instead of conducting their own.
There is no evidence here about emergency triage, liability, deployment across health systems or what happens when a user treats a plausible list as a final answer. Those questions need prospective clinical work, not another benchmark.
For now, SymptomAI is best read as evidence that better questioning can make a health chatbot more useful. It is not evidence that a chatbot can replace the person who examines you.
Sources
- Google Research — SymptomAIPrimary research announcement published 22 July 2026. Source for the study design, reported findings, intended research use and stated limitations.
- Breda et al. — SymptomAI preprintPrimary 55-page preprint. Source for participant counts, clinical comparison, prompt strategies, wearable analysis, consent details and limitations.
- Breda et al. — SymptomAI full paperPrimary full-text paper used to verify the current reported odds ratio, subgroup sizes, evaluation procedure and the fact that outputs did not affect clinical care.



