Source-linked AI summary
SymptomAI: Toward a Conversational AI Agent for Everyday Symptom Assessment
Joseph Breda, Fadi Yousif, Beszel Hawkins, Marinela Cotoi, Miao Liu, Ray Luo, Po-Hsuan Cameron Chen, Mike Schaekermann, Samuel Schmidgall, Xin Liu, Girish Narayanswamy, Samuel Solomon, Maxwell A. Xu, Xiaoran Fan, Longfei Shangguan, Anran Wang, Bhavna Daryani, Buddy Herkenham, Cara Tan, Mark Malhotra, Shwetak Patel, John B. Hernandez, Quang Duong, Yun Liu, Zach Wasson, Dimitrios Antos, Bob Lou, Matthew Thompson, Jonathan Richina, Anupam Pathak, Nichole Young-Lin, Jake Sunshine, Daniel McDuff
TL;DR
Everyday symptom assessment remains underrepresented in evaluations centered on curated, information-rich cases. SymptomAI evaluated conversational interviewing and differential diagnosis at population scale, finding that its diagnoses outperformed clinician differentials and that organized interviews exceeded user-guided conversations.
Problem
Existing evaluations rarely represent everyday patient conversations or population-level symptom and illness distributions, limiting evidence about conversational AI for real-world symptom assessment.
Method
SymptomAI conducted dynamic symptom interviews and produced differential diagnoses in a national-scale deployment, with clinician validation and naturalistic population data.
Results
SymptomAI differential diagnoses outperformed clinician differentials, while performance degraded when agents failed to elicit additional symptom information.
Takeaways & Limitations
Organized symptom interviews offer an opportunity for higher-quality, safer, and more accurate differential diagnoses than entirely user-guided conversations.
Takeaways & Limitations
Clinician differentials were based on conversations conducted by SymptomAI rather than physician-led interviews, limiting a fully controlled end-to-end comparison.
Abstract
from arXiv · showhide
Language models excel at diagnostic assessments on curated medical case-studies and vignettes, performing on par with, or better than, clinical professionals. However, existing studies focus on complex scenarios with rich context making it difficult to draw conclusions about how these systems perform for patients reporting symptoms in everyday life. We deployed SymptomAI, a set of conversational AI agents for end-to-end patient interviewing and differential diagnosis (DDx), via the Fitbit app in a study that randomized participants (N=13,917) to interact with five AI agents. This corpus captures diverse communication and a realistic distribution of illnesses from a real world population. A subset of 1,228 participants reported a clinician-provided diagnosis, and 517 of these were further evaluated by a panel of clinicians during over 250 hours of annotation. SymptomAI DDx were significantly more accurate (OR = 2.56, p < 0.001) than those from independent clinicians given the same dialogue in a blinded randomized comparison. Moreover, agentic strategies which conduct a dedicated symptom interview that elicit additional symptom information before providing a diagnosis, perform substantially better than baseline, user-guided conversations (p < 0.001). An auxiliary analysis on 1,509 conversations from a general US population panel validated that these results generalize beyond wearable device users. We used SymptomAI diagnoses as labels for all 13,917 participants to analyze over 500,000 days of wearable metrics across nearly 400 unique conditions. We identified strong associations between acute infections and physiological shifts (e.g., OR > 7 for influenza). While limited by self-reported ground truth, these results demonstrate the benefits of a dedicated and complete symptom interview compared to a user-guided symptom discussion, which is the default of most consumer LLMs.
1. Introduction
SymptomAI addresses the limited real-world representativeness of prior conversational-AI evaluations by studying end-to-end symptom interviewing and differential diagnosis with people reporting real health needs. The study combines randomized agent strategies, clinician-validated diagnoses, and wearable data to evaluate performance and explore associations across a large cohort.
- Study design: N=13,917 participants contributed naturalistic symptom conversations paired with recent wearable data, while participant-reported healthcare-provider diagnoses supported benchmarking.The study also used clinical expert validation and assessed SymptomAI differential-diagnosis accuracy across participants who self-reported a diagnosis.
- Motivation and gap: 20-40% diagnostic accuracy reported for traditional symptom checkers highlights a limitation in initial assessments that often precede downstream medical care.These assessments can serve as a primary entry point for downstream medical care.
- Motivation and gap: Prior language-model evaluations use synthetic examples, single-turn question-answering, or highly detailed atypical cases that poorly represent everyday patient conversations and population illness distributions.The paper identifies this lack of representativeness as a barrier to understanding conversational AI performance in real-world contexts.
- Study design: SymptomAI was deployed through Fitbit Labs from June 2025 to April 2026 and randomized participants across five study arms using different symptom-interview prompting strategies.Strategies ranged from structured history-of-present-illness questions to a fully dynamic conversational agent.
- Study contributions: N=1,228 self-reporting participants supported automated differential-diagnosis assessment, and nearly 400 diagnoses across over 500,000 wearable-data days enabled a phenome-wide association study.The top SymptomAI diagnosis was treated as a reference label for the full cohort.
2. Results
SymptomAI performed better when agents actively elicited information through dedicated interviews rather than relying on user-guided conversations. Its differential diagnoses also outperformed independent clinicians, generalized to a broader population, and were associated with wearable-signal shifts around respiratory illness onset.
- Interview strategy: 27.57% higher accuracy resulted from model-guided interviews that elicited more information than the base user-guided strategy.Each experimental arm independently outperformed the base condition, with p=0.003, p<0.001, p=0.003, and p=0.022 for arms 2–5, respectively.
- Interview strategy: 72.07% combined accuracy for flexible, non-predefined questioning was comparable to 77.51% for canonical clinician-defined questions (p=0.235).The findings support autonomous history-taking through natural conversational trajectories without statistically significant loss relative to canonical questions.
- Clinical evaluation: 2.34 odds ratio indicated that clinical raters ranked SymptomAI’s DDx first in over 50% of 517 cases (p<0.001).The result represented a significant preference over the expected one-third chance level and had Cohen’s h=0.42.
- Clinical evaluation: 2.56 median odds ratio favored SymptomAI over baseline clinicians for top-5 DDx accuracy in blinded evaluation (95% CI Cohen’s g [0.18, 0.26], p<0.001).The comparison used 517 cases reviewed alongside ground-truth diagnoses by clinical raters blinded to DDx authorship.
- Clinical evaluation: p<0.001 marked SymptomAI’s significant advantage over clinicians when clinicians were neutral or not confident in their own DDx, while performance was similar on confident cases.This result came from McNemar’s test on conversations stratified by clinicians’ self-confidence.
- Generalization: 1,509 people from a broad US population panel were used to assess whether SymptomAI’s reported-symptom and DDx results generalized beyond Fitbit users.The study also found low demographic distributional shifts in the clinical-evaluation subsample, with D=0.026 (p=0.490).
- Wearable associations: Strong associations linked wearable biosignal shifts with acute respiratory infections, especially during the days preceding SymptomAI engagement.A cohort of 1,546 participants diagnosed by SymptomAI with respiratory infection showed distinct biosignal shifts before symptom reporting.
3. Discussion
The discussion emphasizes SymptomAI’s population-scale evaluation on naturally occurring symptom reports, showing that dedicated information elicitation improves diagnostic assessment and that AI-generated differentials can outperform clinician lists. It also highlights how AI-labeled diagnoses enable large-scale analysis of wearable biosignals preceding illness-related help-seeking.
- Population-scale evaluation: SymptomAI is presented as the first population-scale evaluation of symptom assessment on real users seeking digital medical guidance.Its national deployment captured symptom conversations naturalistically across a broad sample, reflecting how people interact with online assessment tools.
- Interview strategy: Diagnostic performance declined when the agent failed to ask questions that elicited additional symptom information.The discussion identifies this as a limitation of the user-guided approach used by major consumer-facing LLMs.
- Diagnostic performance: SymptomAI differential diagnoses outperformed clinical differential lists across all conditions in blinded preference assessments and comparisons using clinician-provided diagnoses.The findings suggest a substantial opportunity to improve consumer symptom checking by requiring more complete symptom elicitation.
- Diagnostic performance: SymptomAI’s strong performance across diverse conditions suggests that broad distributional priors alone may not explain its assessment of atypical or confounded cases.The discussion notes that these cases can include irrelevant symptoms or multiple conditions reported simultaneously.
- Wearable biosignals: Wearable biosignals involving cardiovascular function, respiration, sleep, skin temperature, and physical activity changed notably in the days before users engaged with SymptomAI.These signals may provide supplemental physiological information about illness progression and potentially support earlier warning or check-ins.
- Wearable biosignals: Immediate access to SymptomAI may improve the accuracy of reported symptom onset by allowing users to describe symptoms while they remain fresh.The discussion contrasts this immediacy with traditional diagnosis, which can be delayed by clinician availability.
4. Limitations
The study’s limitations arise from the ambiguity of symptom-to-diagnosis assessment, uncertainty in second- or third-hand ground-truth labels, and constraints in clinician comparison and symptom-report timing. These limitations may affect evaluation accuracy and the representativeness of reported symptoms.
- Diagnostic ambiguity: 10-15% of clinical encounters result in diagnostic error, underscoring the inherent ambiguity of symptom assessment leading to definitive diagnosis.The analysis is similarly affected by this ambiguity.
- Ground-truth accuracy: Second- or third-hand diagnosis reports may be inaccurate, outdated, mistaken, misrepresented, or invalid despite filtering for reliability.Scaling the evaluation necessitated using these reported diagnoses as labels, creating uncertainty in the ground truth.
- Clinician comparison: Clinician baseline DDx were based on participant–SymptomAI conversations rather than clinician-led patient interviews, potentially limiting the information clinicians could obtain.Clinicians may have sourced different information had they directed the symptom interview.
- Reporting timing: Symptom reports were snapshots in time, and deployment scale prevented controlling reporting frequency and timing relative to symptom development.Some participants reported symptoms before representative indicators developed, whereas others reported obvious indicators informed by years of chronic illness experience.
5. Conclusion
The conclusion presents SymptomAI as an investigational conversational agent whose end-to-end differential-diagnosis performance exceeded that of board-certified clinicians in a population sample. It emphasizes dedicated interviewing and reasoning while noting that the system is not validated for diagnostic use and that dataset access is privacy-controlled.
- SymptomAI is an investigational conversational AI agent for conducting real-world patient interviews and symptom assessments.
- SymptomAI’s end-to-end differential-diagnosis accuracy was superior to that of board-certified clinicians in a population sample.
- The system both inquires about and reasons over patients’ medical histories, using information it elicits naturally to produce a final differential diagnosis.
- The released de-identified conversation and clinician-evaluation dataset is available to qualified researchers under protocols, policies, infrastructure, and controls governing access and use.These measures balance open-data goals with participant privacy and health-data protection.
- SymptomAI is a research prototype, not a medical device, has not undergone regulatory validation, and does not provide clinical diagnoses or confirmed medical status.Its labels, associations, and categories are model-generated for research purposes.
A. Methods … A.3. Agent Arm Designs
SymptomAI was evaluated in a large, randomized Fitbit deployment that compared five language-model symptom-interview designs. The study collected real-world symptom conversations, participant-reported diagnoses, and structured agent behavior ranging from user-driven chat to dynamic interviewing.
- A.1. SymptomAI Study Objective: SymptomAI studied interactions with, and diagnostic accuracy of, agentic and non-agentic language models for addressing symptom-checking queries.The evaluation compared clinician preference and DDx accuracy, stratified by information quality, appropriateness, completeness, and clinical harm.
- A.2. Randomized In-Situ Population Deployment: The study recruited US Fitbit users virtually through an in-app notification between June 2025 and April 17th 2026.The Fitbit Labs environment was used to capture a naturalistic distribution of self-reported symptoms.
- A.2. Randomized In-Situ Population Deployment: Participants consented electronically before interacting with the SymptomAI chat, which asked arm-dependent symptom questions and displayed matching candidate diagnoses.The application then collected a user-experience survey and diagnosis self-report.
- A.2. Randomized In-Situ Population Deployment: 13,917 participants completed at least one conversation, and 1,228 were linked to a patient-reported diagnosis.The initial enrolled Fitbit-user cohort contained 40,000 users.
- A.3. Agent Arm Designs: Five study arms varied the level of instruction given to agents for conducting symptom interviews, with participant geography summarized against the US National Census.Exact prompt language for each arm was provided in Section E.
- A.3. Agent Arm Designs: Arm 1 provided user-driven baseline chat, while Arm 2 used fixed canonical HPI questions covering symptom characteristics and risk factors within six turns.Arm 1 restricted discussion to health topics and required a DDx list; Arm 2 specified location, timing, severity, quality, frequency, continuity, modifiers, and preexisting risks.
- A.3. Agent Arm Designs: Arm 3 flexibly sourced canonical HPI information without fixed questions, whereas Arm 4 dynamically narrowed possibilities and provided a DDx list at every turn.Arm 4 used a prompt optimizer and did not specify particular questions or information targets.
- A.3. Agent Arm Designs: Arm 5 used dynamic interviewing similar to Arm 4 but instructed the agent to provide a DDx only in the final output.Unlike Arm 4, it did not require a DDx list at each conversational turn.
A.4. Evaluation
The evaluation grounded SymptomAI’s differential diagnoses in expert annotation of 517 diagnosed conversations and measured whether reported diagnoses appeared among the top five candidates. After validation, the study used SymptomAI’s top-1 diagnosis as a “silver standard” for wearable biosignal analyses.
- Three clinicians provided expert annotations to establish a baseline clinical differential diagnosis and ground participant symptom-reporting conversations.
- N=517 diagnosed conversations were sampled for human annotation, prioritizing complete conversations that ended with a SymptomAI differential diagnosis.
- Conversations required at least 10 cumulative user words and were filtered to exclude cases deemed impossible for medical assessment from conversation data.
- Accuracy was the percentage of cases where the participant-reported diagnosis appeared among the clinical rater’s top-5 differential-diagnosis candidates.
- After clinical validation, the top-1 SymptomAI differential-diagnosis candidate served as a “silver standard” label for associating diagnoses with wearable biosignal trends.
A.4.1. Clinical Evaluation Tasks … Q9. On a scale of 1–5, how similar are these DDx lists to each other?
The clinical evaluation used two sequential, blinded tasks to compare SymptomAI with clinician-generated differential diagnoses, then assessed diagnosis matching, appropriateness, completeness, confidence, similarity, and potential harms through structured survey questions.
- A.4.1. Clinical Evaluation Tasks: The study comprised two independent evaluation tasks conducted sequentially, with the complete workflow shown in Figure 10.Task 1 generated blinded clinician baselines, while Task 2 ranked and evaluated three differential-diagnosis lists.
- A.4.2. Survey Questions: In Task 1, baseline clinicians independently reviewed conversations with SymptomAI diagnoses redacted and supplied best-effort five-candidate differential diagnoses without external tools.They also reported confidence, information sufficiency, and crucial additional questions they would have asked.
- A.4.1. Clinical Evaluation Tasks: Task 1 distributed conversations round-robin across clinician pairings so each clinician reviewed two-thirds of conversations and every third received each possible pairing.The three pairings were 1-2, 2-3, and 3-1.
- A.4.1. Clinical Evaluation Tasks: In Task 2, a held-out clinical rater blindly ranked and scored two clinician DDx lists and the original SymptomAI DDx after standardizing their format.Supplemental questions were administered in ordered batches alongside relevant transcripts, DDx lists, and diagnoses.
- A.4.1. Clinical Evaluation Tasks: Clinical raters compared each DDx with the self-reported healthcare-provider diagnosis, assessing appropriateness, symptom support, specificity, and whether clearly inappropriate diagnoses could be identified.They also recorded the position of the clinically identical matching candidate, if present, enabling clinically defined accuracy and equivalent-diagnosis matching.
- A.4.2. Survey Questions: The evaluation additionally asked clinicians to assess the potential, likelihood, and severity of harms caused by the SymptomAI experience.The clinical annotation summary included harmful-response assessments alongside DDx and self-reported-diagnosis ratings.
- Q9. On a scale of 1–5, how similar are these DDx lists to each other?: Similarity follow-ups classified whether a DDx contained all, most, or some reasonable candidates, or had major candidates missing, before revealing the self-reported diagnosis.The self-reported-diagnosis assessment followed disclosure of the ground truth.
Q14. Please provide details. … Q41. I know how to find helpful health resources on the Internet.
The evaluation materials assess DDx support, misclassification, ranking, and safety, while the auxiliary survey collects symptom, diagnosis, demographic, healthcare, internet-use, and health-literacy information. The survey also measures participants’ confidence, perceived usefulness, and ability to evaluate online and LLM-derived health information.
- Q20. Please provide any comments on unexpected mappings or other issues.: Safety assessment asked evaluators to estimate expected harm if users accepted the interaction as true and the likelihood that the information would cause harm.Additional comments could address unexpected mappings, the conversation, self-reported diagnosis, DDx lists, or the harm assessment.
- A.5. Auxiliary Study Objective: 1,509 general-population participants joined an auxiliary study designed to address population skew from Fitbit users, with survey responses reformatted into hypothetical SymptomAI chat logs.Combining auxiliary and SymptomAI data produced 2,737 participants with self-reported diagnoses, including 962 assessed participants; the auxiliary data evaluated broader-population generalizability.
- A.5. Auxiliary Study Objective: The auxiliary study recruited US adults through an IRB-approved Toluna panel, covered 39 primary-care symptom categories, and collected provider diagnoses after structured symptom reporting.Participants provided consent, received $4 USD, reported health events and internet-seeking behavior, and had de-identified data stored with restricted access.
- A.6. Auxiliary Study Survey Questions: The survey captured primary and associated symptoms across a broad checklist, alongside gender, ethnicity, race, education, and insurance characteristics.Primary symptoms included categories such as abdominal pain, cough, fever, headache, nausea/vomiting, numbness, and pain in foot, leg, or arm.
- Q12. What type of health insurance do you have?: Participants described symptom onset, triggers, location, severity, duration, frequency, quality, progression, modifiers, risk factors, functional effects, and the provider or laboratory diagnosis.The questionnaire also asked about diagnostic certainty, the diagnosing professional, and medications taken during the illness episode.
- Q29. Which internet web search engine did you use?: Participants reported whether they used web searches or LLM chatbots for symptom information, which services they used, and what queries they entered.Examples included Google, Bing, Yahoo, DuckDuckGo, Google Gemini, OpenAI ChatGPT, and Perplexity.
- Q36. What steps did you take to manage your symptoms AFTER receiving your diagnosis.: The questionnaire measured perceived internet usefulness and importance, knowledge of available resources, ability to find and use health information, and confidence evaluating its quality.It also assessed participants’ reported ability to use LLM chatbots for health questions and make health decisions from internet information.
A.7. Clinical Evaluation Statistics
The clinical evaluation used exact paired tests and resampling to compare SymptomAI with clinicians, while separate tests assessed ranking preferences, prompt strategies, and stratified performance. Representativeness analyses compared demographic and diagnostic-category distributions across cohorts.
- Clinician comparison: 1,000 resampling iterations sampled one clinician baseline accuracy per case and used exact McNemar tests, with significance based on the median p-value.The resampling was performed with replacement across cases.
- Clinician ranking: Clinician ranking preferences were tested against random-choice probability p = 0.333 using an exact two-sided binomial test.Preference magnitude was quantified with Odds Ratios and Cohen’s h effect sizes using arcsine transformation.
- Prompt strategy: Top-5 accuracy for each prompt strategy was compared with the Base prompt arm using Chi-square tests of independence.Comparisons covered individual strategies and pooled conceptual categories.
- Stratified evaluation: Exact McNemar’s tests for paired nominal data assessed SymptomAI outperformance over baseline clinicians within each conversation-quality and confidence stratum.
A.8. Wearable Biosignals Analysis
The analysis linked daily wearable biosignals collected before and after SymptomAI conversations to assigned diagnoses using a temporal phenome-wide association study. It tested nine biosignals, excluded skin temperature for lacking significant associations, and separately visualized aligned timeseries for acute infection onset.
- Phenome Exploration: Wearable cardiovascular, sleep, respiratory, temperature, activity, and stress metrics were collected for 30 days before and 7 days after SymptomAI conversations.These data were analyzed as daily downstream metrics.
- Phenome Exploration: A temporal phenome-wide association study evaluated associations between wearable biosignal presentation and SymptomAI-assigned diagnoses.Continuous predictors and covariates were standardized before modeling to produce directly comparable odds ratios.
- Phenome Exploration: 9 biosignals were tested, while skin temperature showed no significant associations and was excluded from the results.The tested signals included resting heart rate, RMSSD heart rate variability, sleep respiratory rate, sleep wake minutes, total sleep minutes, non-REM heart rate, skin temperature during sleep, active minutes, and daily steps.
- Acute Infection Onset Analysis: Acute infection onset was investigated by visualizing aggregated, temporally aligned daily timeseries for the infectious disease cohort versus the remaining population.This analysis further examined the timing of acute infection-related changes.
B. Supplemental Experiments · B.1. Performance of LLMs on Existing Diagnosis Benchmarks
SymptomAI’s use of Gemini was supported by evaluations on existing diagnostic benchmarks. Gemini performed highly on curated symptom-checker vignettes and complex diagnostic cases, while performance decreased on the study’s own data.
- B.1. Performance of LLMs on Existing Diagnosis Benchmarks: Gemini was evaluated on existing diagnostic benchmark datasets to justify its selection as SymptomAI’s base model.The evaluation included online symptom-checker case vignettes and more complex diagnostic cases.
- B.1. Performance of LLMs on Existing Diagnosis Benchmarks: The benchmark evaluation covered curated case vignettes designed to assess online symptom checkers.These vignettes were attributed to Aissaoui Ferhi et al. (2024).
- B.1. Performance of LLMs on Existing Diagnosis Benchmarks: The evaluation also included more complex diagnostic cases from McDuff et al. (2025).This provided a second existing benchmark setting beyond symptom-checker vignettes.
- B.1. Performance of LLMs on Existing Diagnosis Benchmarks: Current language models performed well on both benchmark types described in the supplemental evaluation.The passage characterizes performance on these existing datasets as strong.
- B.1. Performance of LLMs on Existing Diagnosis Benchmarks: Gemini showed high performance on the curated datasets but decreased performance on the study’s own data.The supplied passage introduces this contrast without providing the corresponding numerical results.
- B.1. Performance of LLMs on Existing Diagnosis Benchmarks: Table 2 reports Gemini’s performance against the existing benchmark datasets and the study data.It places curated case reports and symptom-checker vignettes alongside the study’s data.
B.2. Auto-rater Consistency Across Studies
An LLM auto-rater extended evaluation beyond clinician-reviewed conversations and was validated against 517 clinically rated cases. Its performance was consistent across the SymptomAI and auxiliary populations for both SymptomAI and clinician DDx.
- Auto-rater validation: An LLM auto-rater using Gemini 2.5 Pro extended evaluation beyond the conversations manually reviewed by clinicians.The auto-rater prompt is provided in Section E.
- Auto-rater validation: 517 clinically rated SymptomAI conversations were used to validate auto-rater alignment with clinician labels for exact positional and top-5 matches.Alignment was represented in confusion matrices in Figure 13a and Figure 13b.
- Auto-rater validation: Most misalignments arose when auto-raters identified a match but clinicians were more conservative.The supplied passage truncates the remainder of this comparison.
- Cross-study consistency: Auto-rated performance was consistent across the SymptomAI and auxiliary populations for both SymptomAI DDx and baseline clinician DDx.This consistency qualitatively supports generalizability across populations and robustness of auto-rater performance.
B.3. Auto-rated Performance of Gemini DDx Across Models … E. Prompts
SymptomAI’s DDx performance was broadly consistent across Gemini variants and demographic groups, while the study used a custom illness taxonomy, category-level analyses, conversation examples, and multiple structured prompting strategies. The prompts emphasized targeted symptom interviews, bounded questioning, differential generation, and auto-rater evaluation.
- B.3. Auto-rated Performance of Gemini DDx Across Models: Accuracy was specific to Gemini 2.0 Flash, while cross-model evaluation used DDx from 2,737 diagnosed conversations and found only slight improvements with model recency and size.The common evaluation set combined SymptomAI and auxiliary-study conversations; prompt comparisons used 1,228 diagnosed SymptomAI participants.
- B.4. DDx Accuracy Across Demographics: DDx performance increased subtly among older adults, female participants, and participants with advanced degrees, while medical literacy showed no performance difference.Self-reported online health-resource literacy was associated with slightly increased performance.
- C. Diagnosed Condition Taxonomy: The study’s custom diagnostic taxonomy organized illnesses into 12 parent categories and 42 granular categories, coalescing participant-reported diagnoses with Gemini.The taxonomy provided a coarser data view than ICD-10 and Phecode and was supplied to Gemini for categorization.
- C.1. DDx Accuracy Across Illness Categories: SymptomAI’s auto-rater measured top-1 and top-5 DDx accuracy across illness categories using diagnosed participants from the SymptomAI and auxiliary studies.The analysis included the study-arm-1 subset for posterity.
- D. Conversation Examples: A poor-quality conversation failed to gather sufficient information for the participant-reported Hashimoto’s disease diagnosis.The example began with fatigue and headache and elicited general symptom details without supporting the reported diagnosis.
- D. Conversation Examples: A high-quality conversation gathered symptom timing, frequency, severity, and associated urinary details, making urinary tract infection the top-1 candidate.The participant-reported diagnosis was urinary tract infection.
- E. Prompts: Dynamic prompts prioritized specificity, used symptom denials to eliminate possibilities, explored less frequent explanations when appropriate, acknowledged uncertainty, and applied an auto-rater to locate the reported diagnosis in the differential.The auto-rater returned the matching differential position or -1 when the reported diagnosis was absent.