Source-linked AI summary
Do AI chatbots find what experts would? Effects of model, user role, and sample size on study retrieval for medical questions
Qingfang Liu, Qiao Jin, Joe D. Menke, Thorsten Kahnt, Zhiyong Lu
TL;DR
The quality and determinants of clinical-study retrieval by general-purpose chatbots remain insufficiently characterized, particularly across models and user roles. This study benchmarks three chatbots against Cochrane review studies and finds that retrieval varies by model and role, with larger sample size the only significant predictor after adjustment.
Problem
General-purpose chatbots’ ability to identify clinical studies and the effects of user role on retrieval remain insufficiently characterized.
Method
The study evaluated three chatbots on questions from 20 Cochrane reviews, using patient, clinician, and researcher prompts repeated four times and benchmarked against primary-study sets.
Results
GPT-5.5 had the highest recall, researcher framing outperformed clinician and patient framing, and larger sample size remained the only significant retrieval predictor after adjustment.
Takeaways & Limitations
Chatbot evidence retrieval varies by model and user role, while retrieval favors larger-sample studies and requires accurate application of PICO-S criteria.
Takeaways & Limitations
The exploratory test set comprised only 20 reviews from two Cochrane issues, limiting the findings’ breadth and generalizability.
Abstract
from arXiv · showhide
Large language model (LLM) chatbots are increasingly used to answer clinical questions with citations to relevant clinical studies. Prior research has largely focused on citation fabrication, leaving a gap in evaluating the quality of retrieved studies and the factors driving their selection. In this study, we evaluated three general-purpose LLM chatbots: Claude Sonnet 5, Gemini 3.1 Pro, and ChatGPT GPT-5.5. We prompted the models with clinical questions adapted from 20 review questions in Issues 6 and 7 of the 2026 Cochrane Database of Systematic Reviews, simulating patient, clinician, and evidence-synthesis researcher roles. Each chatbot was queried under each user role with four independent repetitions, yielding 720 responses. Each chatbot was asked to support its answers with primary clinical citations, which we benchmarked against the included and excluded study sets of the Cochrane reviews. On average, a chatbot response retrieved 39.2% $\pm$ 29.8% of Cochrane included studies, while citing 5.0% $\pm$ 9.4% of excluded studies. Recall of Cochrane included studies varied significantly by model and user role. ChatGPT achieved higher recall than Claude or Gemini (63.1% $\pm$ 29.5% vs. 37.0% $\pm$ 23.8% vs. 17.3% $\pm$ 13.1%; $p=2.0\times10^{-5}$). The researcher role yielded higher recall than the clinician or patient roles (42.8% $\pm$ 30.8% vs. 38.6% $\pm$ 28.9% vs. 36.1% $\pm$ 29.3%; $p=2.0\times10^{-5}$). Controlling for publication year, citations per year, and open-access status, sample size was the only independently significant predictor of retrieval (odds ratio 1.80 per 1-unit increase in log sample size, 95% CI 1.37-2.36, $p=2.34\times10^{-5}$). These findings suggest that while LLM chatbots can retrieve some studies identified by expert reviewers, their performance varies by model and user role, and they exhibit a bias toward clinical trials with larger sample sizes.
1. Introduction
As people increasingly use chatbots for medical evidence retrieval, the ability of general-purpose systems to identify relevant primary studies—and the effect of user role—remain insufficiently characterized. This study evaluates three contemporary chatbots across Cochrane-derived medical questions and finds that GPT-5.5 achieved the highest recall of included studies.
- Motivation: Growing medical chatbot use has increased requests to retrieve evidence from existing literature, while the ability to identify and cite relevant clinical studies remains less well characterized.Prior concerns about fabricated or hallucinated citations further motivate evaluating retrieval quality.
- Prior evidence: Prior evaluations generally report low recall when chatbot-cited primary studies are compared with systematic-review included-study lists.Reported precision and citation fabrication vary across systems and evaluation methods.
- Open question: Whether patients, clinicians, and evidence-synthesis researchers obtain different retrieved evidence remains an open question.Prior work shows that role-play, persona information, and question framing can alter model behavior, but retrieval effects by self-identified user role had not been tested.
- Study design: The study evaluated GPT-5.5, Claude Sonnet 5, and Gemini 3.1 Pro on medical questions adapted from 20 Cochrane review topics.Prompts represented patient, clinician, and evidence-synthesis researcher perspectives, with each role–topic combination repeated four times.
- Contribution: GPT-5.5 achieved the highest recall of Cochrane included-study sets among the three evaluated chatbot systems.All systems were equipped with the ability to perform web searches.
2. Methods
The study evaluated three general-purpose chatbots on clinical questions drawn from 20 recent Cochrane intervention reviews, using standardized prompts for three user roles. Retrieved studies were assessed against Cochrane included and excluded studies, counting uniquely identifiable primary studies as the analysis unit.
- Review sample: 20 Cochrane intervention reviews from Issues 6 and 7 of the 2026 Cochrane Database of Systematic Reviews formed the study sample.The reviews covered diverse clinical areas and were the two most recent complete issues available at the experiment’s time.
- Chatbot evaluation: 3 reasoning-enabled, general-purpose chatbots represented prominent consumer chatbot ecosystems.Humanity’s Last Exam was additionally referenced as external context for model capability.
- Prompt design: 3 prompt variants represented patient, clinician, and evidence-synthesis researcher roles while keeping wording as consistent as possible.Prompts explicitly instructed chatbots to rely on primary evidence, including original research studies and clinical trials, rather than secondary evidence.
- Reference standard: Cochrane included and excluded studies served as the reference standard for analyzing studies retrieved in each chatbot response.The analysis focused on the set of uniquely identifiable primary studies rather than citation or publication counts.
3. Results
Across 720 responses, chatbots retrieved some Cochrane-included studies but also cited excluded studies, with recall varying by chatbot and user role. Larger study sample size was the only independent predictor of recall, while metadata errors affected 3.0% of matched citation rows.
- Overall retrieval: 39.2% ± 29.8% mean recall of Cochrane-included studies contrasted with 5.0% ± 9.4% mean citation of excluded studies per response.Across all 720 responses, the mean number of retrieved studies per response was 10.66.
- Model and user role: 63.1% ± 29.5% recall for ChatGPT exceeded Claude’s 37.0% ± 23.8% and Gemini’s 17.3% ± 13.1% (p = 2.0 × 10^-5).Recall also differed by user role: researcher 42.8% ± 30.8%, clinician 38.6% ± 28.9%, and patient 36.1% ± 29.3%.
- Citation overlap: 328 of 442 Cochrane-included studies (74.2%) appeared in at least one response, whereas 143 of 932 excluded studies (15.3%) were cited.Among cited included studies, 107 (32.6%) were cited by all three chatbots; only 16 of 143 cited excluded studies (11.2%) were cited by all three.
- Predictors of recall: Adjusted OR 1.80 per unit increase in log sample size predicted recall (95% CI, 1.37–2.36; p < 0.001), while publication year, citations per year, and open-access status were not significant.The regression included 330 studies with complete data for all predictors.
- Citation accuracy: 3.0% of 7,676 matched response-study citation rows carried study metadata errors, with rates of 5.9% for Claude, 1.5% for ChatGPT, and 0.2% for Gemini.Errors involved conflicting bibliographic details such as lead author, year, journal, or page range.
4. Discussion
Chatbot study retrieval was incomplete and selective, varying substantially by model and less by user role, with larger trial size the only independent predictor after adjustment. The findings support using chatbots as supplementary tools rather than replacements for reproducible evidence searches and formal review methods.
- 39.2% of Cochrane-included studies were retrieved per response on average, while pooling all responses identified 74.2% at least once.
- 63.1% recall for ChatGPT exceeded Claude’s 37.0% and Gemini’s 17.3%, whereas user-role recall ranged from 36.1% to 42.8%.
- 15.80 studies per response for ChatGPT was approximately four times Gemini’s 3.87, suggesting citation volume largely drove model-level recall differences.
- OR = 1.80 was the only independent retrieval association for larger trial size after adjusting for publication year, citations per year, and open-access status.Citation rate remained positively associated with retrieval, but its clustered confidence interval included the null (p = 0.057).
- Chatbot citations represent model-specific, role-conditioned subsets of literature that can vary across repeated queries and should not replace reproducible database searches or formal evidence synthesis.
- The exploratory test set comprised only 20 reviews from two recent Cochrane issues, limiting breadth and generalizability; future work should independently assess citations outside the included and excluded study sets.Such citations may be irrelevant studies, newly published eligible studies, or other related evidence, requiring independent adjudication and sometimes clinical expertise.