Source-linked AI summary
Towards AI-Assisted Clinical Trial Matching: Practical Considerations, Multicenter Evaluation, and Real-World Deployment
Yin Fang, Qiao Jin, Shubo Tian, Lauren He, Maya Geer, Noor Naffakh, Ryan Huu-Tuan Nguyen, Zifeng Wang, Jimeng Sun, Charalampos S. Floudas, James L. Gulley, Kamilia Moalem, Catarina Martins Maia, Amanda Nottke, Juan W. Valle, Melinda Bachini, Lourdes Rocha-Nussbaum, Kari Ramage, Nikita Curry, Megan Barnes, Mandy Mansaray, Darlene Gabeau, Craig E. Grossman, Heath Skinner, Michael Burczynski, NIH-TrialBench Consortium, Zhiyong Lu
TL;DR
Clinical trial matching remains difficult because recruitment depends on complex eligibility interpretation, while prior AI systems have focused mainly on eligibility and had limited real-world oncology evaluation. This study develops and deploys TrialGPT 2.0, a configurable recommendation system with inspectable explanations, and evaluates it across retrospective, prospective, benchmark, and clinical workflows. The evaluations support improved clinician efficiency and identification of additional trial opportunities, while leaving downstream enrollment impact for future study.
Problem
Patient-trial matching is complex and largely manual, while prior AI systems have emphasized eligibility assessment and provided limited evidence from clinician-adjudicated or prospective oncology workflows.
Method
TrialGPT 2.0 ranks candidate trials by patient-trial fit using configurable workflow priorities and provides structured evidence, incompatibilities, missing information, and rationales for review.
Results
Across retrospective and prospective clinical evaluations, TrialGPT 2.0 retrieved clinician-recommended trials, reduced screening time, and identified additional trial opportunities in an active precision-oncology workflow.
Takeaways & Limitations
AI-assisted matching can help oncology teams identify trial options that clinicians choose to act on and improve the efficiency of clinical trial review.
Takeaways & Limitations
The study did not assess downstream outcomes such as patient contact, referral, formal screening, consent, enrollment, accrual, or patient outcomes.
Abstract
from arXiv · showhide
Clinical trials are essential for advancing cancer care and drug development, but many fail because of insufficient patient enrollment. While there is growing interest in using AI to support patient recruitment, existing systems largely perform eligibility assessment alone and have rarely been evaluated in real-world oncology workflows. Here we present TrialGPT 2.0, an AI-assisted clinical trial recommendation system designed for real-world deployment. Rather than asking only whether a patient may qualify, the system also assesses which trials warrant further consideration given the patient's current clinical needs and local workflow priorities, and provides structured, inspectable explanations for expert review. Importantly, we evaluated TrialGPT 2.0 retrospectively and prospectively across multiple oncology-focused settings, spanning government, academic cancer-center, patient-advocacy, and NIH referral workflows. In retrospective multicenter cohorts comprising 288 cases, TrialGPT 2.0 retrieved at least one clinician-recommended trial in its top 10 recommendations for approximately 91% of cases while reducing clinician screening time by 55.0%. In a six-month prospective evaluation embedded in an active precision oncology tumor board, TrialGPT 2.0 contributed additional trial opportunities missed by the routine workflow, expanding patient access to clinical trial participation by 90.9%. To support scientific reproducibility, we also introduce NIH-TrialBench, a clinician-authored dataset comprising 126 diverse synthetic patient vignettes and matching scenarios from 11 NIH Institutes and Centers. Together, these results support the value of AI to assist clinical trial matching by improving clinician efficiency and identifying frequently overlooked trial opportunities, ultimately helping to expand and accelerate accrual to cancer trials.
Introduction
Patient-trial matching is a manual, complex bottleneck in oncology recruitment, while prior AI systems have emphasized eligibility and lacked real-world clinical evaluation. TrialGPT 2.0 addresses these gaps with configurable recommendation, structured explanations, multicenter evaluation, reproducible benchmarking, and workflow deployment.
- Clinical need: Patient-trial matching remains slow and error-prone because it requires integrating complex eligibility criteria, heterogeneous patient information, and changing site-level trial portfolios.Poor recruitment contributes substantially to premature trial discontinuation, particularly in oncology.
- Prior AI approaches: Prior AI approaches support trial search, cohort identification, and eligibility screening, including TrialGPT’s zero-shot retrieval, criterion assessment, and trial-ranking framework.These methods have expanded automation beyond manual matching but primarily address eligibility-related tasks.
- Unmet need: Existing systems distinguish eligibility poorly from recommendation and have limited clinician-adjudicated or prospective validation using real oncology workflows.Eligibility indicates that enrollment may be possible, whereas recommendation considers clinical appropriateness and current treatment needs.
- TrialGPT 2.0: TrialGPT 2.0 generates ranked trials using configurable portfolios and workflow-specific priorities, with structured evidence, incompatibilities, missing information, and rationales for expert review.The system ranks overall patient-trial fit rather than reporting eligibility alone.
- Evaluation and reproducibility: The study evaluates TrialGPT 2.0 retrospectively and prospectively across heterogeneous oncology and NIH referral workflows, alongside NIH-TrialBench and established public benchmarks.NIH-TrialBench contains 126 synthetic vignettes contributed by investigators from 11 NIH Institutes and Centers.
- Deployment: A web-based interface deploys the system for routine trial review, supporting more efficient identification of clinically appropriate options in real-world workflows.The authors position this deployment as a step toward broadening and accelerating cancer trial recruitment.
Results
TrialGPT 2.0 was evaluated across retrospective, prospective, and reproducible benchmarking settings, showing alignment with clinician recommendations, faster screening, and additional trial opportunities. Its results also illustrate that recommendation requires prioritization beyond eligibility alone.
- Retrospective evaluation: 0.91 hit rate at K=10 across four retrospective cohorts placed at least one clinician-recommended trial in most cases.Eligible precision decreased only from 0.89 at K=1 to 0.87 at K=10.
- Retrospective evaluation: 94.0% clinician agreement with consensus labels occurred with TrialGPT 2.0 assistance, compared with 94.2% without assistance.TrialGPT 2.0 alone reached 89.0% exact agreement; assisted review had 5 severe reversals versus 16 without assistance.
- Retrospective evaluation: 55.0% reduction in mean screening time lowered per-patient decision time from 129.8 to 58.4 seconds in the counterbalanced experiment.The timing excluded patient-context review and automated PDF report generation.
- Prospective evaluation: 90.9% increase in cases with at least one final recommendation occurred when TrialGPT 2.0 expanded recommendations from 11 to 21 across 27 prospective tumor-board cases.TrialGPT 2.0 expanded final trial options in 10 cases, or 37.0% of all reviewed cases.
- Prospective evaluation: 83.3% of 54 final clinician-selected recommendations were identified by TrialGPT 2.0, including 37 recommendations identified by TrialGPT 2.0 alone.The routine workflow identified 17 recommendations in total.
- Reproducible benchmarking: 70% macro-averaged target-trial Recall@10 on NIH-TrialBench exceeded 54% for TrialGPT 1.0 and 43% for ChatGPT-5.5 Thinking.NIH-TrialBench contained 126 vignettes spanning 11 NIH Institutes and Centers.
Discussion
TrialGPT 2.0 addresses limitations of eligibility-only matching through real-world, workflow-aware recommendation evaluation, while acknowledging boundaries around downstream outcomes, generalizability, and residual errors.
- TrialGPT 2.0 was evaluated through retrospective review, prospective precision-oncology tumor-board assessment, reproducible benchmarking, and deployment.
- Its real-world evaluation used clinician adjudication and prospective review rather than relying only on curated technical benchmarks.
- TrialGPT 2.0 frames matching as recommendation beyond eligibility, using graded categories and fit scores to incorporate clinical appropriateness and workflow priorities.
- The study measured early-stage outcomes, but not patient contact, referral, formal screening, consent, enrollment, accrual, or patient outcomes.
- The prospective evaluation covered a single POTB workflow over a limited period with a modest number of cases, leaving generalizability and downstream impact for larger studies.
- Residual errors included overly literal eligibility application and mismatches in complex multi-arm trials or cases with incomplete or clinically ambiguous information.
- AI-assisted matching identified additional trial options that clinicians chose to act on, but broader value requires rigorous workflow integration and outcome-based evaluation.
Methods
TrialGPT 2.0 was evaluated as a configurable, versioned clinical trial-matching system across retrospective, prospective, and synthetic benchmark settings.
- System design: TrialGPT 2.0 takes patient context, a predefined local trial corpus, and a workflow-specific matching policy as inputs.Patient context may be raw clinical records or a clinician-prepared narrative summary.
- System design: The system uses pluggable retrieval to support matching across local trial corpora ranging from 9 to 1,871 candidate trials.For large trial pools, the default retriever uses hybrid-fusion retrieval before detailed assessment.
- System design: Each candidate trial receives a structured assessment combining eligibility, clinical relevance, uncertainty, fit score, confidence, and recommendation category.Trials are ranked primarily by fit score, with confidence breaking ties.
- Evaluation design: Evaluations fixed code, prompts, model, decoding, retrieval, candidate sets, and matching policies before inference to support versioned analysis.Prospective evaluation used the current local trial-corpus snapshot available at each case review.
- Evaluation design: The study combined 288 de-identified retrospective cases, prospective Precision Oncology Tumor Board cases, and 126 clinician-authored synthetic vignettes.NIH-TrialBench vignettes were contributed by investigators from 11 NIH Institutes and Centers.
- Evaluation design: Review strategies varied by cohort, using top-ranked review, exhaustive review for the nine-trial NCI portfolio, or prospective workflow review.Clinician annotations underwent completeness and consistency checks before metric calculation.
Supplementary Information
The supplementary materials describe TrialGPT 2.0’s configurable matching policies, factor coverage, evaluation analyses, interface, runtime, reviewer acceptance, and NIH-TrialBench recovery.
- Supplementary analyses cover score responses to mismatch and missing information, clinician comments, public benchmark comparisons, runtime, and reviewer acceptance.
- The supplementary results include reviewer acceptance responses and NIH-TrialBench target-trial recovery by NIH Institute or Center and overall.NIH-TrialBench recovery categories are Highly Recommended, Possible Match, and Low Fit.
- TrialGPT 2.0 matching policies classify trial-fit factors as dedicated, partial, or triage-level components.All policies receive patient context and raw eligibility criteria, while factor-specific instructions vary by workflow.
- Policies differ in their handling of diagnosis, disease status, biomarkers, line of therapy, prior treatment, cohort or arm applicability, and inclusion or exclusion criteria.Some workflows use dedicated instructions, whereas others consider factors through eligibility criteria or high-level referral triage.
- Supplementary Table 1 reports factor-assessment depth by matching policy using symbols for specific, partial, or triage-level instructions.
Supplementary Note 2: Fit and confidence scores respond to clinical mismatch and missing information
The analyses test whether TrialGPT 2.0’s fit and confidence scores respond appropriately to clinical compatibility and information completeness. They also report benchmark, efficiency, runtime, and reviewer-acceptance results, alongside a scope limitation for therapeutic relevance.
- Clinician-recommended pairs received higher fit scores than partial-match and patient-shuffled controls, indicating patient-specific rather than disease- or biomarker-only matching.Controls included disease-matched/biomarker-mismatched, biomarker-matched/disease-mismatched, and patient-shuffled pairs.
- Single-factor hard mismatches markedly reduced fit scores across evaluated clinical domains while holding the trial and remaining patient profile constant.Perturbations targeted factors such as diagnosis, stage, biomarkers, treatment history, cohort applicability, and eligibility criteria.
- Removing trial-critical information caused larger confidence reductions than removing trial-irrelevant information, supporting confidence as an indicator of assessment-relevant information availability.
- Across SIGIR, TREC 2021, and TREC 2022, TrialGPT 2.0 increased Precision@10 by 4.5%, nDCG@10 by 5.7%, MRR by 9.8%, and MAP by 10.2% versus TrialGPT 1.0.The evaluated metrics capture top-ranked precision, rank-sensitive relevance, early retrieval, and ranking quality across relevant trials.
- Averaged across benchmarks, processing speed increased by 3.20-fold while input and output tokens per patient-trial pair decreased by 58% and 73%, respectively.
- Manual clinician review, downstream chart review, and coordinator follow-up were not included in the runtime analysis.