Source-linked AI summary

Evaluating AI Generated Summaries for Cancer Patients

Muhammad Aurangzeb Ahmad, Kim Shyu, Leon Oliver, Fergus Sleight, Paul Landau

arXiv:2608.26154v1cs.CLcs.AI

TL;DR

AI-generated patient summaries can support cancer-care workflows but require careful evaluation because inaccuracies and omissions may affect downstream clinical communication and safety. This study combines expert review with LLM-as-a-judge assessment in a production cancer-care setting. Results show strong factual accuracy, professional style, and low detected toxicity, alongside greater variability in completeness and usefulness and failures involving subtle distortions or missing information.

  • Problem

    AI-generated clinical summaries may contain unsupported statements, numerical errors, omissions, or altered priorities that can persist across workflows and create potential patient-harm risks.

  • Method

    The study evaluates cancer-app summaries through human domain-expert review and LLM-as-a-judge assessment in a clinical production environment.

  • Results

    Across 103 summaries, FactScore averaged 0.992, ProfessionalStyle averaged 0.979, Toxicity was 0.0, while Completeness showed the greatest variability.

  • Takeaways & Limitations

    Evaluation insights support iterative improvements to prompt design, grounding, deployment practices, and safety guardrails.

Abstract

from arXiv · show

Large language models (LLMs) are increasingly being integrated into digital health platforms to generate summaries of complex medical data. Although these models can improve patient engagement and communication, these systems also raise concerns about accuracy, faithfulness, and safety in clinical contexts. In this study, we evaluate AI-generated summaries within a cancer patient care application using a dual assessment framework. Human domain experts, including oncology clinicians and patient-facing care staff, provided ground-truth evaluations of summary quality along dimensions of accuracy, clinical relevance, and readability. In parallel, we employed LLMs serving as evaluators (LLM-as-a-judge). Some limitations were identified in the generated summaries e.g., occasional omissions and minor inaccuracies. These were systematically analyzed and used to iteratively improve prompt design, grounding, and safety guardrails.

I. INTRODUCTION

The paper frames AI-generated patient summaries as safety-critical clinical intermediaries that require evaluation for accuracy, relevance, consistency, and uncertainty. It uses human experts and LLM-as-a-judge evaluation to guide responsible deployment and improvement.

  • Clinical records are fragmented across documentation types, increasing cognitive burden and information overload.
  • Generative summaries can introduce unsupported statements, numerical errors, omissions, or altered clinical priorities.
  • Small inaccuracies may persist across encounters or subsequent documentation, creating potential patient-harm risks.
  • The study evaluates summaries across coherency, fluency, consistency, relevance, and clinical use using human and LLM-based assessment.
  • Evaluation findings are used to improve deployment and usage of AI-generated summaries in clinical contexts.

II. RELATED WORK

Prior clinical summarization research progressed from rigid template systems to more fluent neural and LLM-based approaches, while retaining concerns about faithfulness, omissions, hallucinations, and patient-facing safety. Evaluation therefore increasingly emphasizes clinically meaningful quality, fairness, and monitoring.

  • Template-based clinical summarizers emphasized factual accuracy and traceability but were limited in flexibility and personalization.
  • Neural summarization improved fluency and coherence while evaluations continued to identify omissions and hallucinated content.
  • Patient-accessible summaries require attention to readability and relevance because these affect comprehension, adherence, and satisfaction.
  • Human-in-the-loop designs, retrieval augmentation, and bounded knowledge sources are common mitigations for hallucinations, inaccuracies, and omissions.
  • Healthcare evaluation increasingly considers fairness because performance can differ across demographic groups, clinical conditions, and language varieties.

III. CANCER CARE MONITORING APP

The study is situated in a cancer-care coordination and remote-monitoring app that collects longitudinal patient-reported and physiologic data. Summarization turns this high-volume record into an overview for clinicians.

  • The cancer-care app supports patients, caregivers, and oncology teams throughout treatment through coordination and remote monitoring.
  • Patients record symptoms, medications, vital signs, mood, well-being, and free-text journals in real time.
  • Longitudinal data outside hospital touchpoints captures treatment experiences and clinically meaningful events between appointments.
  • The app’s data volume makes summarization useful as a clinician overview.

IV. EVALUATION METHODOLOGY

The evaluation combines expert review of patient data and AI-generated summaries with automated LLM-as-a-judge assessment. Human experts rate summary quality on clinical and interpretability dimensions, while open-ended feedback captures risks and improvement needs.

  • Patient-reported and physiologic data are input to Claude Sonnet 3.7, which generates summaries from a predefined prompt.
  • Clinical experts from Guy’s and St Thomas’ Hospital validate summaries through manual review and structured ratings.
  • Experts rate accuracy, completeness, clarity, clinical relevance, faithfulness, and bias-free language on a 1–5 scale.
  • Open-ended questions assess omissions, unsupported assumptions, factual errors, ambiguity, unsafe statements, trust, and improvement needs.
  • Human and LLM evaluations use partially distinct rubrics, preventing direct quantitative comparison of identical metrics.

1) Evaluation Scores Clusters:

Human evaluation responses formed a dominant cluster with uniformly high performance, alongside smaller clusters reflecting slightly lower or moderate, less consistent outputs.

  • Most evaluations formed a dominant cluster with uniformly high performance across all metrics, corresponding to robust and clinically reliable summaries.
  • A second cluster showed slightly lower values, but overall evaluation remained high.
  • A third cluster showed moderate performance across metrics, reflecting partially correct but less consistent outputs.
  • Clarity, faithfulness, and accuracy varied more than other evaluation dimensions, while bias-free ratings remained consistently high.

2) Analysis of Likert Scale Responses:

Clinical users rated the summaries positively across criteria, while thematic analysis identified omissions, dosing errors, ambiguity, and outdated data as distinct quality concerns.

  • More than 75–85% of responses were positive for Accuracy, Completeness, Clinical Relevance, Faithfulness, and Bias-Free language, with no negative ratings.
  • Thematic Clustering: “No Data” cases achieved perfect scores, whereas “Missing Info” caused only minor degradation across evaluation dimensions.
  • Thematic Clustering: “Dosing Errors” sharply reduced overall quality despite high bias-free scores, indicating a clinically unsafe failure mode rather than biased output.
  • Thematic Clustering: “Ambiguity” produced moderate or middling scores, reflecting difficulty resolving uncertainty while leaving some content usable.
  • Thematic Clustering: “Outdated Data” primarily affected accuracy and clinical relevance, with low ratings linked to unclear data time ranges.

B. Automated Evaluation

Automated evaluation found high factual accuracy, professional style, and safety, but greater variability in completeness and helpfulness. Correlation and error analyses indicate that usefulness, verbosity, and faithfulness capture distinct trade-offs and failure modes.

  • Overall scores: FactScore averaged 0.992 and ProfessionalStyle 0.979 across 103 summaries, while Toxicity was 0.0 for all records.Faithfulness was high overall (mean=0.944; median=1.0) but had a minimum of 0.0.
  • Overall scores: Completeness had the greatest variability (SD=0.236; min=0.0) despite a median of 1.0, mainly because patient names were omitted after the first sentence.Readability and Helpfulness means were 0.789 and 0.816, respectively.
  • Metric relationships: Helpfulness correlated most strongly with Completeness (r = 0.61), while Readability had weak correlations with other metrics (r ≤0.22).Helpfulness also correlated with Professional Style (r = 0.47), FactScore (r = 0.42), and Faithfulness (r = 0.34).
  • Metric relationships: Completeness and Faithfulness had a modest correlation (r = 0.26), suggesting a potential trade-off between content coverage and source adherence.The analysis supports evaluating multiple quality dimensions rather than optimizing a single metric.
  • Error analysis: Manual review found temporal errors, unsupported numerical or clinical claims, and over-interpretation among the 23 summaries with Faithfulness below 1.0.Examples included incorrect dates, inferred medication non-adherence, and unsupported claims that symptoms improved or resolved.
  • Length analysis: Longer outputs correlated with Completeness (r ≈0.39–0.44) and Helpfulness (r ≈0.41-0.49), but Faithfulness showed weak negative correlations with length (r ≈-0.08 to -0.13).This pattern indicates that verbosity improved perceived coverage and utility without improving factual correctness.

VI. SUMMARIZATION IMPROVEMENT

Evaluation findings and review of underperforming summaries informed iterative updates to the generation prompt and system presentation. The updated platform combines longitudinal clinical information with grounded AI summaries to support rapid situational awareness and traceability.

  • Prompt refinement: Human expert feedback and automated evaluation results were used to iteratively refine the summary-generation prompt.The stated goal was to improve accuracy, completeness, and clinical relevance.
  • Platform updates: The updated platform integrates longitudinal clinical data, symptoms, medication adherence, and patient-reported inputs into a unified dashboard with an AI-generated summary.The interface presents both structured and narrative information for healthcare personnel.
  • Platform updates: Narrative summaries are grounded in source data and paired with structured metrics to help clinicians assess patient status, track treatment progression, and identify potential risks.The system combines structured data, patient activity logs, and explainable AI outputs while maintaining transparency and traceability.

VII. CURRENT & FUTURE WORK

The evaluation found high factual accuracy, appropriate clinical tone, and minimal safety risk, alongside greater variability in completeness and perceived usefulness. The findings support multi-metric evaluation and conservative, grounding-focused optimization for safety-critical clinical deployment.

  • Evaluation findings: The LLM produced summaries with high factual accuracy, appropriate clinical tone, and minimal safety risk, but completeness and perceived usefulness were more variable.This conclusion summarizes the reported evaluation results.
  • Evaluation findings: Helpfulness aligned with completeness, professional style, and factual correctness, whereas readability remained largely independent of substantive clinical quality.The result supports treating readability separately from other quality dimensions.
  • Trade-offs and errors: Longer inputs and outputs were associated with improved completeness and helpfulness but not with gains in faithfulness, indicating a verbosity–grounding trade-off.The paper reports that most failures involved subtle distortions or over-interpretations rather than overt hallucinations.
  • Implications: The findings support multi-metric evaluation and conservative, grounding-focused optimization for deploying summaries in safety-critical clinical settings.The system updates were internally tested and scheduled for further evaluation.
Loading 2608.26154v1…