Source-linked AI summary

ChatGPT Makes Medicine Easy to Swallow: An Exploratory Case Study on Simplified Radiology Reports

Katharina Jeblick, Balthasar Schachtner, Jakob Dexl, Andreas Mittermeier, Anna Theresa Stüber, Johanna Topalis, Tobias Weber, Philipp Wesp, Bastian Sabel, Jens Ricke, Michael Ingrisch

arXiv:2212.14882v1cs.CLcs.LG

TL;DR

Patients may use ChatGPT to make difficult radiology reports easier to understand, although general-purpose model outputs are not guaranteed to be factually correct or complete. This exploratory case study asked 15 radiologists to evaluate simplified versions of three fictitious reports. Most radiologists judged the outputs correct, complete, and not potentially harmful, but they also identified incorrect statements, missing key findings, and potentially harmful passages.

  • Problem

    Patients need better access to understandable radiology reports, but automated report simplification has received limited attention and ChatGPT’s factual correctness and completeness are not guaranteed.

  • Method

    The study used a questionnaire in which 15 radiologists assessed ChatGPT-generated simplifications of three fictitious radiology reports.

  • Results

    Most radiologists agreed that the simplified reports were factually correct, complete, and not potentially harmful, although errors, omissions, and harmful passages were also identified.

  • Takeaways & Limitations

    The findings indicate potential for LLM-based medical text simplification to support patient-centered care, provided experts remain involved.

  • Takeaways & Limitations

    The study used only three fictitious, moderately complex reports and 15 experts, limiting the scope of its evidence.

Abstract

from arXiv · show

The release of ChatGPT, a language model capable of generating text that appears human-like and authentic, has gained significant attention beyond the research community. We expect that the convincing performance of ChatGPT incentivizes users to apply it to a variety of downstream tasks, including prompting the model to simplify their own medical reports. To investigate this phenomenon, we conducted an exploratory case study. In a questionnaire, we asked 15 radiologists to assess the quality of radiology reports simplified by ChatGPT. Most radiologists agreed that the simplified reports were factually correct, complete, and not potentially harmful to the patient. Nevertheless, instances of incorrect statements, missed key medical findings, and potentially harmful passages were reported. While further studies are needed, the initial insights of this study indicate a great potential in using large language models like ChatGPT to improve patient-centered care in radiology and other medical domains.

1 Introduction

ChatGPT’s convincing language generation creates opportunities for simplifying specialized medical reports, but its outputs may be inaccurate or harmful. This exploratory study therefore examines whether radiology reports can be simplified while remaining correct, complete, and safe for patients.

  • 1 Introduction: LLMs can simplify complex text and make it more accessible to broader audiences, including readers facing specialized medical language.Medical documents often have immediate consequences for non-experts.
  • 1 Introduction: Radiology reports commonly use specialized jargon for clinicians, making them difficult for readers without medical backgrounds to interpret.Delayed doctor-patient conversations may lead patients to seek explanations independently.
  • 1 Introduction: ChatGPT may produce plausible-sounding simplifications whose factual truth is not guaranteed, creating risks of incorrect or harmful patient interpretations.The study frames factual correctness, completeness, and potential harm as questions requiring expert assessment.
  • 1 Introduction: The exploratory case study investigates how patients might use or misuse ChatGPT to simplify radiology reports and the resulting implications for patient-centered care.The authors asked 15 radiologists to rate simplified reports on factual correctness, completeness, and potential harm.

2 Background

The background distinguishes text simplification from summarization and situates ChatGPT within the broader development of large language models. It emphasizes both the need for understandable radiology reports and the limitations of applying general-purpose LLMs in medicine.

  • 2.1 Large Language Models in Natural Language Processing: Large language models are transformer-based systems scaled to billions of parameters and trained on enormous datasets for varied downstream tasks.The background describes foundation models as broadly adaptable models trained on very large data bases.
  • 2.2 ChatGPT: ChatGPT is an OpenAI LLM released on November 30th, 2022, built on the GPT family and adapted for conversational instruction following.Its training setup is closely related to InstructGPT, with GPT-3.5 as the backbone and additional undisclosed safety mechanisms.
  • 2.3 Simplification and Summarization of Radiology Reports: Simplification transforms text to improve readability and understanding, whereas summarization creates a shorter version containing important aspects.Simplification therefore does not necessarily reduce length.
  • 2.3 Simplification and Summarization of Radiology Reports: Patients need better access to understandable radiology reports, but automated report simplification has received less attention than radiology report summarization.ChatGPT and other foundation models are presented as promising while requiring careful consideration of medical limitations.
  • 2.4 Known Limitations of LLMs: Known challenges include hallucinated statements, intrinsic biases, and the fact that ChatGPT was not developed specifically to simplify radiology reports.The authors discuss domain adaptation and reward modeling as possible ways to improve quality, while identifying these as ongoing challenges.
  • 2.4 Known Limitations of LLMs: ChatGPT-generated simplifications may sound plausible even though factual correctness and completeness are not guaranteed.This limitation motivates the authors’ subsequent assessment of simplified radiology reports.

3 Methods

The study uses an exploratory questionnaire to assess ChatGPT-generated simplifications of three fictitious radiology reports. Fifteen radiologists evaluate the reports using criteria covering correctness, completeness, and potential harm.

  • 3 Methods: The exploratory case study asks radiologists for their opinion on the quality of ChatGPT-generated simplified radiology reports.The questionnaire included three fictitious original reports and one unique ChatGPT-generated simplification of each.
  • 3.1 Original Radiology Reports: The study used three fictitious, moderately complex reports covering Knee MRI, Head MRI, and oncological CT cases.The reports mimicked clinical routine by including prior information, imaging findings, and conclusions.
  • 3.2 Simplified Radiology Reports: ChatGPT was prompted with “Explain this medical report to a child using simple language:” followed by each original report.The prompt was selected heuristically after comparing several formulations for simplicity and stability.
  • 3.2 Simplified Radiology Reports: The researchers generated 15 outputs for each of the three reports to account for ChatGPT’s nondeterministic output.The interface did not allow tuning the temperature parameter controlling token-probability handling.
  • 3.3 Questionnaire and Evaluation: Fifteen radiologists independently rated simplified reports using five-point Likert scales for factual correctness, completeness, and potential harm.Follow-up questions asked them to mark incorrect passages and list missing relevant medical information.
  • 3.4 Data Analysis: The questionnaire data were checked for consent and completeness, then summarized with descriptive statistics for each of the three clinical cases.Reported ordinal-scale statistics included medians, quartiles, interquartile ranges, minima, maxima, means, and standard deviations.

4 Results

Across 45 simplified reports, radiologists generally rated ChatGPT outputs as factually correct and complete, and disagreed that they would cause harm. Free-text review nevertheless identified incorrect statements, omitted findings, and potentially harmful wording across cases.

  • Likert Scale Analysis: 75% of ratings marked factual correctness and completeness as Agree or Fully agree, with median ratings of 2 for both criteria.Completeness received more positive ratings than factual correctness, with means of 1.8 and 2.2, respectively.
  • Likert Scale Analysis: Radiologists disagreed that simplified reports would lead patients to harmful wrong conclusions, with a median potential-harm rating of 4.Potential-harm ratings were more broadly distributed than correctness and completeness ratings, and Strongly Disagree was never selected.
  • Likert Scale Analysis: The three cases showed no relevant median differences in correctness, completeness, or potential harm, although Head MRI and Oncol. CT tended toward higher perceived harm risk.The cases were Knee MRI, Brain MRI, and Oncol. CT.
  • Free-text Analysis: 51% of participants highlighted incorrect passages, while 22% identified missing relevant information and 36% identified potentially harmful conclusions.These percentages came from follow-up free-text questions and text evidence.
  • Incorrect Text Passages: Incorrect passages included misinterpreted medical terms, imprecise or odd language, hallucinated findings, and occasional grammatical errors.Examples included treating differential diagnosis as a final diagnosis and introducing findings absent from the original report.
  • Missing Key Medical Information: Simplified reports often omitted key findings or used nonspecific locations, including missing cartilage damage, lesion growth, metastases, or disease sites.These omissions and location problems affected Knee MRI, Head MRI, and Oncol. CT reports.
  • Potentially Harmful: Potentially harmful conclusions arose from misinterpreted terms, hallucinations, imprecise wording, missed findings, and unspecific disease locations.Examples included calling a differential diagnosis cancer, denying brain damage despite a growing mass, and obscuring which lesion was progressing.
  • Word Count: The simplified Knee MRI report had a median word count of 414 (IQR 60), compared with 222 words in the original report.Simplified Head MRI and Oncol. CT reports had word counts comparable to their originals.

5 Discussion

ChatGPT produced generally accurate and complete simplified radiology reports, but individual hallucinations, omissions, and misleading statements created safety concerns. The discussion therefore balances potential gains in patient autonomy with the need for expert oversight, technical safeguards, and further validation.

  • 5.1 Findings: Most radiologists rated the ChatGPT-generated reports as factually correct and containing relevant medical information.Factual correctness had a median rating of 2, and completeness also received overall agreement.
  • 5.1 Findings: Hallucinations add statements not derivable from the original report and remain an intrinsic problem of generative language models.Other observed error categories included medical-term misinterpretation, imprecise or odd language, and grammatical errors.
  • 5.1 Findings: About half of radiologists identified critical statements with high potential to lead patients to wrong conclusions, despite disagreement that harmful conclusions would generally be drawn.Examples included overstating possible cancer and understating disease progression.
  • 5.2 Opportunities and Challenges: Simplified reports may improve comprehension, help patients prepare questions, and support greater autonomy and participation in treatment.These benefits are especially relevant when patients face delayed consultations or language barriers.
  • 5.2 Opportunities and Challenges: Reading simplified reports without expert assistance may produce misinterpretation, mental stress, or unsupervised clinical decisions.Patients might delay appointments, omit further care, or terminate therapy without professional consultation.
  • 5.2 Opportunities and Challenges: Key technical and operational challenges include domain adaptation, rare-pathology accuracy, outdated medical knowledge, nondeterministic outputs, and privacy concerns.The study also notes that ChatGPT was not developed specifically for radiology-report simplification.
  • 5.2 Opportunities and Challenges: The authors envision clinic-based systems that generate simplified reports alongside originals, with radiologist proofreading and privacy-preserving medical-domain adaptation.Both versions would be issued to the patient after necessary corrections.
  • 5.3 Limitations of the Study: The study’s evidence is constrained by three fictitious medium-complexity reports, 15 experts, qualitative language assessment, and an unvalidated prompt.The authors also note that the report writer was not a native English speaker.

6 Conclusion

The study found substantial potential in ChatGPT-generated simplified radiology reports, but also identified accuracy, completeness, and safety concerns requiring medical oversight and further technical work.

  • Most radiologists judged the simplified reports factually correct, complete, and not potentially harmful to patients.
  • Radiologists nevertheless identified incorrect statements, missing key medical information, and passages that could lead patients to harmful conclusions without physician involvement.
  • The exploratory findings suggest that simplified reports could strengthen patient autonomy during treatment.
  • Future clinical use should rely on domain-adapted models, expert approval, and delivery alongside the original report.

A.1.1 Simplified Report Sample #1

The first simplified report explains radiography, MRI, and knee findings in child-accessible language while preserving the report’s main abnormalities.

  • Radiography and MRI are explained as methods for producing internal body images, with MRI offering different imaging techniques.
  • The medial knee compartment is described as having minor cartilage damage, swelling, ligament damage, and a small semimembranosus muscle tear.
  • The lateral compartment includes substantial bone swelling, an outer meniscus tear, and possible injury to surrounding tissues.
  • The report describes tendon wear, a correctly positioned kneecap, and a thinned MPFL around the patellofemoral joint.
  • The anterior cruciate ligament is described as most likely completely torn, while the posterior cruciate ligament is intact.
  • Overall, the simplified conclusion summarizes complete ACL and medial ligament tears, meniscus and muscle injuries, bone swelling, tendon wear, and minor cartilage damage.

A.1.2 Simplified Report Sample #2

The second simplified report presents MRI concepts and knee findings in simple language, including ligament, meniscus, bone, tendon, and cartilage abnormalities.

  • Radiography is explained as internal imaging using machines that use X-rays, while MRI uses magnets and computers for varied views.
  • The medial knee compartment has cartilage damage, swelling, non-torn meniscus, partially damaged ligaments, and a small muscle tear.
  • The lateral compartment includes bone swelling, a small meniscus tear, possible surrounding tissue damage, and intact lateral structures.
  • The patellofemoral region shows tendon wear, a correctly positioned kneecap, and a slightly thinned supporting ligament.
  • The front knee ligament is described as most likely completely torn, while the back ligament is intact.
  • The overall summary reports a complete anterior knee ligament tear alongside medial and muscle tears and bone and cartilage damage.

A.1.3 Simplified Report Sample #3

The third simplified report explains imaging procedures and translates detailed knee findings into accessible descriptions while retaining uncertainty around some injuries.

  • The report explains radiography as internal imaging and MRI as multislice imaging enhanced by additional techniques.
  • The medial knee region includes minor cartilage damage, swelling and bruising, intact meniscus, and damaged inner ligaments.
  • The outer knee shows substantial bone swelling, an outer meniscus tear, an intact lateral collateral ligament, and possible anterolateral ligament damage.
  • The front knee contains mildly sore tendons, a properly positioned kneecap, and a thinned MPFL.
  • The anterior cruciate ligament is described as most likely completely torn, while the posterior cruciate ligament is not damaged.
  • The overall assessment combines complete ACL and higher-grade medial ligament tears with an outer meniscus tear, possible anterolateral injury, muscle damage, and minimal cartilage damage.

A.1.4 Simplified Report Sample #4

The report describes MRI findings across the knee, including a complete anterior cruciate ligament rupture and additional injuries involving ligaments, bone, meniscus, muscle, and cartilage.

  • MRI uses strong magnets and radio waves to produce detailed images, with additional techniques used to visualize tissues more clearly.
  • Medial compartment: The medial compartment shows minor cartilage damage, bone swelling, ligament damage, and a small semimembranosus muscle tear, while the meniscus remains intact.
  • Lateral compartment: The lateral compartment shows bone swelling, a meniscus tear, and possible tissue damage, while the lateral collateral ligament and popliteus muscle remain intact.
  • Patellofemoral joint: The patellofemoral joint and quadriceps tendon show minor wear, the patella is correctly positioned, and the medial patellofemoral ligament is thinned.
  • Intercondylar area: The anterior cruciate ligament is most likely completely ruptured, whereas the posterior cruciate ligament is intact.
  • Conclusion: The report conclusion summarizes a complete anterior cruciate ligament rupture, higher-grade medial ligament rupture, lateral tibial plateau fracture, outer meniscus tear, semimembranosus tear, and minimal cartilage damage.

A.1.5 Simplified Report Sample #5

The simplified report explains MRI examination methods and summarizes injuries across the knee, emphasizing a complete ACL tear alongside ligament, meniscus, muscle, bone, and cartilage findings.

  • An MRI was used to obtain detailed images of the legs, with different techniques applied to improve image quality.
  • Medial compartment: The medial compartment contains minor cartilage damage, slight swelling, partially torn ligaments, and a small semimembranosus muscle tear; the meniscus is intact.
  • Lateral compartment: The lateral compartment contains substantial bone swelling, an outer meniscus tear, and possible nearby tissue damage, while specified lateral structures are intact.
  • Patellofemoral joint: The patellofemoral joint shows minor wear involving tendons, the kneecap is correctly positioned, and the MPFL is thinned.
  • Intercondylar area: The anterior cruciate ligament is most likely completely torn, while the posterior cruciate ligament is intact.
  • Overall findings: Overall, the principal problems are complete ACL and higher-grade medial ligament tears, an outer meniscus tear, a muscle tear, bone swelling, and minimal cartilage damage.

A.1.6 Simplified Report Sample #6

The report explains MRI imaging and describes multiple knee injuries, including a complete ACL tear, partial medial ligament tears, meniscus damage, and minor tendon and cartilage abnormalities.

  • MRI uses strong magnets and radio waves to create detailed images, with multiple techniques and viewing angles used to examine injuries.
  • Medial knee: The medial knee shows cartilage damage and swelling, partial tears of inner ligaments, and a small semimembranosus muscle tear, while the meniscus is intact.
  • Outer knee: The outer knee shows bone swelling and meniscus damage, with possible injury to other tissues; the outer ligaments and popliteus muscle are not torn or damaged.
  • Anterior knee: The quadriceps and patellar tendons are slightly damaged, the patella is correctly positioned, and the patella-stabilizing ligament is slightly damaged.
  • Central knee: The anterior cruciate ligament is completely torn, while the posterior cruciate ligament is not damaged.
  • Summary: In summary, the patient has a complete ACL tear, partial inner-ligament tears, inner cartilage damage, outer bone swelling, outer meniscus tearing, and slight anterior tendon and patellar damage.

A.1.7 Simplified Report Sample #7

The simplified explanation introduces MRI and knee anatomy before describing findings across the compartments and joints, with a complete ACL tear as the main injury.

  • MRI uses magnets and radio waves rather than X-rays and may use extra techniques or angles for clearer, more detailed images.
  • Knee anatomy: The medial compartment is the inside part of the knee, while the lateral compartment is the outside part.
  • Medial compartment: The medial compartment has minor thigh-bone cartilage damage, swelling, inner-knee cartilage injury, ligament damage, and a small semimembranosus tear; the meniscus is intact.
  • Lateral compartment: The lateral compartment has substantial lower-leg bone swelling, an outer meniscus tear, and possible tissue damage, while the outer ligament and popliteus muscle are intact.
  • Patellofemoral joint: The patellofemoral joint has minor tendon damage and a thinned kneecap-stabilizing ligament, while the kneecap remains correctly positioned.
  • Intercondylar area: The ACL is completely torn in the intercondylar area, while the PCL is not damaged.
  • Overall findings: Overall, the report identifies a complete ACL tear plus other ligament, bone, cartilage, muscle, and tendon injuries around the knee.

A.1.8 Simplified Report Sample #8

This simplified knee report explains findings across the medial, lateral, patellofemoral, and intercondylar regions. It identifies a complete anterior cruciate ligament tear alongside additional tissue damage and swelling.

  • The medial compartment contains cartilage damage, bone swelling and bruising, ligament damage, and a small muscle tear, while the meniscus remains intact.
  • The lateral compartment shows substantial bone swelling, meniscus damage, and possible injury to other tissues, although several named structures remain intact.
  • The patellofemoral joint has minor tendon irritation, a correctly positioned kneecap, and a thinned medial patellofemoral ligament.
  • The anterior cruciate ligament is most likely completely torn, whereas the posterior cruciate ligament is intact.
  • Overall, the report describes a complete anterior cruciate ligament tear with additional ligament, muscle, bone, and cartilage abnormalities.

A.1.9 Simplified Report Sample #9

This knee report translates imaging findings into accessible descriptions of anatomy and injury. Its overall assessment includes a probable complete ACL tear, meniscus and ligament damage, swelling, and minor cartilage or tendon abnormalities.

  • The report introduces radiography and MRI as imaging tests, noting that MRI uses magnets, radio waves, computers, and varied techniques to produce detailed pictures.
  • It identifies the medial, lateral, patellofemoral, and intercondylar regions as the principal knee areas examined.
  • The medial compartment has cartilage and ligament damage, swelling, an intact meniscus, and a semimembranosus muscle tear.
  • The lateral compartment contains bone swelling and a small meniscus tear, with possible nearby tissue damage but an intact popliteus muscle.
  • The patellofemoral joint has tendon wear, while the patella is correctly positioned and the medial patellofemoral ligament is thinned.
  • The anterior cruciate ligament is most likely completely torn, while the posterior cruciate ligament is not damaged.
  • The report’s overall conclusion combines a complete ACL tear with other ligament and muscle injuries, bone swelling, a meniscus tear, and minimal medial cartilage damage.
  • A further simplified version describes the knee as injured and needing treatment, while retaining the complete ACL tear and associated abnormalities.

B Questionnaire

The questionnaire documents consent, presents meta and report-specific materials, and records radiologists’ answers and comments. The appendix includes responses concerning factual relevance, omissions, wording, and report content.

  • B.1 Questionnaire Design: The questionnaire is titled “Quality of Simplified Radiological Reports” and includes an anonymized participation-consent statement.
  • B.1 Questionnaire Design: Figure 5 presents the questionnaire’s meta questions and cover.
  • B.1 Questionnaire Design: Figure 6 presents questions concerning one simplified report supplied to a radiologist.
  • B.2 Answers to the Questionnaire: The recorded answers include comments about abnormal findings, missing information, wording, communication, and report conclusions.
  • B.2 Answers to the Questionnaire: Additional responses mention brain masses, glioblastoma, cancer spread, lesion growth, and thyroid-related concerns.
  • B.2 Answers to the Questionnaire: The appendix labels answer records by participant experience and question fields, including Q1, Q1a, Q2, Q2a, Q3, and Q3a.
  • B.2 Answers to the Questionnaire: Other responses address factual relevance, report specificity, omitted communication, and changes in lesion size or spread.
  • B.2 Answers to the Questionnaire: Table 5 is identified as containing the radiologists’ questionnaire answers.
Loading 2212.14882v1…