Source-linked AI summary
Translating Radiology Reports into Plain Language using ChatGPT and GPT-4 with Prompt Learning: Promising Results, Limitations, and Potential
Qing Lyu, Josh Tan, Michael E. Zapadka, Janardhana Ponnatapura, Chuang Niu, Kyle J. Myers, Ge Wang, Christopher T. Whitlow
TL;DR
Radiology reports can be difficult for patients to understand, motivating evaluation of ChatGPT for plain-language translation and report-based suggestions. The study tests these capabilities on chest CT and brain MRI screening reports and compares ChatGPT with GPT-4. ChatGPT receives high radiologist-rated translation scores, while GPT-4 significantly improves translated-report quality; incompleteness and response variability remain limitations.
Problem
Radiology reports contain terminology that patients without medical backgrounds may find difficult to understand, motivating plain-language translation for patient education.
Method
The study evaluates ChatGPT translations and patient- and provider-directed suggestions on chest CT and brain MRI screening reports, then compares translation results with GPT-4.
Results
4.268 was ChatGPT’s overall translation score on a five-point scale, with 0.097 places of information missing and 0.065 places incorrect per translation; GPT-4 significantly improved results.
Takeaways & Limitations
The results support the feasibility of using large language models for radiology-report translation and clinical education.
Takeaways & Limitations
ChatGPT can produce incomplete, inconsistent, over-simplified, or information-losing translations for the same report and prompt.
Abstract
from arXiv · showhide
The large language model called ChatGPT has drawn extensively attention because of its human-like expression and reasoning abilities. In this study, we investigate the feasibility of using ChatGPT in experiments on using ChatGPT to translate radiology reports into plain language for patients and healthcare providers so that they are educated for improved healthcare. Radiology reports from 62 low-dose chest CT lung cancer screening scans and 76 brain MRI metastases screening scans were collected in the first half of February for this study. According to the evaluation by radiologists, ChatGPT can successfully translate radiology reports into plain language with an average score of 4.27 in the five-point system with 0.08 places of information missing and 0.07 places of misinformation. In terms of the suggestions provided by ChatGPT, they are general relevant such as keeping following-up with doctors and closely monitoring any symptoms, and for about 37% of 138 cases in total ChatGPT offers specific suggestions based on findings in the report. ChatGPT also presents some randomness in its responses with occasionally over-simplified or neglected information, which can be mitigated using a more detailed prompt. Furthermore, ChatGPT results are compared with a newly released large model GPT-4, showing that GPT-4 can significantly improve the quality of translated reports. Our results show that it is feasible to utilize large language models in clinical education, and further efforts are needed to address limitations and maximize their potential.
1 Introduction
The study examines whether ChatGPT can translate radiology reports into patient-friendly language and provide report-based suggestions. It also compares ChatGPT with GPT-4 for this clinical education task.
- Background: ChatGPT attracted attention because of its human-like expression and reasoning abilities.The passage describes ChatGPT as a large natural language processing model released by OpenAI.
- Related work: Earlier work had explored ChatGPT for downstream tasks including medical education, patient discharge summaries, language translation, and mathematical problem solving.The cited passage situates this study among emerging investigations of ChatGPT in clinical usage.
- Motivation: Radiology reports contain medical terminology that can be difficult for patients without medical backgrounds to understand.Plain-language re-expression is presented as potentially helpful for patient understanding, anxiety, compliance, and outcomes.
- Study aim: The study evaluates ChatGPT’s translation of radiology reports into lay language and its suggestions for patients and healthcare providers.The study also assesses suggestion quality and compares ChatGPT with GPT-4.
2 Methodology
The study tested ChatGPT on de-identified chest CT and brain MRI screening reports using prompts for translation and patient or provider suggestions. Radiologists evaluated the resulting translations for completeness, correctness, and overall quality.
- Data: 62 chest CT and 76 brain MRI screening reports were collected from a clinical database and de-identified before evaluation.The chest CT and brain MRI datasets followed distinct screening protocols and included reports classified by clinical categories.
- Prompts: ChatGPT received prompts to translate reports into plain language and provide suggestions for patients and healthcare providers.Responses were collected in mid-February.
- Evaluation: Two experienced radiologists evaluated the quality of ChatGPT’s responses.The evaluators had 21 and 8 years of experience, respectively.
- Evaluation: Translation quality was assessed using information loss, information misinterpretation, and a five-point overall score.The score ranged from 1 for worst quality to 5 for best quality, while missing and incorrect information were recorded per report.
- Evaluation: Suggestion evaluation recorded frequent suggestions, finding-specific suggestions, and inappropriate suggestions unrelated to report findings.
3 Key results
ChatGPT generally shortened and simplified chest CT and brain MRI reports while integrating information across report sections. Radiologist evaluations found few missing or incorrect details overall, although the supplied passages provide fuller quantitative evaluation for chest CT than brain MRI.
- Report comparison: 85.5% of chest CT translations were shorter than the originals, with an overall length reduction of 26.7%.For brain MRI reports, 72.4% were shorter, with an overall reduction of 21.1%.
- Report comparison: ChatGPT shortened multiple negative findings by summarizing them together in a single sentence.A chest CT example combines normal pleural, heart, blood-vessel, and lymph-node findings.
- Plain-language translation: ChatGPT replaced medical terminology with common words and explained a granuloma’s meaning and severity for patients.The example translates a 1mm granuloma into a small area of inflammation that is usually not concerning.
- Information integration: ChatGPT integrated findings with comparison information, stating that a 6mm granuloma had not changed since an August 2021 CT scan.
3.2 Evaluation of ChatGPT translations by radiologists
Radiologists evaluated ChatGPT translations using information loss, information misinterpretation, and a five-point overall quality score. Across chest CT and brain MRI reports, translations generally received favorable evaluations.
- 4.268 was the average overall score across all translated reports on the five-point evaluation scale.A score of 5 represented the best quality and 1 the worst.
- 0.080 missing-information places and 0.065 incorrect-information places occurred on average across all translated reports.These corresponded to roughly one occurrence every 12.5 and 15.4 reports, respectively.
- 76% of chest CT translations received an overall score of 5.
- Brain MRI translations received overall scores of 4 and 5 in 37% and 32% of cases, respectively.
3.3 Evaluation of ChatGPT-generated suggestions
ChatGPT generated highly relevant general suggestions for patients and healthcare providers, while providing report-specific suggestions in a subset of cases. Repeated translations also showed omissions, partial translations, and occasional misinterpretation.
- Suggestion quality: 37% of all cases received specific suggestions based on findings in the radiology report.An example addressed paranasal sinus disease by suggesting management of sinus symptoms and further discussion with clinicians.
- Suggestion quality: “Follow up with doctors” and “communicate the findings clearly to patient” were frequent suggestions for patients and healthcare providers, respectively.
- Translation robustness: 55.2% of information points were translated well across ten repeated chest CT translations, while 19.2% were omitted, 24.8% partially translated, and 0.8% misinterpreted.The repeated outputs used 25 key information points for point-by-point evaluation.
3.5 Optimized prompt for improved translation
The study addressed response variability by replacing a vague translation request with a detailed prompt specifying report structure and requiring complete coverage of findings. This clearer prompt substantially improved translation quality.
- ChatGPT produced different responses for the same input, and prompt ambiguity was identified as one reason for this variability.
- The optimized prompt specified four paragraphs covering screening context, detailed findings, conclusions, and incidental findings.It explicitly instructed ChatGPT not to omit information about findings.
- 77.2% of information points were translated well with the clearer prompt, up from 55.2%.Completely omitted, partially translated, and misinterpreted information fell to 9.2%, 13.6%, and 0%, respectively.
- 8 of 10 detailed-prompt translations preserved information about lung nodule 1, whereas none did so with the vague prompt.
3.6 Different prompts on ChatGPT’s performance
The study tested five modified prompts targeting different education levels or prompt designs and compared their results with the original and optimized prompts. All modified prompts performed similarly to the original and substantially worse than the optimized prompt.
- Five further-modified prompts were evaluated against the original and optimized prompts using the same response-assessment method.The modified prompts included education-level instructions and prompts designed by ChatGPT or PromptPerfect.
- All five modified prompts produced results similar to the original prompt and far worse results than the optimized prompt.
- The fourth, ChatGPT-designed prompt performed slightly better than the other four modified prompts, with a higher good rate and lower missing and inaccurate rates.It still performed significantly worse than the optimized prompt.
3.7 ChatGPT’s ensemble learning results
Ensemble learning combined five randomly selected ChatGPT translations into one report, but the resulting performance was generally limited and could lose minor findings.
- 3.7 ChatGPT’s ensemble learning results: 5 translated reports were randomly selected and integrated by ChatGPT into a single report for each case.The ensemble procedure was evaluated across 10 results.
- 3.7 ChatGPT’s ensemble learning results: The ensemble results were summarized as percentage changes relative to non-ensemble results.These comparisons are presented in Table 10.
- 3.7 ChatGPT’s ensemble learning results: The optimized prompt reduced the good rate while increasing missing and accurate rates in integrated results.The reported decline was attributed mainly to over-simplified lung-nodule findings and overlooked minor findings.
3.8 Comparison with GPT-4
The study compared GPT-4 with ChatGPT using the same radiology-report translation methodology and both original and optimized prompts. GPT-4 substantially improved translation quality, although some response randomness remained.
- 3.8 Comparison with GPT-4: GPT-4 was evaluated against ChatGPT using the original and optimized prompts with the same methodology.The comparison is reported in Table 11.
- 3.8 Comparison with GPT-4: GPT-4 significantly improved translated-report quality, with higher good rates and lower other rates under both prompts.The improvement was reported for both the original and optimized prompts.
- 3.8 Comparison with GPT-4: GPT-4 with the optimized prompt almost achieved a 100% good rate.GPT-4 using the original prompt was competitive with ChatGPT using the optimized prompt.
- 3.8 Comparison with GPT-4: GPT-4 still showed randomness, including one translation that placed an incidental finding in the wrong paragraph.The chronic obstructive pulmonary disease finding appeared in the third paragraph rather than the required fourth paragraph.
4 Discussions
ChatGPT produced concise, clear, and comprehensive plain-language radiology translations, but variability, omissions, and inconsistent formatting remain important barriers to clinical deployment.
- 4 Discussions: ChatGPT’s translations demonstrated conciseness, clarity, and comprehensiveness for radiology reports.It removed redundant wording, replaced complicated terminology, and integrated findings into understandable sentences.
- 4 Discussions: ChatGPT integrated information from different report sections into easily understandable sentences for readers with different education backgrounds.This capability was described as supporting comprehensibility.
- 4 Discussions: The same prompt and report could produce distinctive responses because of model randomness and vague instructions about which information to preserve.The vague prompt was associated with over-simplification and omitted important information, while detailed prompts improved results.
- 4 Discussions: ChatGPT lacks a built-in translation template, so prompts without format instructions can yield variable paragraph structures that are harder for patients to read.Specifying paragraph and word counts can improve consistency and readability.
- 4 Discussions: Radiologist evaluations found few missing or misinterpreted items, while GPT-4 produced a significant improvement over ChatGPT.The discussion presents these findings as evidence relevant to translation reliability and model progress.
- 4 Discussions: Incomplete, inconsistent, and potentially over-simplified translations remain concerns before clinical deployment.The optimized prompt improved completeness but did not make results perfect.
- 4 Discussions: The study presents radiology-report translation and automated suggestions as examples of potential large-language-model applications in healthcare.The passage also describes possible future uses, including report generation and treatment-option analysis.
- 4 Discussions: Safety evidence for medical-information tools will depend on intended use and the risks and benefits of those uses.Communication-support tools may be easier to demonstrate as safe than higher-risk applications.
5 Conclusion
The study found ChatGPT feasible for translating radiology reports into plain language and generating recommendations, while GPT-4 significantly improved translation quality and optimized prompts reduced omissions.
- 5 Conclusion: ChatGPT translations scored 4.268 on a five-point scale, with 0.097 missing and 0.065 incorrect information items per translation.The evaluation used professional assessors.
- 5 Conclusion: 55.2% of key points were completely translated with a vague prompt, increasing to 77.2% with an optimized prompt.The optimized prompt provided clearer and more specific instructions about information preservation.
- 5 Conclusion: GPT-4 significantly improved the quality of translated radiology reports compared with ChatGPT.This was the paper’s reported model comparison.
- 5 Conclusion: The findings support the feasibility of large language models for clinical education through plain-language translation and recommendations.The conclusion frames these applications as feasible while acknowledging the need for further effort.