Source-linked AI summary
Japanese Stroke LLM Evaluation: A Conversational Benchmark for Safe Stroke Care in Japanese Using Large Language Models
Keisuke Masuda, Kazutaka Yatsushiro, Hirohumi Iwamoto, Hirofumi Hirano, Ryosuke Hanaya
TL;DR
LLMs need evaluation beyond knowledge tests because clinical history taking, urgency assessment, and safety remain insufficiently assessed. This study introduces a Japanese, specialist-scored conversational stroke benchmark using standardized cases and action phases. Claude Fable 5 and Claude Opus 4.7 met the prespecified safety threshold, while critical mistakes remained in several models and broader real-world evaluation is still needed.
Problem
Clinical history taking, urgency assessment, and safety of LLMs remain insufficiently evaluated beyond medical knowledge tests.
Method
The benchmark evaluated 18 LLMs in standardized Japanese, multi-turn conversations across 10 stroke-related cases, using specialist scoring of history-taking and action phases with critical-mistake identification.
Results
Claude Fable 5 scored 87.4% with zero critical mistakes and Claude Opus 4.7 scored 80.3% with zero critical mistakes, making both meet the safety threshold.
Takeaways & Limitations
Japanese Stroke LLM Evaluation provides a practice-oriented benchmark, and the latest evaluated models showed improved performance with some exceeding the safety threshold.
Takeaways & Limitations
The pilot used only 10 cases, one run per model, and one physician as both simulated patient and evaluator.
Abstract
from arXiv · showhide
Background: Large language models (LLMs) have achieved physician-comparable performance on multiple-choice medical knowledge examinations, but their capabilities in clinical history taking, urgency assessment, and safety remain insufficiently evaluated. We proposed Japanese Stroke LLM Evaluation, a multi-turn conversational benchmark for stroke care in Japanese, and evaluated LLM performance and safety under practice-oriented conditions. Methods: We created 10 stroke and related-condition cases and evaluated LLMs in multi-turn Japanese conversations. The LLM acted as physician, while a board-certified neurosurgeon acted as simulated patient and evaluator. Each case comprised history-taking and action phases scored using pre-specified criteria. Errors that could directly threaten life were defined as critical mistakes. The safety threshold was at least 80% overall with zero critical mistakes. Eighteen models were evaluated in October 2025 and June 2026. Results: Claude Fable 5 achieved the highest score (87.4%) with zero critical mistakes, followed by Claude Opus 4.7 (80.3%) and GLM-5.2 (75.6%). Two leaders met the safety threshold. Eleven models made 17 critical mistakes, including failure to confirm laboratory results or blood glucose before t-PA, surgery before airway stabilization, omission of cervical vascular evaluation, and t-PA outside its indication. History-taking question count correlated with history-taking score (r = 0.648, p = 0.007). Conclusions: Japanese Stroke LLM Evaluation provides a benchmark for LLM performance under practice-oriented conditions, including a cap on history-taking questions. Cases and evaluations were created by neurosurgical specialists rather than using an LLM-as-judge approach. Performance improved across cloud-based and on-premise models in 2026, with some exceeding the safety threshold. Further evaluation using real-world cases is required.
3 Iwamoto Neurosurgery, Kanoya, Kagoshima, Japan
The paper concerns large language models, medical benchmarking, stroke care, patient safety, and Japanese evaluation.
- The study addresses large language models in a medical benchmark context.
- Stroke care is evaluated as a clinical application domain.
- Patient safety and Japanese-language evaluation are central themes.
1. Introduction
Existing LLM medical evaluations often emphasize reproducible knowledge tests or automated judging, while practical, safety-focused Japanese clinical assessment remains necessary. Japanese Stroke LLM Evaluation addresses this gap with a specialist-scored, multi-turn benchmark for stroke care.
- 1. Introduction: MCQs are reproducible but poorly reflect real clinical scenarios, while LLM-as-judge methods can diverge from specialist human evaluations because of systematic biases.
- 1. Introduction: LLMs may perform well on knowledge tests yet substantially worse on practical and safety evaluations, motivating multiturn clinical assessment.
- 1. Introduction: Japanese Stroke LLM Evaluation assesses stroke-care performance through Japanese conversations, with a physician acting as simulated patient and a specialist performing scoring.
- 1. Introduction: Japanese evaluation is needed because language, medication dosing, insurance systems, and patient-explanation customs differ from other clinical contexts.
- 1. Introduction: Stroke care requires rapid integration of history taking, treatment, urgency assessment, and safety judgments under time pressure.
2. Methods
The benchmark evaluates LLMs as physicians in standardized Japanese conversations across ten stroke-related cases, separating history taking from action selection and explicitly scoring life-threatening errors. Its design uses capped questioning, sequential actions, predefined criteria, and specialist-created cases.
- 2.1 Benchmark overview: Ten Japanese cases were scored across 50-point history-taking and 50-point action phases, with 80% plus zero critical mistakes defining safe support.
- 2.2 History-taking and action phases: Models asked one history-taking item per response, up to 20 questions, then summarized the record before proposing sequential tests or actions.
- 2.3 Critical mistakes: Critical mistakes were case-specific errors that could directly affect survival, including unsafe thrombolysis, airway neglect, and omitted vascular evaluation.
- 2.4 Cases: The study used newly created fictional cases, publicly available case materials and scoring criteria, and model cohorts evaluated in October 2025 and June 2026.
- 2.2 History-taking and action phases: All models received the same minimal prompt, including one-question and one-action response rules, a 20-question cap, and explicit critical-mistake recognition.
- 2.5 Statistical analysis: The question-count analysis used Pearson correlation across eight non-emergency cases, excluding two emergency cases and models with incomplete records.
3. Results
Claude Fable 5 led the 18-model benchmark and was one of two models meeting the safety threshold. Critical mistakes remained common, while history-taking performance was associated with asking more questions but also depended on prioritization and information integration.
- 3.1 Overall performance: 87.4%: Claude Fable 5 achieved the highest overall score with zero critical mistakes, followed by Claude Opus 4.7 at 80.3% with zero critical mistakes.Only these two models met the prespecified safety threshold of at least 80% overall with zero critical mistakes.
- 3.1 Overall performance: 69.5% versus 48.7%: the mean overall score was higher for models evaluated in June 2026 than for those evaluated in October 2025.Each group contained nine models.
- 3.1 Overall performance: Models varied in history-taking and action performance, and higher overall scores did not necessarily indicate safety.For example, Gemini 2.5 Pro scored 71.0% in action but 46.8% in history-taking, while DeepSeek-V4-Pro scored 64.3% overall yet committed critical mistakes.
- 3.2 Critical mistakes: 17 critical errors occurred across 11 of 18 models, including unsafe t-PA decisions, surgery before airway stabilization, and omitted cervical vascular evaluation.Reported t-PA errors included missing laboratory or blood glucose checks, Western rather than Japanese dosing, and use outside indication.
- 3.4 Association between the number of history-taking questions and history-taking score: Performance depended on question prioritization and information integration, not merely on asking more questions.Some models asked many questions without achieving high history-taking scores, while others appeared to reach diagnoses or treatment plans prematurely.
4. Discussion
The benchmark addresses practical and safety gaps in Japanese stroke-care evaluation through constrained, expert-scored conversations. Results show improvement, but critical mistakes and pilot-design limitations prevent claims of clinical readiness.
- 4.1 Significance of the benchmark: Models evaluated in June 2026 averaged approximately 21 percentage points higher overall scores than models evaluated in October 2025.Open-weight models deployable on premises may support operations without transmitting medical information outside the institution.
- 4.2 Premature closure and the quality of history-taking: Question count correlated positively with history-taking score, but longer questioning alone did not ensure high performance.The benchmark limits questions to address verbosity bias and evaluates question prioritization and information integration alongside count.
- 4.3 Critical mistakes and safety evaluation: Critical mistakes mainly involved omitted safety checks, incorrect thrombolytic dosing, and failure to prioritize airway stabilization.The discussion argues that critical mistakes require assessment independent of mean scores, with multiple trials needed because model responses are nondeterministic.
- 4.4 Regional and language-specific safety: Japanese and Western clinical differences, particularly t-PA dosing, make English-language benchmark performance insufficient to establish capability in Japanese practice.The discussion supports region- and language-specific evaluation because translated criteria may mismatch Japanese clinical context.
- 4.5 Limitations: The pilot used 10 cases and one evaluator, without assessing inter-rater reliability.The authors recommend larger case sets, repeated evaluations, independent dual scoring, and agreement testing for critical-mistake judgments.
5. Conclusions
Japanese Stroke LLM Evaluation measures integrated stroke-care capabilities through expert-evaluated, constrained Japanese conversations. Although newer models improved and two reached the safety threshold, critical mistakes remain, requiring region-specific interactive safety evaluation before clinical support use.
- 5. Conclusions: The benchmark measures history-taking, urgency assessment, action selection, and avoidance of major errors using human evaluation by a stroke-care physician.It incorporates a question cap, expert assessment, and explicit critical-mistake identification.
- 5. Conclusions: Claude Fable 5 and Claude Opus 4.7 reached the prespecified safety threshold, but even the latest models produced life-threatening critical mistakes.The conclusion highlights differences in Japanese and Western practice, particularly tPA dosing, as a source of serious errors.
- 5. Conclusions: Interactive safety evaluations specific to regional, language, and disease domains are necessary before LLMs support stroke care.The study is presented as a preparatory investigation for such evaluations.
9. Supplementary
The supplementary materials specify the Japanese conversational protocol, scoring structure, deductions, and critical-mistake handling. They operationalize a 100-point-per-case evaluation with capped history taking and sequential actions.
- 9. Supplementary: Each case assigns 50 points to history taking and 50 points to action, for 100 points per case and 1,000 points across 10 cases.Scores are calculated across the history-taking and action phases, with the case score unable to fall below zero.
- 9. Supplementary: All conversations are conducted in Japanese, and the evaluator standardizes responses to neurological-examination requests according to case-specific disclosure rules.The full prompt and deduction criteria are provided in supplementary tables.
- 9. Supplementary: The history-taking phase permits one item per question and at most 20 questions, with questions organized by importance and unnecessary questions avoided.In highly urgent situations, early termination of history taking is evaluated.
- 9. Supplementary: During the action phase, models propose one test or action at a time and receive the corresponding result from the physician.Encounters end at case-specific criteria such as referral, admission, discharge, or emergency surgery.
- 9. Supplementary: A critical mistake zeros only the phase in which it occurs, while the other phase is scored independently.Critical mistakes are case-defined major errors that could directly affect survival, including unsafe t-PA decisions and airway-stabilization failures.
- 9. Supplementary: Deductions cover excess questions, contradictory or incomplete records, incorrect diagnoses, unavailable or low-value tests, multiple simultaneous actions, language violations, and premature completion.The supplementary criteria also specify penalties for technical terms, typographical errors, loops, and other operational failures.