상세 보기
The Effects of Prompt Engineering on a Multimodal Large Language Model in Neuroradiology: A Case Study of GPT-4o for Intracranial Hemorrhage Diagnosis from Brain CT
- Kim, Young-Tak;
- Kim, Hayom;
- Kim, Hyunji;
- Lee, Seho;
- Lim, Dong-Kwon;
- ... Han, Jae-Ho;
- 외 2명
WEB OF SCIENCE
0SCOPUS
0초록
Multimodal large language models (MLLMs) connect vision and language, yet their diagnostic output can vary with prompt design. We examined how prompting influences GPT-4o for zero-shot intracranial hemorrhage (ICH) classification on non-contrast head CT. Using 60 axial slices from the CQ500 dataset (10 each: intraparenchymal [IPH], intraventricular [IVH], subdural [SDH], epidural [EDH], subarachnoid [SAH]; and 10 normal; one patient per case), we compared five prompts: free response, minimal lesion query, forced-choice among ICH subtypes, chain-of-thought, and self-correction. Performance was measured on a 3-class task (intra-axial, extra-axial, normal) and a 6-class task (five ICH subtypes plus normal). The forced-choice prompt yielded the best results: accuracy of 0.67 and macro-F1 of 0.65 (3-class), and accuracy of 0.42 and macro-F1 of 0.39 (6-class). Chain-of-thought and self-correction did not improve macro-F1 over forced-choice; macro precision rose slightly while accuracy, macro-F1, and the normal-class F1 decreased. Gains were concentrated in IPH and IVH; EDH/SDH improved little, and SAH remained below chance level. Furthermore, an analysis of linguistic sensitivity revealed a strong correlation (R2 = 0.94) between prediction consistency and diagnostic performance, suggesting that the model's robustness to prompt phrasing is intrinsically linked to the visual saliency of the pathology. These findings indicate that constrained, task-aligned prompting can raise zero-shot performance of a general MLLM, but errors among visually similar subtypes persist. Safe clinical integration will require domain-specific fine-tuning beyond prompt design. © 2026 KSII.
키워드
- 제목
- The Effects of Prompt Engineering on a Multimodal Large Language Model in Neuroradiology: A Case Study of GPT-4o for Intracranial Hemorrhage Diagnosis from Brain CT
- 저자
- Kim, Young-Tak; Kim, Hayom; Kim, Hyunji; Lee, Seho; Lim, Dong-Kwon; Han, Jae-Ho; Kim, Jung Bin; Do, Synho
- 발행일
- 2026-04-30
- 유형
- Article
- 권
- 20
- 호
- 4
- 페이지
- 1966 ~ 1983