상세 보기
초록
In response to the rapid onset of super-aged societies and the growing prevalence of dementia, our research proposes a multimodal, on-device large language model (LLM)-based agent designed for comprehensive elderly care. Unlike traditional models focused only on dementia diagnosis, our approach leverages a compact, privacy-preserving LLM as the core technology, enabling a wide range of support functions-including cognitive assessment-directly on the device. The pipeline begins with automatic speech recognition (ASR) to convert user speech into text, followed by a vision-language model (VLM) that generates contextual cues from provided images through visual question answering (VQA). We then employ Chain-of-Thought (CoT) reasoning during supervised fine-tuning (SFT) of the LLM, using the generated cues to improve the classification of Alzheimer's disease (AD) versus non-AD cases. A linear layer is added to the LLM for final binary classification. Our results show a 16.7% relative improvement in performance when CoT reasoning is applied compared to when it is not. This multimodal, on-device strategy enables real-time, comprehensive support for the elderly-such as early intervention and integrated cognitive assessment-without reliance on cloud services, thereby laying the groundwork for advanced, privacy-friendly care solutions tailored to super-aged societies.
키워드
- 제목
- A Novel Chain-of-Thought Reasoning Approach for Alzheimer's Disease Detection Using Large Language and Vision-Language Models
- 저자
- Park, Chanwoo; Kim, Chanwoo
- 발행일
- 2025
- 유형
- Article
- 권
- 33
- 페이지
- 4386 ~ 4395