ReAx: Resource-Efficient Asynchronous Execution for Accelerating LLM Fine-Tuning at the Edge

Citations

WEB OF SCIENCE

1
Citations

SCOPUS

1

초록

With the widespread use of large language models (LLMs), there is an increasing demand for personalized models that meet diverse needs of users. To enable private, network-independent personalization of LLMs, on-device fine-tuning is receiving much attention. However, on-device fine-tuning faces efficiency and scalability challenges as sequential execution of compute- and memory-intensive operations often underuses resources. In this letter, we propose ReAx, a framework that accelerates on-device fine-tuning through resource-efficient asynchronous parallel execution of memory- and compute-intensive operations. Without increasing memory usage, ReAx improves the average fine-tuning performance and energy consumption by 10.42% and 5.55%, respectively, compared with the baseline. As a positive side effect, asynchronous parameter updates induce gradient noise due to slight delays between streams, which act as a regularizer for adverse updates minimizing accuracy drops.

키워드

Adaptation models; Resource management; Computational modeling; Training; Synchronization; Graphics processing units; Memory management; Servers; Data models; Accuracy; Asynchronous parallel execution of operations; fine-tuning; large language model (LLM); on-device AI
제목
ReAx: Resource-Efficient Asynchronous Execution for Accelerating LLM Fine-Tuning at the Edge
저자
Na, Hyukju; Choi, Daeseon; Gong, Young-Ho; Kim, Young Geun
DOI
10.1109/LES.2025.3640585
발행일
2026-08
유형
Article
저널명
IEEE Embedded Systems Letters
권
18
호
4
페이지
296 ~ 299