상세 보기
SADDLE: A runtime feedback control architecture for adaptive distributed deep in GPU clusters
- Kim, Hyungjun;
- Lee, Eunyoung;
- Yu, Heonchang
WEB OF SCIENCE
0SCOPUS
1초록
Adaptive training in heterogeneous GPU clusters requires more than isolated heuristics-it demands a realtime, feedback-driven control system. SADDLE is a self-adaptive framework that unifies global batch scaling, local throughput balancing, and transient straggler mitigation into a fully coordinated runtime. It combines scaling guided by the Gradient Noise Scale (GNS), z-score detection over Exponentially Weighted Moving Average (EWMA)-smoothed iteration times, and responsiveness tuned via a Proportional-Integral-Derivative (PID) controller into a single, event-driven control loop. Across vision and language tasks, SADDLE improves training time by up to 2.84xand accuracy by up to 5.26% over strong baselines, while maintaining under 6% runtime overhead. This reframing positions adaptive training as dynamic system regulation, enabling deep learning frameworks to self-optimize under real-world heterogeneity.
키워드
- 제목
- SADDLE: A runtime feedback control architecture for adaptive distributed deep in GPU clusters
- 저자
- Kim, Hyungjun; Lee, Eunyoung; Yu, Heonchang
- 발행일
- 2025-11
- 유형
- Article
- 권
- 168