SADDLE: A runtime feedback control architecture for adaptive distributed deep in GPU clusters

Citations

WEB OF SCIENCE

0
Citations

SCOPUS

1

초록

Adaptive training in heterogeneous GPU clusters requires more than isolated heuristics-it demands a realtime, feedback-driven control system. SADDLE is a self-adaptive framework that unifies global batch scaling, local throughput balancing, and transient straggler mitigation into a fully coordinated runtime. It combines scaling guided by the Gradient Noise Scale (GNS), z-score detection over Exponentially Weighted Moving Average (EWMA)-smoothed iteration times, and responsiveness tuned via a Proportional-Integral-Derivative (PID) controller into a single, event-driven control loop. Across vision and language tasks, SADDLE improves training time by up to 2.84xand accuracy by up to 5.26% over strong baselines, while maintaining under 6% runtime overhead. This reframing positions adaptive training as dynamic system regulation, enabling deep learning frameworks to self-optimize under real-world heterogeneity.

키워드

Feedback-driven adaptation; Control-theoretic training; Global-local batch optimization; Event-triggered rebalancing; Online parameter tuning; Heterogeneous GPU clusters
제목
SADDLE: A runtime feedback control architecture for adaptive distributed deep in GPU clusters
저자
Kim, Hyungjun; Lee, Eunyoung; Yu, Heonchang
DOI
10.1016/j.sysarc.2025.103573
발행일
2025-11
유형
Article
저널명
Journal of Systems Architecture
권
168