REX: Adaptive distributed training with statistically validated batch scaling and spline-guided Bayesian optimization on heterogeneous GPU clusters

Citations

WEB OF SCIENCE

0
Citations

SCOPUS

0

초록

Training deep neural networks on heterogeneous GPU clusters often suffers from idle time and delayed convergence due to worker speed imbalance and transient slowdowns. We present REX, a runtime framework for adaptive distributed training that jointly adapts global and per-worker batch sizes using lightweight, statistically robust signals. REX combines exponential global batch scaling guided by gradient noise scale (GNS), spline-guided Bayesian optimization for per-worker scheduling, and event-triggered reallocation via a two-level nonparametric test to handle transient stragglers. Across CIFAR-10/100, ImageNet/ResNet-50, and WikiText-2 on mixed T4/V100 clusters, REX reduces time-to-accuracy by up to 50.4% under static heterogeneity and 44.2% under transient slowdowns; on ImageNet it achieves 15.9-25.3% faster convergence on 4-16 GPUs with accuracy gains up to +0.27%. REX also lowers the Resource Inefficiency Indicator (RII) while keeping control overhead under 9%. These results position REX as a low-overhead, statistically grounded alternative to heuristic adaptive training, providing robust convergence and scalability in heterogeneous GPU clouds.

키워드

Adaptive distributed training; Heterogeneous GPU clusters; Batch size scaling; Gradient noise scale; Spline regression; Bayesian optimization; Nonparametric statistical tests; Event-triggered scheduling; Straggler mitigation; Resource Inefficiency Indicator (RII)
제목
REX: Adaptive distributed training with statistically validated batch scaling and spline-guided Bayesian optimization on heterogeneous GPU clusters
저자
Kim, HyungJun; Lee, Hwamin; Yu, Heonchang
DOI
10.1016/j.future.2026.108659
발행일
2026-12
유형
Article
저널명
Future Generation Computer Systems
권
185