상세 보기
REX: Adaptive distributed training with statistically validated batch scaling and spline-guided Bayesian optimization on heterogeneous GPU clusters
- Kim, HyungJun;
- Lee, Hwamin;
- Yu, Heonchang
WEB OF SCIENCE
0SCOPUS
0초록
Training deep neural networks on heterogeneous GPU clusters often suffers from idle time and delayed convergence due to worker speed imbalance and transient slowdowns. We present REX, a runtime framework for adaptive distributed training that jointly adapts global and per-worker batch sizes using lightweight, statistically robust signals. REX combines exponential global batch scaling guided by gradient noise scale (GNS), spline-guided Bayesian optimization for per-worker scheduling, and event-triggered reallocation via a two-level nonparametric test to handle transient stragglers. Across CIFAR-10/100, ImageNet/ResNet-50, and WikiText-2 on mixed T4/V100 clusters, REX reduces time-to-accuracy by up to 50.4% under static heterogeneity and 44.2% under transient slowdowns; on ImageNet it achieves 15.9-25.3% faster convergence on 4-16 GPUs with accuracy gains up to +0.27%. REX also lowers the Resource Inefficiency Indicator (RII) while keeping control overhead under 9%. These results position REX as a low-overhead, statistically grounded alternative to heuristic adaptive training, providing robust convergence and scalability in heterogeneous GPU clouds.
키워드
- 제목
- REX: Adaptive distributed training with statistically validated batch scaling and spline-guided Bayesian optimization on heterogeneous GPU clusters
- 저자
- Kim, HyungJun; Lee, Hwamin; Yu, Heonchang
- 발행일
- 2026-12
- 유형
- Article
- 권
- 185