Toward Transparent and Reproducible Machine Learning-Based Research in Urban Water Networks

Citations

WEB OF SCIENCE

2
Citations

SCOPUS

2

초록

Numerous machine learning (ML) models have been developed for various applications in urban water networks (i.e., water distribution systems and urban drainage networks), including state estimation and failure detection. However, although pre-model setup procedures, such as data handling and model configuration, are crucial for ensuring the optimal performance of ML models, less attention has been paid to them. These steps are also important for promoting the transparency and reproducibility of ML models applied to urban water networks. To fill these gaps, this study reviews ML-related papers published in the Journal of Water Resources Planning and Management over the past five years and summarizes the critical limitations identified in the pre-model setup processes for ML modeling. This procedure can be categorized into three main steps: data collection, data preprocessing, and model optimization. Data preprocessing is further categorized into four steps: outlier detection and removal, missing data management, data splitting, and feature engineering and scaling. An evaluation table is provided to highlight the common mistakes and gaps in each pre-model setup category. Finally, this paper presents comprehensive guidelines for reporting ML workflows, offering insights into the effective management and stewardship of ML data and models to enhance transparency and reproducibility.

키워드

Deep learningReproducibilityTransparencyUrban drainage networkWater distribution system
제목
Toward Transparent and Reproducible Machine Learning-Based Research in Urban Water Networks
저자
Jun, SanghoonYoo, Do GuenJung, Donghwi
DOI
10.1061/JWRMD5.WRENG-6868
발행일
2025-10-01
유형
Article
저널명
Journal of Water Resources Planning and Management - ASCE
151
10