Interpretable reinforcement-learning energy dispatch co-optimized with system sizing for stable and efficient off-grid green-hydrogen systems

  • Yun, Byeongchan; 
  • Lee, Suin; 
  • Yoo, Yeon-Shick; 
  • Shim, Jaegyu; 
  • Cho, Kyung Hwa
Citations

WEB OF SCIENCE

0
Citations

SCOPUS

0

초록

In off-grid green-hydrogen systems, the absence of a grid buffer couples component sizing to operation, demanding joint optimization. Existing energy management relies on rule-based control built upon expert heuristics or on model predictive control (MPC), whose forecast dependence limits autonomous deployment. We propose a bi-level framework that jointly designs an off-grid photovoltaic-wind turbine-energy storage system-proton-exchange membrane water electrolysis (PV-WT-ESS-PEMWE) system and its operating policy. The outer loop sizes the ESS, PV, and WT capacities via black-box optimization (random search and covariance matrix adaptation evolution strategy), while the inner loop trains a proximal-policy-optimization reinforcement-learning (RL) agent to dispatch power within a high-fidelity electrochemical-thermal, zero-dimensional differential-algebraic PEMWE simulation. A model-fidelity ablation confirmed the necessity of this thermal fidelity. The framework was demonstrated using three years of meteorological data from Jeju City, Republic of Korea. Against the rule-based baseline at statistically indistinguishable hydrogen output (p = 0.052) across twelve 60-day scenarios, the RL agent reduced on-off cycling by similar to 94% (p < 0.001) and curtailment by similar to 41% (p = 0.002). Against a receding-horizon MPC suite, the forecast-free RL agent surpassed both persistence-forecast variants. SHapley Additive exPlanations interpreted the learned strategies, including anticipatory overnight energy reserve, pre-dawn stack pre-heating, and thermal self-regulation, thereby supporting transparent, trustworthy AI-based energy management. The optimized configuration achieved a levelized hydrogen cost of 13.1 USD kg(-1)H(2), and Sobol' analysis identified WT sizing as the dominant cost driver (S-1 = 0.745). The framework offers a reproducible route to co-designing capacity and operation for autonomous green hydrogen production.

키워드

Green hydrogen; Renewable energy; Proton-exchange membrane water electrolysis; Bi-level optimization; Reinforcement learning; Techno-economic analysis; ELECTROLYZER; PERFORMANCE; STORAGE
제목
Interpretable reinforcement-learning energy dispatch co-optimized with system sizing for stable and efficient off-grid green-hydrogen systems
저자
Yun, Byeongchan; Lee, Suin; Yoo, Yeon-Shick; Shim, Jaegyu; Cho, Kyung Hwa
DOI
10.1016/j.apenergy.2026.128682
발행일
2026-12
유형
Article
저널명
Applied Energy
권
426