Hierarchical End-to-end Control Policy for Multi-degree-of-freedom Manipulators

  • Min, Cheol-Hui
  • Song, Jae-Bok
Citations

WEB OF SCIENCE

8
Citations

SCOPUS

9

초록

In recent years, several control policies for a multi-degree-of-freedom (DOF) manipulator using deep reinforcement learning have been proposed. To avoid complexity, previous studies have applied a number of constraints on the high-dimensional state-action space, thus hindering generalized policy function learning. In this study, the control problem is addressed by in-troducing a hierarchical reinforcement learning method that can learn the end-to-end control policy of a multi-DOF manipula-tor without any constraints on the state-action space. The proposed method learns hierarchical policy using two off-policy methods. Using human demonstration data and a newly proposed data-correction method, controlling the multi-DOF manipu-lator in an end-to-end manner is shown to outperform the non-hierarchical deep reinforcement learning methods.

키워드

Deep reinforcement learningdemonstration-based learningend-to-end robot controlhierarchical reinforcement learning
제목
Hierarchical End-to-end Control Policy for Multi-degree-of-freedom Manipulators
저자
Min, Cheol-HuiSong, Jae-Bok
DOI
10.1007/s12555-021-0511-4
발행일
2022-10
유형
Article
저널명
International Journal of Control, Automation, and Systems
20
10
페이지
3296 ~ 3311