Multiple model-based reinforcement learning

被引:301
作者
Doya, K [1 ]
Samejima, K
Katagiri, K
Kawato, M
机构
[1] ATR Int, Human Informat Sci Labs, Sora Ku, Kyoto 6190288, Japan
[2] Japan Sci & Technol Corp, ERATO, Kawato Dynam Brain Project, Sora Ku, Kyoto 6190288, Japan
[3] Nara Inst Sci & Technol, Nara 6300101, Japan
关键词
D O I
10.1162/089976602753712972
中图分类号
TP18 [人工智能理论];
学科分类号
081104 ; 0812 ; 0835 ; 1405 ;
摘要
We propose a modular reinforcement learning architecture for nonlinear, nonstationary control tasks, which we call multiple model-based reinforcement learning (MMRL). The basic idea is to decompose a complex task into multiple domains in space and time based on the predictability of the environmental dynamics. The system is composed of multiple modules, each of which consists of a state prediction model and a reinforcement learning controller. The "responsibility signal," which is given by the softmax function of the prediction errors, is used to weight the outputs of multiple modules, as well as to gate the learning of the prediction models and the reinforcement learning controllers. We formulate MMRL for both discrete-time, finite-state case and continuous-time, continuous-state case. The performance of MMRL was demonstrated for discrete case in a nonstationary hunting task in a grid world and for continuous case in a nonlinear, nonstationary control task of swinging up a pendulum with variable physical parameters.
引用
收藏
页码:1347 / 1369
页数:23
相关论文
共 26 条
  • [1] NEURONLIKE ADAPTIVE ELEMENTS THAT CAN SOLVE DIFFICULT LEARNING CONTROL-PROBLEMS
    BARTO, AG
    SUTTON, RS
    ANDERSON, CW
    [J]. IEEE TRANSACTIONS ON SYSTEMS MAN AND CYBERNETICS, 1983, 13 (05): : 834 - 846
  • [2] Bertsekas DP, 2012, DYNAMIC PROGRAMMING, V2
  • [3] CACCIATORE TW, 1994, ADV NEURAL INFORMATI, V6
  • [4] Dayan P., 1992, Advances in Neural Information Processing Systems, V5, P271, DOI DOI 10.5555/2987061.2987095
  • [5] Reinforcement learning in continuous time and space
    Doya, K
    [J]. NEURAL COMPUTATION, 2000, 12 (01) : 219 - 245
  • [6] RECOGNITION OF MANIPULATED OBJECTS BY MOTOR LEARNING WITH MODULAR ARCHITECTURE NETWORKS
    GOMI, H
    KAWATO, M
    [J]. NEURAL NETWORKS, 1993, 6 (04) : 485 - 497
  • [7] Haruno M, 1999, ADV NEUR IN, V11, P31
  • [8] MOSAIC model for sensorimotor learning and control
    Haruno, M
    Wolpert, DM
    Kawato, M
    [J]. NEURAL COMPUTATION, 2001, 13 (10) : 2201 - 2220
  • [9] Human cerebellar activity reflecting an acquired internal model of a new tool
    Imamizu, H
    Miyauchi, S
    Tamada, T
    Sasaki, Y
    Takino, R
    Pütz, B
    Yoshioka, T
    Kawato, M
    [J]. NATURE, 2000, 403 (6766) : 192 - 195
  • [10] IMAMIZU H, 1997, NEUROIMAGE, V5