Multiple model-based reinforcement learning

被引：301

作者：

Doya, K ^{[1
]}

Samejima, K

Katagiri, K

Kawato, M

机构：

[1] ATR Int, Human Informat Sci Labs, Sora Ku, Kyoto 6190288, Japan

[2] Japan Sci & Technol Corp, ERATO, Kawato Dynam Brain Project, Sora Ku, Kyoto 6190288, Japan

[3] Nara Inst Sci & Technol, Nara 6300101, Japan

来源：

NEURAL COMPUTATION | 2002年 / 14卷 / 06期

关键词：

D O I：

10.1162/089976602753712972

中图分类号：

TP18 [人工智能理论];

学科分类号：

081104 ; 0812 ; 0835 ; 1405 ;

摘要：

We propose a modular reinforcement learning architecture for nonlinear, nonstationary control tasks, which we call multiple model-based reinforcement learning (MMRL). The basic idea is to decompose a complex task into multiple domains in space and time based on the predictability of the environmental dynamics. The system is composed of multiple modules, each of which consists of a state prediction model and a reinforcement learning controller. The "responsibility signal," which is given by the softmax function of the prediction errors, is used to weight the outputs of multiple modules, as well as to gate the learning of the prediction models and the reinforcement learning controllers. We formulate MMRL for both discrete-time, finite-state case and continuous-time, continuous-state case. The performance of MMRL was demonstrated for discrete case in a nonstationary hunting task in a grid world and for continuous case in a nonlinear, nonstationary control task of swinging up a pendulum with variable physical parameters.

引用

页码：1347 / 1369

页数：23

共 26 条

[1] NEURONLIKE ADAPTIVE ELEMENTS THAT CAN SOLVE DIFFICULT LEARNING CONTROL-PROBLEMS
BARTO, AG
SUTTON, RS
ANDERSON, CW
[J]. IEEE TRANSACTIONS ON SYSTEMS MAN AND CYBERNETICS, 1983, 13 (05): : 834 - 846
[2] Bertsekas DP, 2012, DYNAMIC PROGRAMMING, V2
[3] CACCIATORE TW, 1994, ADV NEURAL INFORMATI, V6
[4] Dayan P., 1992, Advances in Neural Information Processing Systems, V5, P271, DOI DOI 10.5555/2987061.2987095
[5] Reinforcement learning in continuous time and space
Doya, K
[J]. NEURAL COMPUTATION, 2000, 12 (01) : 219 - 245
[6] RECOGNITION OF MANIPULATED OBJECTS BY MOTOR LEARNING WITH MODULAR ARCHITECTURE NETWORKS
GOMI, H
KAWATO, M
[J]. NEURAL NETWORKS, 1993, 6 (04) : 485 - 497
[7] Haruno M, 1999, ADV NEUR IN, V11, P31
[8] MOSAIC model for sensorimotor learning and control
Haruno, M
Wolpert, DM
Kawato, M
[J]. NEURAL COMPUTATION, 2001, 13 (10) : 2201 - 2220
[9] Human cerebellar activity reflecting an acquired internal model of a new tool
Imamizu, H
Miyauchi, S
Tamada, T
Sasaki, Y
Takino, R
Pütz, B
Yoshioka, T
Kawato, M
[J]. NATURE, 2000, 403 (6766) : 192 - 195
[10] IMAMIZU H, 1997, NEUROIMAGE, V5

← 1 2 3 →