Multiple model-based reinforcement learning

被引：301

作者：

Doya, K ^{[1
]}

Samejima, K

Katagiri, K

Kawato, M

机构：

[1] ATR Int, Human Informat Sci Labs, Sora Ku, Kyoto 6190288, Japan

[2] Japan Sci & Technol Corp, ERATO, Kawato Dynam Brain Project, Sora Ku, Kyoto 6190288, Japan

[3] Nara Inst Sci & Technol, Nara 6300101, Japan

来源：

NEURAL COMPUTATION | 2002年 / 14卷 / 06期

关键词：

D O I：

10.1162/089976602753712972

中图分类号：

TP18 [人工智能理论];

学科分类号：

081104 ; 0812 ; 0835 ; 1405 ;

摘要：

We propose a modular reinforcement learning architecture for nonlinear, nonstationary control tasks, which we call multiple model-based reinforcement learning (MMRL). The basic idea is to decompose a complex task into multiple domains in space and time based on the predictability of the environmental dynamics. The system is composed of multiple modules, each of which consists of a state prediction model and a reinforcement learning controller. The "responsibility signal," which is given by the softmax function of the prediction errors, is used to weight the outputs of multiple modules, as well as to gate the learning of the prediction models and the reinforcement learning controllers. We formulate MMRL for both discrete-time, finite-state case and continuous-time, continuous-state case. The performance of MMRL was demonstrated for discrete case in a nonstationary hunting task in a grid world and for continuous case in a nonlinear, nonstationary control task of swinging up a pendulum with variable physical parameters.

引用

页码：1347 / 1369

页数：23

共 26 条

[11] Adaptive Mixtures of Local Experts
Jacobs, Robert A.
Jordan, Michael I.
Nowlan, Steven J.
Hinton, Geoffrey E.
[J]. NEURAL COMPUTATION, 1991, 3 (01) : 79 - 87
[12] Littman M. L., 1995, Machine Learning. Proceedings of the Twelfth International Conference on Machine Learning, P362
[13] Acquisition of stand-up behavior by a real robot using hierarchical reinforcement learning
Morimoto, J
Doya, K
[J]. ROBOTICS AND AUTONOMOUS SYSTEMS, 2001, 36 (01) : 37 - 51
[14] ADAPTATION AND LEARNING USING MULTIPLE MODELS, SWITCHING, AND TUNING
NARENDRA, KS
BALAKRISHNAN, J
CILIZ, MK
[J]. IEEE CONTROL SYSTEMS MAGAZINE, 1995, 15 (03): : 37 - 51
[15] Parr R, 1998, ADV NEUR IN, V10, P1043
[16] Annealed competition of experts for a segmentation and classification of switching dynamics
Pawelzik, K
Kohlmorgen, J
Muller, KR
[J]. NEURAL COMPUTATION, 1996, 8 (02) : 340 - 356
[17] Schaal S, 1996, ADV NEUR IN, V8, P605
[18] SINGH SP, 1992, MACH LEARN, V8, P323, DOI 10.1007/BF00992700
[19] Sutton R. S., 1988, Machine Learning, V3, P9, DOI 10.1023/A:1022633531479
[20] Sutton R. S., 1998, Reinforcement Learning: An Introduction, V22447

← 1 2 3 →