Modular deep reinforcement learning from reward and punishment for robot navigation

Jiexin Wang¹, Stefan Elfwing², Eiji Uchibe³

Affiliations

¹ Department of Brain Robot Interface, ATR Computational Neuroscience Laboratories, 2-2-2 Hikaridai, Seikacho, Soraku-gun, Kyoto 619-0288, Japan. Electronic address: wang-j@atr.jp.
² ContextVision AB, Storgatan 24, 582 23 Linkoping, Sweden. Electronic address: stefan.elfwing@contextvision.se.
³ Department of Brain Robot Interface, ATR Computational Neuroscience Laboratories, 2-2-2 Hikaridai, Seikacho, Soraku-gun, Kyoto 619-0288, Japan. Electronic address: uchibe@atr.jp.

PMID: 33383526
DOI: 10.1016/j.neunet.2020.12.001

Free article

Modular deep reinforcement learning from reward and punishment for robot navigation

Jiexin Wang et al. Neural Netw. 2021 Mar.

Free article

. 2021 Mar:135:115-126.

doi: 10.1016/j.neunet.2020.12.001. Epub 2020 Dec 8.

Authors

Jiexin Wang¹, Stefan Elfwing², Eiji Uchibe³

Affiliations

¹ Department of Brain Robot Interface, ATR Computational Neuroscience Laboratories, 2-2-2 Hikaridai, Seikacho, Soraku-gun, Kyoto 619-0288, Japan. Electronic address: wang-j@atr.jp.
² ContextVision AB, Storgatan 24, 582 23 Linkoping, Sweden. Electronic address: stefan.elfwing@contextvision.se.
³ Department of Brain Robot Interface, ATR Computational Neuroscience Laboratories, 2-2-2 Hikaridai, Seikacho, Soraku-gun, Kyoto 619-0288, Japan. Electronic address: uchibe@atr.jp.

PMID: 33383526
DOI: 10.1016/j.neunet.2020.12.001

Abstract

Modular Reinforcement Learning decomposes a monolithic task into several tasks with sub-goals and learns each one in parallel to solve the original problem. Such learning patterns can be traced in the brains of animals. Recent evidence in neuroscience shows that animals utilize separate systems for processing rewards and punishments, illuminating a different perspective for modularizing Reinforcement Learning tasks. MaxPain and its deep variant, Deep MaxPain, showed the advances of such dichotomy-based decomposing architecture over conventional Q-learning in terms of safety and learning efficiency. These two methods differ in policy derivation. MaxPain linearly unified the reward and punishment value functions and generated a joint policy based on unified values; Deep MaxPain tackled scaling problems in high-dimensional cases by linearly forming a joint policy from two sub-policies obtained from their value functions. However, the mixing weights in both methods were determined manually, causing inadequate use of the learned modules. In this work, we discuss the signal scaling of reward and punishment related to discounting factor γ, and propose a weak constraint for signaling design. To further exploit the learning models, we propose a state-value dependent weighting scheme that automatically tunes the mixing weights: hard-max and softmax based on a case analysis of Boltzmann distribution. We focus on maze-solving navigation tasks and investigate how two metrics (pain-avoiding and goal-reaching) influence each other's behaviors during learning. We propose a sensor fusion network structure that utilizes lidar and images captured by a monocular camera instead of lidar-only and image-only sensing. Our results, both in the simulation of three types of mazes with different complexities and a real robot experiment of an L-maze on Turtlebot3 Waffle Pi, showed the improvements of our methods.

Keywords: Deep reinforcement learning; Max pain; Maze solving; Modular reinforcement learning; Robot navigation.

PubMed Disclaimer

Conflict of interest statement

Declaration of Competing Interest The authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper.

Cited by

Advances in non-invasive biosensing measures to monitor wound healing progression.
Short WD, Olutoye OO 2nd, Padon BW, Parikh UM, Colchado D, Vangapandu H, Shams S, Chi T, Jung JP, Balaji S. Short WD, et al. Front Bioeng Biotechnol. 2022 Sep 23;10:952198. doi: 10.3389/fbioe.2022.952198. eCollection 2022. Front Bioeng Biotechnol. 2022. PMID: 36213059 Free PMC article. Review.
Enhancing reinforcement learning models by including direct and indirect pathways improves performance on striatal dependent tasks.
Blackwell KT, Doya K. Blackwell KT, et al. PLoS Comput Biol. 2023 Aug 18;19(8):e1011385. doi: 10.1371/journal.pcbi.1011385. eCollection 2023 Aug. PLoS Comput Biol. 2023. PMID: 37594982 Free PMC article.
Application of an adapted FMEA framework for robot-inclusivity of built environments.
Ng YJ, Yeo MSK, Ng QB, Budig M, Muthugala MAVJ, Samarakoon SMBP, Mohan RE. Ng YJ, et al. Sci Rep. 2022 Mar 1;12(1):3408. doi: 10.1038/s41598-022-06902-4. Sci Rep. 2022. PMID: 35233018 Free PMC article.
Having multiple selves helps learning agents explore and adapt in complex changing worlds.
Dulberg Z, Dubey R, Berwian IM, Cohen JD. Dulberg Z, et al. Proc Natl Acad Sci U S A. 2023 Jul 11;120(28):e2221180120. doi: 10.1073/pnas.2221180120. Epub 2023 Jul 3. Proc Natl Acad Sci U S A. 2023. PMID: 37399387 Free PMC article.
Mobile Robot Application with Hierarchical Start Position DQN.
Erkan E, Arserim MA. Erkan E, et al. Comput Intell Neurosci. 2022 Sep 5;2022:4115767. doi: 10.1155/2022/4115767. eCollection 2022. Comput Intell Neurosci. 2022. PMID: 36105641 Free PMC article.

MeSH terms

Actions
Actions
Actions
Actions
Actions
Actions
Actions
Actions
Actions

LinkOut - more resources

Full Text Sources
- Elsevier Science
Other Literature Sources
- scite Smart Citations
Research Materials
- NCI CPTC Antibody Characterization Program
Miscellaneous
- NCI CPTAC Assay Portal

Save citation to file

Email citation

Add to Collections

Add to My Bibliography

Your saved search

Create a file for external citation management software

Your RSS Feed

Modular deep reinforcement learning from reward and punishment for robot navigation

Affiliations

Modular deep reinforcement learning from reward and punishment for robot navigation

Authors

Affiliations

Abstract

Conflict of interest statement

Similar articles

Cited by

MeSH terms

LinkOut - more resources

Full Text Sources

Other Literature Sources

Research Materials

Miscellaneous