Invariant Policy Optimization: Towards Stronger Generalization in Reinforcement Learning

June 01, 2020 ยท Declared Dead ยท ๐Ÿ› Conference on Learning for Dynamics & Control

๐Ÿ‘ป CAUSE OF DEATH: Ghosted
No code link whatsoever

"No code URL or promise found in abstract"

Evidence collected by the PWNC Scanner

Authors Anoopkumar Sonar, Vincent Pacelli, Anirudha Majumdar arXiv ID 2006.01096 Category cs.LG: Machine Learning Cross-listed cs.AI, cs.RO, stat.ML Citations 60 Venue Conference on Learning for Dynamics & Control Last Checked 5 months ago
Abstract
A fundamental challenge in reinforcement learning is to learn policies that generalize beyond the operating domains experienced during training. In this paper, we approach this challenge through the following invariance principle: an agent must find a representation such that there exists an action-predictor built on top of this representation that is simultaneously optimal across all training domains. Intuitively, the resulting invariant policy enhances generalization by finding causes of successful actions. We propose a novel learning algorithm, Invariant Policy Optimization (IPO), that implements this principle and learns an invariant policy during training. We compare our approach with standard policy gradient methods and demonstrate significant improvements in generalization performance on unseen domains for linear quadratic regulator and grid-world problems, and an example where a robot must learn to open doors with varying physical properties.
Community shame:
Not yet rated
Community Contributions

Found the code? Know the venue? Think something is wrong? Let us know!

๐Ÿ“œ Similar Papers

In the same crypt โ€” Machine Learning

Died the same way โ€” ๐Ÿ‘ป Ghosted