Principled Methods for Advising Reinforcement Learning Agents
An important issue in reinforcement learning is how to incorporate expert knowledge in a principled manner, especially as we scale up to real-world tasks. In this paper, we present a method for incorporating arbitrary advice into the reward structure of a reinforcement learning agent without altering the optimal policy. This method extends the potential-based shaping method proposed by Ng et al. (1999) to the case of shaping functions based on both states and actions. This allows for much more specic information to guide the agent { which action to choose { without re-quiring the agent to discover this from the re-wards on states alone. We develop two qual-itatively dierent methods for converting a potential function into advice for the agent. We also provide theoretical and experimen-tal justications for choosing between these advice-giving algorithms based on the prop-erties of the potential function. 1.
