Dream to Control: Learning Behaviors by Latent Imagination

Learned world models summarize an agent's experience to facilitate learning complex behaviors. While learning world models from high-dimensional sensory inputs is becoming feasible through deep learning, there are many potential ways for deriving behaviors from them. We present Dreamer, a reinforcement learning agent that solves long-horizon tasks from images purely by latent imagination. We efficiently learn behaviors by propagating analytic gradients of learned state values back through trajectories imagined in the compact state space of a learned world model. On 20 challenging visual control tasks, Dreamer exceeds existing approaches in data-efficiency, computation time, and final performance.

Learning and QueryingFast Generative Models…Learning and Querying Fast Generative Models for Reinforcement LearningDeepMind Control SuiteDeepMind Control SuiteSoft Actor-Critic:Off-Policy Maximum…Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic ActorModel-EnsembleTrust-Region Policy…Model-Ensemble Trust-Region Policy OptimizationLearning Latent Dynamicsfor Planning from PixelsLearning Latent Dynamics for Planning from PixelsDeepMDP: LearningContinuous Latent Space…DeepMDP: Learning Continuous Latent Space Models for Representation LearningImagined ValueGradients: Model-Based…Imagined Value Gradients: Model-Based Policy Optimization with Transferable Latent Dynamics ModelsShaping Belief Stateswith Generative…Shaping Belief States with Generative Environment Models for RLModel BasedReinforcement Learning…Model Based Reinforcement Learning for AtariStochastic LatentActor-Critic: Deep…Stochastic Latent Actor-Critic: Deep Reinforcement Learning with a Latent Variable ModelMastering Atari, Go,Chess and Shogi by…Mastering Atari, Go, Chess and Shogi by Planning with a Learned ModelExploring Model-basedPlanning with Policy…Exploring Model-based Planning with Policy NetworksCURL: ContrastiveUnsupervised…CURL: Contrastive Unsupervised Representations for Reinforcement LearningReinforcement Learningthrough Active InferenceReinforcement Learning through Active InferenceLocal Search for PolicyIteration in Continuous…Local Search for Policy Iteration in Continuous ControlMasked ContrastiveRepresentation Learning…Masked Contrastive Representation Learning for Reinforcement LearningAuxiliary-task BasedDeep Reinforcement…Auxiliary-task Based Deep Reinforcement Learning for Participant Selection Problem in Mobile CrowdsourcingMastering Atari Gameswith Limited DataMastering Atari Games with Limited DataModel-Based OfflinePlanningModel-Based Offline PlanningVector Quantized Modelsfor PlanningVector Quantized Models for PlanningLearning Vision-GuidedQuadrupedal Locomotion…Learning Vision-Guided Quadrupedal Locomotion End-to-End with Cross-Modal TransformersSAMBA: Safe Model-Based& Active Reinforcement…SAMBA: Safe Model-Based & Active Reinforcement LearningLearning explainabletask-relevant state…Learning explainable task-relevant state representation for model-free deep reinforcement learningWorld Action Models areZero-shot PoliciesWorld Action Models are Zero-shot PoliciesDream to Control:Learning Behaviors by…Dream to Control: Learning Behaviors by Latent ImaginationEarlier referencesFocus paperCiting papersOlderNewer

Click a node to pin it, click the empty canvas to go back to this paper, or hover to preview. Open a node’s page from its title.