Mastering Atari with Discrete World Models

Intelligent agents need to generalize from past experience to achieve goals in complex environments. World models facilitate such generalization and allow learning behaviors from imagined outcomes to increase sample-efficiency. While learning world models from image inputs has recently become feasible for some tasks, modeling Atari games accurately enough to derive successful behaviors has remained an open challenge for many years. We introduce DreamerV2, a reinforcement learning agent that learns behaviors purely from predictions in the compact latent space of a powerful world model. The world model uses discrete representations and is trained separately from the policy. DreamerV2 constitutes the first agent that achieves human-level performance on the Atari benchmark of 55 tasks by learning behaviors inside a separately trained world model. With the same computational budget and wall-clock time, Dreamer V2 reaches 200M frames and surpasses the final performance of the top single-GPU agents IQN and Rainbow. DreamerV2 is also applicable to tasks with continuous actions, where it learns an accurate world model of a complex humanoid robot and solves stand-up and walking from only pixel inputs.

Learning and QueryingFast Generative Models…Learning and Querying Fast Generative Models for Reinforcement LearningModel-EnsembleTrust-Region Policy…Model-Ensemble Trust-Region Policy OptimizationLearning Latent Dynamicsfor Planning from PixelsLearning Latent Dynamics for Planning from PixelsIs Deep ReinforcementLearning Really…Is Deep Reinforcement Learning Really Superhuman on Atari?Benchmarking Model-BasedReinforcement LearningBenchmarking Model-Based Reinforcement LearningDream to Control:Learning Behaviors by…Dream to Control: Learning Behaviors by Latent ImaginationModel BasedReinforcement Learning…Model Based Reinforcement Learning for AtariCURL: ContrastiveUnsupervised…CURL: Contrastive Unsupervised Representations for Reinforcement LearningPlanning to Explore viaSelf-Supervised World…Planning to Explore via Self-Supervised World ModelsMastering Atari, Go,Chess and Shogi by…Mastering Atari, Go, Chess and Shogi by Planning with a Learned ModelStochastic LatentActor-Critic: Deep…Stochastic Latent Actor-Critic: Deep Reinforcement Learning with a Latent Variable ModelImage Augmentation IsAll You Need…Image Augmentation Is All You Need: Regularizing Deep Reinforcement Learning from PixelsMastering Atari Gameswith Limited DataMastering Atari Games with Limited DataDecision Transformer:Reinforcement Learning…Decision Transformer: Reinforcement Learning via Sequence ModelingImaginary HindsightExperience Replay…Imaginary Hindsight Experience Replay: Curious Model-based Learning for Sparse Reward TasksFractional TransferLearning for Deep…Fractional Transfer Learning for Deep Model-Based Reinforcement LearningRepresentation Matters:Offline Pretraining for…Representation Matters: Offline Pretraining for Sequential Decision MakingCuriosity-DrivenExploration via Latent…Curiosity-Driven Exploration via Latent Bayesian SurpriseRepresentation Learningfor Continuous Action…Representation Learning for Continuous Action Spaces is Beneficial for Efficient Policy LearningStructured World Modelsfrom Human VideosStructured World Models from Human VideosMoDem: AcceleratingVisual Model-Based…MoDem: Accelerating Visual Model-Based Reinforcement Learning with DemonstrationsDream to Generalize:Zero-Shot Model-Based…Dream to Generalize: Zero-Shot Model-Based Reinforcement Learning for Unseen Visual DistractionsLearning InteractiveReal-World SimulatorsLearning Interactive Real-World SimulatorsDINO-WM: World Models onPre-trained Visual…DINO-WM: World Models on Pre-trained Visual Features enable Zero-shot PlanningMastering Atari withDiscrete World ModelsMastering Atari with Discrete World Models過去の参考文献中心の論文この論文を引用する論文古い新しい

ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。