DreamGen: Unlocking Generalization in Robot Learning through Neural Trajectories

We introduce DreamGen, a simple yet highly effective 4-stage pipeline for training robot policies that generalize across behaviors and environments through neural trajectories - synthetic robot data generated from video world models. DreamGen leverages state-of-the-art image-to-video generative models, adapting them to the target robot embodiment to produce photorealistic synthetic videos of familiar or novel tasks in diverse environments. Since these models generate only videos, we recover pseudo-action sequences using either a latent action model or an inverse-dynamics model (IDM). Despite its simplicity, DreamGen unlocks strong behavior and environment generalization: a humanoid robot can perform 22 new behaviors in both seen and unseen environments, while requiring teleoperation data from only a single pick-and-place task in one environment. To evaluate the pipeline systematically, we introduce DreamGen Bench, a video generation benchmark that shows a strong correlation between benchmark performance and downstream policy success. Our work establishes a promising new axis for scaling robot learning well beyond manual data collection. Code available at https://github.com/NVIDIA/GR00T-Dreams.

Gen2Act: Human VideoGeneration in Novel…Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot ManipulationRoboDreamer: LearningCompositional World…RoboDreamer: Learning Compositional World Models for Robot ImaginationAny-point TrajectoryModeling for Policy…Any-point Trajectory Modeling for Policy LearningGR-2: A GenerativeVideo-Language-Action…GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot ManipulationGR00T N1: An OpenFoundation Model for…GR00T N1: An Open Foundation Model for Generalist Humanoid RobotsUnified World Models:Coupling Video and…Unified World Models: Coupling Video and Action Diffusion for Pretraining on Large Robotic DatasetsAdaWorld: LearningAdaptable World Models…AdaWorld: Learning Adaptable World Models with Latent ActionsUnified Video ActionModelUnified Video Action ModelUniVLA: Learning to ActAnywhere with…UniVLA: Learning to Act Anywhere with Task-centric Latent ActionsAgiBot World Colosseo: ALarge-scale Manipulatio…AgiBot World Colosseo: A Large-scale Manipulation Platform for Scalable and Intelligent Embodied SystemsMoto: Latent MotionToken as the Bridging…Moto: Latent Motion Token as the Bridging Language for Learning Robot Manipulation from VideosCosmos-Transfer1:Conditional World…Cosmos-Transfer1: Conditional World Generation with Adaptive Multimodal ControlGenie Envisioner: AUnified World Foundatio…Genie Envisioner: A Unified World Foundation Platform for Robotic ManipulationGigaWorld-0: WorldModels as Data Engine t…GigaWorld-0: World Models as Data Engine to Empower Embodied AIGigaBrain-0: A WorldModel-Powered…GigaBrain-0: A World Model-Powered Vision-Language-Action ModelWorld Simulation withVideo Foundation Models…World Simulation with Video Foundation Models for Physical AIRynnVLA-002: A UnifiedVision-Language-Action…RynnVLA-002: A Unified Vision-Language-Action and World ModelA Step Toward WorldModels: A Survey on…A Step Toward World Models: A Survey on Robotic ManipulationFast-WAM: Do WorldAction Models Need…Fast-WAM: Do World Action Models Need Test-time Future Imagination?Rethinking VideoGeneration Model for th…Rethinking Video Generation Model for the Embodied WorldCosmos Policy:Fine-Tuning Video Model…Cosmos Policy: Fine-Tuning Video Models for Visuomotor Control and PlanningGigaBrain-0.5M*: a VLAThat Learns From World…GigaBrain-0.5M*: a VLA That Learns From World Model-Based Reinforcement LearningVAG: Dual-StreamVideo-Action Generation…VAG: Dual-Stream Video-Action Generation for Embodied Data SynthesisWorldArena: A UnifiedBenchmark for Evaluatin…WorldArena: A Unified Benchmark for Evaluating Perception and Functional Utility of Embodied World ModelsDreamGen: UnlockingGeneralization in Robot…DreamGen: Unlocking Generalization in Robot Learning through Neural Trajectories過去の参考文献中心の論文この論文を引用する論文古い新しい

ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。