Genie Envisioner: A Unified World Foundation Platform for Robotic Manipulation

We introduce Genie Envisioner (GE), a unified world foundation platform for robotic manipulation that integrates policy learning, evaluation, and simulation within a single video-generative framework. At its core, GE-Base is a large-scale, instruction-conditioned video diffusion model that captures the spatial, temporal, and semantic dynamics of real-world robotic interactions in a structured latent space. Built upon this foundation, GE-Act maps latent representations to executable action trajectories through a lightweight, flow-matching decoder, enabling precise and generalizable policy inference across diverse embodiments with minimal supervision. To support scalable evaluation and training, GE-Sim serves as an action-conditioned neural simulator, producing high-fidelity rollouts for closed-loop policy development. The platform is further equipped with EWMBench, a standardized benchmark suite measuring visual fidelity, physical consistency, and instruction-action alignment. Together, these components establish Genie Envisioner as a scalable and practical foundation for instruction-driven, general-purpose embodied intelligence. All code, models, and benchmarks will be released publicly.

Visual Foresight:Model-Based Deep…Visual Foresight: Model-Based Deep Reinforcement Learning for Vision-Based Robotic Controlπ0: AVision-Language-Action…π0: A Vision-Language-Action Flow Model for General Robot ControlEnerVerse-AC:Envisioning Embodied…EnerVerse-AC: Envisioning Embodied Environments with Action ConditionEWMBench: EvaluatingScene, Motion, and…EWMBench: Evaluating Scene, Motion, and Semantic Quality in Embodied World ModelsEnerVerse: EnvisioningEmbodied Future Space…EnerVerse: Envisioning Embodied Future Space for Robotics ManipulationDreamGen: UnlockingGeneralization in Robot…DreamGen: Unlocking Generalization in Robot Learning through Neural TrajectoriesGR00T N1: An OpenFoundation Model for…GR00T N1: An Open Foundation Model for Generalist Humanoid RobotsRoboTwin 2.0: A ScalableData Generator and…RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic ManipulationUniVLA: Learning to ActAnywhere with…UniVLA: Learning to Act Anywhere with Task-centric Latent ActionsAgiBot World Colosseo: ALarge-scale Manipulatio…AgiBot World Colosseo: A Large-scale Manipulation Platform for Scalable and Intelligent Embodied SystemsCosmos World FoundationModel Platform for…Cosmos World Foundation Model Platform for Physical AITowards World Simulator:Crafting Physical…Towards World Simulator: Crafting Physical Commonsense-Based Benchmark for Video GenerationCtrl-World: AControllable Generative…Ctrl-World: A Controllable Generative World Model for Robot ManipulationWorld Simulation withVideo Foundation Models…World Simulation with Video Foundation Models for Physical AIGigaWorld-0: WorldModels as Data Engine t…GigaWorld-0: World Models as Data Engine to Empower Embodied AIInternVLA-M1: ASpatially Guided…InternVLA-M1: A Spatially Guided Vision-Language-Action Framework for Generalist Robot PolicyAct2Goal: From WorldModel To General…Act2Goal: From World Model To General Goal-conditioned PolicyFast-WAM: Do WorldAction Models Need…Fast-WAM: Do World Action Models Need Test-time Future Imagination?Cosmos Policy:Fine-Tuning Video Model…Cosmos Policy: Fine-Tuning Video Models for Visuomotor Control and PlanningInternVLA-A1: UnifyingUnderstanding…InternVLA-A1: Unifying Understanding, Generation and Action for Robotic ManipulationWorldArena: A UnifiedBenchmark for Evaluatin…WorldArena: A Unified Benchmark for Evaluating Perception and Functional Utility of Embodied World ModelsGigaBrain-0.5M*: a VLAThat Learns From World…GigaBrain-0.5M*: a VLA That Learns From World Model-Based Reinforcement LearningFrom Imagined Futures toExecutable Actions…From Imagined Futures to Executable Actions: Mixture of Latent Actions for Robot ManipulationWALL-WM: Carving WorldAction Modeling at the…WALL-WM: Carving World Action Modeling at the Event JointsGenie Envisioner: AUnified World Foundatio…Genie Envisioner: A Unified World Foundation Platform for Robotic Manipulation過去の参考文献中心の論文この論文を引用する論文古い新しい

ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。