Wan: Open and Advanced Large-Scale Video Generative Models

This report presents Wan, a comprehensive and open suite of video foundation models designed to push the boundaries of video generation. Built upon the mainstream diffusion transformer paradigm, Wan achieves significant advancements in generative capabilities through a series of innovations, including our novel VAE, scalable pre-training strategies, large-scale data curation, and automated evaluation metrics. These contributions collectively enhance the model's performance and versatility. Specifically, Wan is characterized by four key features: Leading Performance: The 14B model of Wan, trained on a vast dataset comprising billions of images and videos, demonstrates the scaling laws of video generation with respect to both data and model size. It consistently outperforms the existing open-source models as well as state-of-the-art commercial solutions across multiple internal and external benchmarks, demonstrating a clear and significant performance superiority. Comprehensiveness: Wan offers two capable models, i.e., 1.3B and 14B parameters, for efficiency and effectiveness respectively. It also covers multiple downstream applications, including image-to-video, instruction-guided video editing, and personal video generation, encompassing up to eight tasks. Consumer-Grade Efficiency: The 1.3B model demonstrates exceptional resource efficiency, requiring only 8.19 GB VRAM, making it compatible with a wide range of consumer-grade GPUs. Openness: We open-source the entire series of Wan, including source code and all models, with the goal of fostering the growth of the video generation community. This openness seeks to significantly expand the creative possibilities of video production in the industry and provide academia with high-quality video foundation models. All the code and models are available at https://github.com/Wan-Video/Wan2.1.

Stable Video Diffusion:Scaling Latent Video…Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large DatasetsFlow Matching forGenerative ModelingFlow Matching for Generative ModelingModelScope Text-to-VideoTechnical ReportModelScope Text-to-Video Technical ReportI2VGen-XL: High-QualityImage-to-Video Synthesi…I2VGen-XL: High-Quality Image-to-Video Synthesis via Cascaded Diffusion ModelsVideoLCM: Video LatentConsistency ModelVideoLCM: Video Latent Consistency ModelHunyuanVideo: ASystematic Framework Fo…HunyuanVideo: A Systematic Framework For Large Video Generative ModelsLTX-Video: RealtimeVideo Latent DiffusionLTX-Video: Realtime Video Latent DiffusionDreamVideo-2: Zero-ShotSubject-Driven Video…DreamVideo-2: Zero-Shot Subject-Driven Video Customization with Precise Motion ControlVACE: All-in-One VideoCreation and EditingVACE: All-in-One Video Creation and EditingDreamRelation:Relation-Centric Video…DreamRelation: Relation-Centric Video CustomizationPyramidal Flow Matchingfor Efficient Video…Pyramidal Flow Matching for Efficient Video Generative ModelingTimestep EmbeddingTells: It's Time to…Timestep Embedding Tells: It's Time to Cache for Video Diffusion ModelAutoregressiveAdversarial…Autoregressive Adversarial Post-Training for Real-Time Interactive Video GenerationWan-Move:Motion-controllable…Wan-Move: Motion-controllable Video Generation via Latent Trajectory GuidanceRELIC: Interactive VideoWorld Model with…RELIC: Interactive Video World Model with Long-Horizon MemoryDiCache: Let DiffusionModel Determine Its Own…DiCache: Let Diffusion Model Determine Its Own CacheVideoScore2: Thinkbefore You Score in…VideoScore2: Think before You Score in Generative Video EvaluationA Survey of InteractiveGenerative VideoA Survey of Interactive Generative VideoHunyuan-GameCraft:High-dynamic Interactiv…Hunyuan-GameCraft: High-dynamic Interactive Game Video Generation with Hybrid History ConditionAquarius: A Family ofIndustry-Level Video…Aquarius: A Family of Industry-Level Video Generation Models for Marketing ScenariosLoRA-Edit: ControllableFirst-Frame-Guided Vide…LoRA-Edit: Controllable First-Frame-Guided Video Editing via Mask-Aware LoRA Fine-TuningVChain:Chain-of-Visual-Thought…VChain: Chain-of-Visual-Thought for Reasoning in Video GenerationVerseCrafter: DynamicRealistic Video World…VerseCrafter: Dynamic Realistic Video World Model with 4D Geometric ControlBridging Brain andSemantics: A…Bridging Brain and Semantics: A Hierarchical Framework for Semantically Enhanced fMRI-to-Video ReconstructionWan: Open and AdvancedLarge-Scale Video…Wan: Open and Advanced Large-Scale Video Generative ModelsEarlier referencesFocus paperCiting papersOlderNewer

Click a node to pin it, click the empty canvas to go back to this paper, or hover to preview. Open a node’s page from its title.