Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

We present Stable Video Diffusion - a latent video diffusion model for high-resolution, state-of-the-art text-to-video and image-to-video generation. Recently, latent diffusion models trained for 2D image synthesis have been turned into generative video models by inserting temporal layers and finetuning them on small, high-quality video datasets. However, training methods in the literature vary widely, and the field has yet to agree on a unified strategy for curating video data. In this paper, we identify and evaluate three different stages for successful training of video LDMs: text-to-image pretraining, video pretraining, and high-quality video finetuning. Furthermore, we demonstrate the necessity of a well-curated pretraining dataset for generating high-quality videos and present a systematic curation process to train a strong base model, including captioning and filtering strategies. We then explore the impact of finetuning our base model on high-quality data and train a text-to-video model that is competitive with closed-source video generation. We also show that our base model provides a powerful motion representation for downstream tasks such as image-to-video generation and adaptability to camera motion-specific LoRA modules. Finally, we demonstrate that our model provides a strong multi-view 3D-prior and can serve as a base to finetune a multi-view diffusion model that jointly generates multiple views of objects in a feedforward fashion, outperforming image-based methods at a fraction of their compute budget. We release code and model weights at https://github.com/Stability-AI/generative-models .

Imagen Video: HighDefinition Video…Imagen Video: High Definition Video Generation with Diffusion ModelsMagicVideo: EfficientVideo Generation With…MagicVideo: Efficient Video Generation With Latent Diffusion ModelsClassifier-FreeDiffusion GuidanceClassifier-Free Diffusion GuidanceProgressive Distillationfor Fast Sampling of…Progressive Distillation for Fast Sampling of Diffusion ModelsHierarchicalText-Conditional Image…Hierarchical Text-Conditional Image Generation with CLIP LatentsModelScope Text-to-VideoTechnical ReportModelScope Text-to-Video Technical ReportAlign Your Latents:High-Resolution Video…Align Your Latents: High-Resolution Video Synthesis with Latent Diffusion ModelsMake-A-Video:Text-to-Video Generatio…Make-A-Video: Text-to-Video Generation without Text-Video DataI2VGen-XL: High-QualityImage-to-Video Synthesi…I2VGen-XL: High-Quality Image-to-Video Synthesis via Cascaded Diffusion ModelsAnimateDiff: AnimateYour Personalized…AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific TuningSDXL: Improving LatentDiffusion Models for…SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisLAVIE: High-QualityVideo Generation with…LAVIE: High-Quality Video Generation with Cascaded Latent Diffusion ModelsDimensionX: Create Any3D and 4D Scenes from a…DimensionX: Create Any 3D and 4D Scenes from a Single Image with Controllable Video DiffusionChronoMagic-Bench: ABenchmark for…ChronoMagic-Bench: A Benchmark for Metamorphic Evaluation of Text-to-Time-lapse Video GenerationStoryAgent: CustomizedStorytelling Video…StoryAgent: Customized Storytelling Video Generation via Multi-Agent CollaborationxGen-VideoSyn-1:High-Fidelity…xGen-VideoSyn-1: High-Fidelity Text-to-Video Synthesis with Compressed RepresentationsContextualStory:Consistent Visual…ContextualStory: Consistent Visual Storytelling with Spatially-Enhanced and Storyline ContextOpenVid-1M: ALarge-Scale High-Qualit…OpenVid-1M: A Large-Scale High-Quality Dataset for Text-to-video GenerationDiffusion Model-BasedImage Editing: A SurveyDiffusion Model-Based Image Editing: A SurveyYou See it, You Got it:Learning 3D Creation on…You See it, You Got it: Learning 3D Creation on Pose-Free Videos at ScaleText2PDE: LatentDiffusion Models for…Text2PDE: Latent Diffusion Models for Accessible Physics SimulationEnhancing Human-ComputerInteraction Through…Enhancing Human-Computer Interaction Through Decoupling Motion and Camera Control in Human-Centric Video GenerationWan-Move:Motion-controllable…Wan-Move: Motion-controllable Video Generation via Latent Trajectory GuidanceEnerVerse: EnvisioningEmbodied Future Space…EnerVerse: Envisioning Embodied Future Space for Robotics ManipulationStable Video Diffusion:Scaling Latent Video…Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large DatasetsEarlier referencesFocus paperCiting papersOlderNewer

Click a node to pin it, click the empty canvas to go back to this paper, or hover to preview. Open a node’s page from its title.