CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

We present CogVideoX, a large-scale text-to-video generation model based on diffusion transformer, which can generate 10-second continuous videos aligned with text prompt, with a frame rate of 16 fps and resolution of 768 * 1360 pixels. Previous video generation models often had limited movement and short durations, and is difficult to generate videos with coherent narratives based on text. We propose several designs to address these issues. First, we propose a 3D Variational Autoencoder (VAE) to compress videos along both spatial and temporal dimensions, to improve both compression rate and video fidelity. Second, to improve the text-video alignment, we propose an expert transformer with the expert adaptive LayerNorm to facilitate the deep fusion between the two modalities. Third, by employing a progressive training and multi-resolution frame pack technique, CogVideoX is adept at producing coherent, long-duration, different shape videos characterized by significant motions. In addition, we develop an effective text-video data processing pipeline that includes various data preprocessing strategies and a video captioning method, greatly contributing to the generation quality and semantic alignment. Results show that CogVideoX demonstrates state-of-the-art performance across both multiple machine metrics and human evaluations. The model weight of both 3D Causal VAE, Video caption model and CogVideoX are publicly available at https://github.com/THUDM/CogVideo.

Imagen Video: HighDefinition Video…Imagen Video: High Definition Video Generation with Diffusion ModelsProgressive Distillationfor Fast Sampling of…Progressive Distillation for Fast Sampling of Diffusion ModelsStable Video Diffusion:Scaling Latent Video…Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large DatasetsCogVideo: Large-scalePretraining for…CogVideo: Large-scale Pretraining for Text-to-Video Generation via TransformersMake-A-Video:Text-to-Video Generatio…Make-A-Video: Text-to-Video Generation without Text-Video DataGPT-4 Technical ReportGPT-4 Technical ReportAnimateDiff: AnimateYour Personalized…AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific TuningLanguage Model BeatsDiffusion - Tokenizer i…Language Model Beats Diffusion - Tokenizer is Key to Visual GenerationSDXL: Improving LatentDiffusion Models for…SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisLAVIE: High-QualityVideo Generation with…LAVIE: High-Quality Video Generation with Cascaded Latent Diffusion ModelsChronoMagic-Bench: ABenchmark for…ChronoMagic-Bench: A Benchmark for Metamorphic Evaluation of Text-to-Time-lapse Video GenerationShow-1: Marrying Pixeland Latent Diffusion…Show-1: Marrying Pixel and Latent Diffusion Models for Text-to-Video GenerationDimensionX: Create Any3D and 4D Scenes from a…DimensionX: Create Any 3D and 4D Scenes from a Single Image with Controllable Video DiffusionChronoMagic-Bench: ABenchmark for…ChronoMagic-Bench: A Benchmark for Metamorphic Evaluation of Text-to-Time-lapse Video GenerationOpenVid-1M: ALarge-Scale High-Qualit…OpenVid-1M: A Large-Scale High-Quality Dataset for Text-to-video GenerationMimir: Improving VideoDiffusion Models for…Mimir: Improving Video Diffusion Models for Precise Text UnderstandingIMAGEdit: Let AnySubject TransformIMAGEdit: Let Any Subject TransformT2VEval: Benchmarkdataset and objective…T2VEval: Benchmark dataset and objective evaluation method for T2V-generated videosPhantom:Subject-Consistent Vide…Phantom: Subject-Consistent Video Generation via Cross-Modal AlignmentDynamic ConceptsPersonalization from…Dynamic Concepts Personalization from Single VideosMitigating Surgical DataImbalance with…Mitigating Surgical Data Imbalance with Dual-Prediction Video Diffusion ModelDreamDance: AnimatingHuman Images by…DreamDance: Animating Human Images by Enriching 3D Geometry Cues from 2D PosesFlashVideo: FlowingFidelity to Detail for…FlashVideo: Flowing Fidelity to Detail for Efficient High-Resolution Video GenerationVChain:Chain-of-Visual-Thought…VChain: Chain-of-Visual-Thought for Reasoning in Video GenerationCogVideoX: Text-to-VideoDiffusion Models with A…CogVideoX: Text-to-Video Diffusion Models with An Expert TransformerEarlier referencesFocus paperCiting papersOlderNewer

Click a node to pin it, click the empty canvas to go back to this paper, or hover to preview. Open a node’s page from its title.