Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling

In this work, we introduce Janus-Pro, an advanced version of the previous work Janus. Specifically, Janus-Pro incorporates (1) an optimized training strategy, (2) expanded training data, and (3) scaling to larger model size. With these improvements, Janus-Pro achieves significant advancements in both multimodal understanding and text-to-image instruction-following capabilities, while also enhancing the stability of text-to-image generation. We hope this work will inspire further exploration in the field. Code and models are publicly available.

SEED-X: MultimodalModels with Unified…SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and GenerationEmu3: Next-TokenPrediction is All You…Emu3: Next-Token Prediction is All You NeedELLA: Equip DiffusionModels with LLM for…ELLA: Equip Diffusion Models with LLM for Enhanced Semantic AlignmentChameleon: Mixed-ModalEarly-Fusion Foundation…Chameleon: Mixed-Modal Early-Fusion Foundation ModelsAutoregressive ModelBeats Diffusion: Llama…Autoregressive Model Beats Diffusion: Llama for Scalable Image GenerationJanus: Decoupling VisualEncoding for Unified…Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and GenerationShow-o: One SingleTransformer to Unify…Show-o: One Single Transformer to Unify Multimodal Understanding and GenerationTokenFlow: Unified ImageTokenizer for Multimoda…TokenFlow: Unified Image Tokenizer for Multimodal Understanding and GenerationVILA-U: a UnifiedFoundation Model…VILA-U: a Unified Foundation Model Integrating Visual Understanding and GenerationTransfusion: Predict theNext Token and Diffuse…Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal ModelMetaMorph: MultimodalUnderstanding and…MetaMorph: Multimodal Understanding and Generation via Instruction TuningILLUME: IlluminatingYour LLMs to See, Draw…ILLUME: Illuminating Your LLMs to See, Draw, and Self-EnhanceUniWorld-V1:High-Resolution Semanti…UniWorld-V1: High-Resolution Semantic Encoders for Unified Visual Understanding and GenerationSemHiTok: A UnifiedImage Tokenizer via…SemHiTok: A Unified Image Tokenizer via Semantic-Guided Hierarchical Codebook for Multimodal Understanding and GenerationMuddit: LiberatingGeneration Beyond…Muddit: Liberating Generation Beyond Text-to-Image with a Unified Discrete Diffusion ModelInterleaving Reasoningfor Better Text-to-Imag…Interleaving Reasoning for Better Text-to-Image GenerationQwen-Image TechnicalReportQwen-Image Technical ReportMing-Lite-Uni:Advancements in Unified…Ming-Lite-Uni: Advancements in Unified Architecture for Natural Multimodal InteractionUniFork: ExploringModality Alignment for…UniFork: Exploring Modality Alignment for Unified Multimodal Understanding and GenerationAn Empirical Study ofGPT-4o Image Generation…An Empirical Study of GPT-4o Image Generation CapabilitiesUni-X: MitigatingModality Conflict with…Uni-X: Mitigating Modality Conflict with a Two-End-Separated Architecture for Unified Multimodal ModelsUniRL-Zero:Reinforcement Learning…UniRL-Zero: Reinforcement Learning on Unified Models with Joint Language Model and Diffusion Model ExpertsReconstruction AlignmentImproves Unified…Reconstruction Alignment Improves Unified Multimodal ModelsComfyMind: TowardGeneral-Purpose…ComfyMind: Toward General-Purpose Generation via Tree-Based Planning and Reactive FeedbackJanus-Pro: UnifiedMultimodal Understandin…Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling過去の参考文献中心の論文この論文を引用する論文古い新しい

ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。