Qwen-Image Technical Report

We present Qwen-Image, an image generation foundation model in the Qwen series that achieves significant advances in complex text rendering and precise image editing. To address the challenges of complex text rendering, we design a comprehensive data pipeline that includes large-scale data collection, filtering, annotation, synthesis, and balancing. Moreover, we adopt a progressive training strategy that starts with non-text-to-text rendering, evolves from simple to complex textual inputs, and gradually scales up to paragraph-level descriptions. This curriculum learning approach substantially enhances the model's native text rendering capabilities. As a result, Qwen-Image not only performs exceptionally well in alphabetic languages such as English, but also achieves remarkable progress on more challenging logographic languages like Chinese. To enhance image editing consistency, we introduce an improved multi-task training paradigm that incorporates not only traditional text-to-image (T2I) and text-image-to-image (TI2I) tasks but also image-to-image (I2I) reconstruction, effectively aligning the latent representations between Qwen2.5-VL and MMDiT. Furthermore, we separately feed the original image into Qwen2.5-VL and the VAE encoder to obtain semantic and reconstructive representations, respectively. This dual-encoding mechanism enables the editing module to strike a balance between preserving semantic consistency and maintaining visual fidelity. Qwen-Image achieves state-of-the-art performance, demonstrating its strong capabilities in both image generation and editing across multiple benchmarks.

ELLA: Equip DiffusionModels with LLM for…ELLA: Equip Diffusion Models with LLM for Enhanced Semantic AlignmentStep1X-Edit: A PracticalFramework for General…Step1X-Edit: A Practical Framework for General Image EditingUniWorld-V1:High-Resolution Semanti…UniWorld-V1: High-Resolution Semantic Encoders for Unified Visual Understanding and GenerationFLUX.1 Kontext: FlowMatching for In-Context…FLUX.1 Kontext: Flow Matching for In-Context Image Generation and Editing in Latent SpaceOmniGen2: Exploration toAdvanced Multimodal…OmniGen2: Exploration to Advanced Multimodal GenerationBLIP3-o: A Family ofFully Open Unified…BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and DatasetImgEdit: A Unified ImageEditing Dataset and…ImgEdit: A Unified Image Editing Dataset and BenchmarkIn-Context Edit:Enabling Instructional…In-Context Edit: Enabling Instructional Image Editing with In-Context Generation in Large Scale Diffusion TransformerEmerging Properties inUnified Multimodal…Emerging Properties in Unified Multimodal PretrainingFlow-GRPO: Training FlowMatching Models via…Flow-GRPO: Training Flow Matching Models via Online RLShow-o: One SingleTransformer to Unify…Show-o: One Single Transformer to Unify Multimodal Understanding and GenerationJanus-Pro: UnifiedMultimodal Understandin…Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model ScalingLongCat-Image TechnicalReportLongCat-Image Technical ReportUniGenBench++: A UnifiedSemantic Evaluation…UniGenBench++: A Unified Semantic Evaluation Benchmark for Text-to-Image GenerationReconstruction AlignmentImproves Unified…Reconstruction Alignment Improves Unified Multimodal ModelsEasier Painting ThanThinking: Can…Easier Painting Than Thinking: Can Text-to-Image Models Set the Stage, but Not Direct the Play?EMMA: EfficientMultimodal…EMMA: Efficient Multimodal Understanding, Generation, and Editing with a Unified ArchitectureEditMGT: UnleashingPotentials of Masked…EditMGT: Unleashing Potentials of Masked Generative Transformers in Image EditingVisionDirector:Vision-Language Guided…VisionDirector: Vision-Language Guided Closed-Loop Refinement for Generative Image SynthesisPico-Banana-400K: ALarge-Scale Dataset for…Pico-Banana-400K: A Large-Scale Dataset for Text-Guided Image EditingWiseEdit: BenchmarkingCognition- and…WiseEdit: Benchmarking Cognition- and Creativity-Informed Image EditingDreamLite: A LightweightOn-Device Unified Model…DreamLite: A Lightweight On-Device Unified Model for Image Generation and EditingPromptRL: Prompt Mattersin RL for Flow-Based…PromptRL: Prompt Matters in RL for Flow-Based Image GenerationSTARFlow2: BridgingLanguage Models and…STARFlow2: Bridging Language Models and Normalizing Flows for Unified Multimodal GenerationQwen-Image TechnicalReportQwen-Image Technical Report過去の参考文献中心の論文この論文を引用する論文古い新しい

ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。