Qwen3-VL Technical Report

We introduce Qwen3-VL, the most capable vision-language model in the Qwen series to date, achieving superior performance across a broad range of multimodal benchmarks. It natively supports interleaved contexts of up to 256K tokens, seamlessly integrating text, images, and video. The model family includes both dense (2B/4B/8B/32B) and mixture-of-experts (30B-A3B/235B-A22B) variants to accommodate diverse latency-quality trade-offs. Qwen3-VL delivers three core pillars: (i) markedly stronger pure-text understanding, surpassing comparable text-only backbones in several cases; (ii) robust long-context comprehension with a native 256K-token window for both text and interleaved multimodal inputs, enabling faithful retention, retrieval, and cross-referencing across long documents and videos; and (iii) advanced multimodal reasoning across single-image, multi-image, and video tasks, demonstrating leading performance on comprehensive evaluations such as MMMU and visual-math benchmarks (e.g., MathVista and MathVision). Architecturally, we introduce three key upgrades: (i) an enhanced interleaved-MRoPE for stronger spatial-temporal modeling across images and video; (ii) DeepStack integration, which effectively leverages multi-level ViT features to tighten vision-language alignment; and (iii) text-based time alignment for video, evolving from T-RoPE to explicit textual timestamp alignment for more precise temporal grounding. Under comparable token budgets and latency constraints, Qwen3-VL achieves superior performance in both dense and Mixture-of-Experts (MoE) architectures. We envision Qwen3-VL serving as a foundational engine for image-grounded reasoning, agentic decision-making, and multimodal code intelligence in real-world workflows.

Qwen2-VL: EnhancingVision-Language Model's…Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any ResolutionAre We on the Right Wayfor Evaluating Large…Are We on the Right Way for Evaluating Large Vision-Language Models?MMBench: Is YourMulti-modal Model an…MMBench: Is Your Multi-modal Model an All-Around Player?Gemini 2.5: Pushing theFrontier with Advanced…Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic CapabilitiesDeepEyes: Incentivizing"Thinking with Images"…DeepEyes: Incentivizing "Thinking with Images" via Reinforcement LearningOCRBench v2: An ImprovedBenchmark for Evaluatin…OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and ReasoningMobile-Agent-v3:Fundamental Agents for…Mobile-Agent-v3: Fundamental Agents for GUI AutomationZeroBench: An ImpossibleVisual Benchmark for…ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal ModelsOpenCUA: OpenFoundations for…OpenCUA: Open Foundations for Computer-Use AgentsChartMimic: EvaluatingLMM's Cross-Modal…ChartMimic: Evaluating LMM's Cross-Modal Reasoning Capability via Chart-to-Code GenerationVideo-MMMU: EvaluatingKnowledge Acquisition…Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional VideosSigLIP 2: MultilingualVision-Language Encoder…SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense FeaturesVision-DeepResearchBenchmark: Rethinking…Vision-DeepResearch Benchmark: Rethinking Visual and Textual Search for Multimodal Large Language ModelsAgentOCR: ReimaginingAgent History via…AgentOCR: Reimagining Agent History via Optical Self-CompressionVideoAuto-R1: Video AutoReasoning via Thinking…VideoAuto-R1: Video Auto Reasoning via Thinking Once, Answering TwiceCo-Evolving PolicyDistillationCo-Evolving Policy DistillationP1-VL: Bridging VisualPerception and…P1-VL: Bridging Visual Perception and Scientific Reasoning in Physics OlympiadsGUIGuard: Toward aGeneral Framework for…GUIGuard: Toward a General Framework for Privacy-Preserving GUI AgentsRubiCap: Rubric-GuidedReinforcement Learning…RubiCap: Rubric-Guided Reinforcement Learning for Dense Image CaptioningSelf-Prophetic Decodingto Unlock Visual Search…Self-Prophetic Decoding to Unlock Visual Search in LVLMsReading ≠ Seeing:Diagnosing and Closing…Reading ≠ Seeing: Diagnosing and Closing the Typography Gap in Vision-Language ModelsETCHR: Editing ToClarify and Harness…ETCHR: Editing To Clarify and Harness ReasoningHiSpatial: TamingHierarchical 3D Spatial…HiSpatial: Taming Hierarchical 3D Spatial Understanding in Vision-Language ModelsBenchmarking andImproving GUI Agents in…Benchmarking and Improving GUI Agents in High-Dynamic EnvironmentsQwen3-VL TechnicalReportQwen3-VL Technical Report過去の参考文献中心の論文この論文を引用する論文古い新しい

ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。