Qwen2.5-VL Technical Report

We introduce Qwen2.5-VL, the latest flagship model of Qwen vision-language series, which demonstrates significant advancements in both foundational capabilities and innovative functionalities. Qwen2.5-VL achieves a major leap forward in understanding and interacting with the world through enhanced visual recognition, precise object localization, robust document parsing, and long-video comprehension. A standout feature of Qwen2.5-VL is its ability to localize objects using bounding boxes or points accurately. It provides robust structured data extraction from invoices, forms, and tables, as well as detailed analysis of charts, diagrams, and layouts. To handle complex inputs, Qwen2.5-VL introduces dynamic resolution processing and absolute time encoding, enabling it to process images of varying sizes and videos of extended durations (up to hours) with second-level event localization. This allows the model to natively perceive spatial scales and temporal dynamics without relying on traditional normalization techniques. By training a native dynamic-resolution Vision Transformer (ViT) from scratch and incorporating Window Attention, we reduce computational overhead while maintaining native resolution. As a result, Qwen2.5-VL excels not only in static image and document understanding but also as an interactive visual agent capable of reasoning, tool usage, and task execution in real-world scenarios such as operating computers and mobile devices. Qwen2.5-VL is available in three sizes, addressing diverse use cases from edge AI to high-performance computing. The flagship Qwen2.5-VL-72B model matches state-of-the-art models like GPT-4o and Claude 3.5 Sonnet, particularly excelling in document and diagram understanding. Additionally, Qwen2.5-VL maintains robust linguistic performance, preserving the core language competencies of the Qwen2.5 LLM.

ChartQA: A Benchmark forQuestion Answering abou…ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical ReasoningQwen2-VL: EnhancingVision-Language Model's…Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any ResolutionDeepSeek-VL2:Mixture-of-Experts…DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal UnderstandingExpanding PerformanceBoundaries of…Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time ScalingAre We on the Right Wayfor Evaluating Large…Are We on the Right Way for Evaluating Large Vision-Language Models?MMBench: Is YourMulti-modal Model an…MMBench: Is Your Multi-modal Model an All-Around Player?MMMU: A MassiveMulti-Discipline…MMMU: A Massive Multi-Discipline Multimodal Understanding and Reasoning Benchmark for Expert AGIQwen2.5 Technical ReportQwen2.5 Technical ReportLLaVA-OneVision: EasyVisual Task TransferLLaVA-OneVision: Easy Visual Task TransferOCRBench v2: An ImprovedBenchmark for Evaluatin…OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and ReasoningMME-RealWorld: CouldYour Multimodal LLM…MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans?Video-MMMU: EvaluatingKnowledge Acquisition…Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional VideosVideoChat-Flash:Hierarchical Compressio…VideoChat-Flash: Hierarchical Compression for Long-Context Video ModelingMR. Video: "MapReduce"is the Principle for…MR. Video: "MapReduce" is the Principle for Long Video UnderstandingMM-InstructEval:Zero-Shot Evaluation of…MM-InstructEval: Zero-Shot Evaluation of (Multimodal) Large Language Models on Multimodal Reasoning TasksLaV-CoT: Language-AwareVisual CoT with…LaV-CoT: Language-Aware Visual CoT with Multi-Aspect Reward Optimization for Real-World Multilingual VQASAILViT: Towards Robustand Generalizable Visua…SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature RefinementUncertainty-Aware GUIAgent: Adaptive…Uncertainty-Aware GUI Agent: Adaptive Perception through Component Recommendation and Human-in-the-Loop RefinementSelf-RewardingVision-Language Model…Self-Rewarding Vision-Language Model via Reasoning DecompositionChartPoint: GuidingMLLMs with Grounding…ChartPoint: Guiding MLLMs with Grounding Reflection for Chart ReasoningVision Remember:Alleviating Visual…Vision Remember: Alleviating Visual Forgetting in Efficient MLLM with Vision Feature ResampleSherlock:Self-Correcting…Sherlock: Self-Correcting Reasoning in Vision-Language ModelsReconstruction AlignmentImproves Unified…Reconstruction Alignment Improves Unified Multimodal ModelsSE-GA: Memory-AugmentedSelf-Evolution for GUI…SE-GA: Memory-Augmented Self-Evolution for GUI AgentsQwen2.5-VL TechnicalReportQwen2.5-VL Technical ReportEarlier referencesFocus paperCiting papersOlderNewer

Click a node to pin it, click the empty canvas to go back to this paper, or hover to preview. Open a node’s page from its title.