InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency

We introduce InternVL 3.5, a new family of open-source multimodal models that significantly advances versatility, reasoning capability, and inference efficiency along the InternVL series. A key innovation is the Cascade Reinforcement Learning (Cascade RL) framework, which enhances reasoning through a two-stage process: offline RL for stable convergence and online RL for refined alignment. This coarse-to-fine training strategy leads to substantial improvements on downstream reasoning tasks, e.g., MMMU and MathVista. To optimize efficiency, we propose a Visual Resolution Router (ViR) that dynamically adjusts the resolution of visual tokens without compromising performance. Coupled with ViR, our Decoupled Vision-Language Deployment (DvD) strategy separates the vision encoder and language model across different GPUs, effectively balancing computational load. These contributions collectively enable InternVL3.5 to achieve up to a +16.0\% gain in overall reasoning performance and a 4.05$\times$ inference speedup compared to its predecessor, i.e., InternVL3. In addition, InternVL3.5 supports novel capabilities such as GUI interaction and embodied agency. Notably, our largest model, i.e., InternVL3.5-241B-A28B, attains state-of-the-art results among open-source MLLMs across general multimodal, reasoning, text, and agentic tasks -- narrowing the performance gap with leading commercial models like GPT-5. All models and code are publicly released.

Expanding PerformanceBoundaries of…Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time ScalingEnhancing the ReasoningAbility of Multimodal…Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference OptimizationQwen2-VL: EnhancingVision-Language Model's…Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any ResolutionHow Far Are We toGPT-4V? Closing the Gap…How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source SuitesInternVL3: ExploringAdvanced Training and…InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal ModelsVisualPRM: An EffectiveProcess Reward Model fo…VisualPRM: An Effective Process Reward Model for Multimodal ReasoningKimi-VL Technical ReportKimi-VL Technical ReportSeed1.5-VL TechnicalReportSeed1.5-VL Technical ReportMolmo and PixMo: OpenWeights and Open Data…Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language ModelsDynaMath: A DynamicVisual Benchmark for…DynaMath: A Dynamic Visual Benchmark for Evaluating Mathematical Reasoning Robustness of Vision Language ModelsBMMR: A Large-ScaleBilingual Multimodal…BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning DatasetMM-Eureka: ExploringVisual Aha Moment with…MM-Eureka: Exploring Visual Aha Moment with Rule-based Large-scale Reinforcement LearningNVIDIA Nemotron Nano V2VLNVIDIA Nemotron Nano V2 VLBee: A High-QualityCorpus and Full-Stack…Bee: A High-Quality Corpus and Full-Stack Suite to Unlock Advanced Fully Open MLLMsMindGPT-4ov: An EnhancedMLLM via a Multi-Stage…MindGPT-4ov: An Enhanced MLLM via a Multi-Stage Post-Training ParadigmMixing Importance withDiversity: Joint…Mixing Importance with Diversity: Joint Optimization for KV Cache Compression in Large Vision-Language ModelsVKnowU: EvaluatingVisual Knowledge…VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMsLaV-CoT: Language-AwareVisual CoT with…LaV-CoT: Language-Aware Visual CoT with Multi-Aspect Reward Optimization for Real-World Multilingual VQASpinBench: Perspectiveand Rotation as a Lens…SpinBench: Perspective and Rotation as a Lens on Spatial Reasoning in VLMsEfficient Inference forLarge Vision-Language…Efficient Inference for Large Vision-Language Models: Bottlenecks, Techniques, and ProspectsXuanwu: Evolving GeneralMultimodal Models into…Xuanwu: Evolving General Multimodal Models into an Industrial-Grade Foundation for Content EcosystemsP1-VL: Bridging VisualPerception and…P1-VL: Bridging Visual Perception and Scientific Reasoning in Physics OlympiadsInstinct vs. Reflection:Unifying Token and…Instinct vs. Reflection: Unifying Token and Verbalized Confidence in Multimodal Large ModelsMinerU-Diffusion:Rethinking Document OCR…MinerU-Diffusion: Rethinking Document OCR as Inverse Rendering via Diffusion DecodingInternVL3.5: AdvancingOpen-Source Multimodal…InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and EfficiencyEarlier referencesFocus paperCiting papersOlderNewer

Click a node to pin it, click the empty canvas to go back to this paper, or hover to preview. Open a node’s page from its title.