SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features

We introduce SigLIP 2, a family of new multilingual vision-language encoders that build on the success of the original SigLIP. In this second iteration, we extend the original image-text training objective with several prior, independently developed techniques into a unified recipe -- this includes captioning-based pretraining, self-supervised losses (self-distillation, masked prediction) and online data curation. With these changes, SigLIP 2 models outperform their SigLIP counterparts at all model scales in core capabilities, including zero-shot classification, image-text retrieval, and transfer performance when extracting visual representations for Vision-Language Models (VLMs). Furthermore, the new training recipe leads to significant improvements on localization and dense prediction tasks. We also train variants which support multiple resolutions and preserve the input's native aspect ratio. Finally, we train on a more diverse data-mixture that includes de-biasing techniques, leading to much better multilingual understanding and improved fairness. To allow users to trade off inference cost with performance, we release model checkpoints at four sizes: ViT-B (86M), L (303M), So400m (400M), and g (1B).

Fixing Weight DecayRegularization in AdamFixing Weight Decay Regularization in AdamQwen-VL: A FrontierLarge Vision-Language…Qwen-VL: A Frontier Large Vision-Language Model with Versatile AbilitiesEVA-CLIP: ImprovedTraining Techniques for…EVA-CLIP: Improved Training Techniques for CLIP at ScalePaLI: A Jointly-ScaledMultilingual…PaLI: A Jointly-Scaled Multilingual Language-Image ModelPaliGemma: A versatile3B VLM for transferPaliGemma: A versatile 3B VLM for transferPaliGemma 2: A Family ofVersatile VLMs for…PaliGemma 2: A Family of Versatile VLMs for TransferMM1: Methods, Analysis &Insights from Multimoda…MM1: Methods, Analysis & Insights from Multimodal LLM Pre-trainingCambrian-1: A FullyOpen, Vision-Centric…Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMsWayfinding Stages: TheRole of Familiarity…Wayfinding Stages: The Role of Familiarity, Gaze Events, and Visual AttentionGemma: Open Models Basedon Gemini Research and…Gemma: Open Models Based on Gemini Research and TechnologyGemma 2: Improving OpenLanguage Models at a…Gemma 2: Improving Open Language Models at a Practical SizeMultimodalAutoregressive…Multimodal Autoregressive Pre-training of Large Vision EncodersQwen-Image TechnicalReportQwen-Image Technical ReportUniWorld-V1:High-Resolution Semanti…UniWorld-V1: High-Resolution Semantic Encoders for Unified Visual Understanding and GenerationEMMA: EfficientMultimodal…EMMA: Efficient Multimodal Understanding, Generation, and Editing with a Unified ArchitectureUni-X: MitigatingModality Conflict with…Uni-X: Mitigating Modality Conflict with a Two-End-Separated Architecture for Unified Multimodal ModelsSAILViT: Towards Robustand Generalizable Visua…SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature RefinementFUSION: FullyIntegration of…FUSION: Fully Integration of Vision-Language Representations for Deep Cross-Modal UnderstandingBeyond LanguageModeling: An Exploratio…Beyond Language Modeling: An Exploration of Multimodal Pretraining20/20 Vision LanguageModels: A Prescription…20/20 Vision Language Models: A Prescription for Better VLMs through Data Curation AloneAligning Forest andTrees in Images and Lon…Aligning Forest and Trees in Images and Long Captions for Visually Grounded UnderstandingLookBench: A Live andHolistic Open Benchmark…LookBench: A Live and Holistic Open Benchmark for Fashion Image RetrievalInnovator-VL: AMultimodal Large…Innovator-VL: A Multimodal Large Language Model for Scientific DiscoverySCAM: A Real-WorldTypographic Robustness…SCAM: A Real-World Typographic Robustness Evaluation for Multimodal Foundation ModelsSigLIP 2: MultilingualVision-Language Encoder…SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense FeaturesEarlier referencesFocus paperCiting papersOlderNewer

Click a node to pin it, click the empty canvas to go back to this paper, or hover to preview. Open a node’s page from its title.