Sigmoid Loss for Language Image Pre-Training

We propose a simple pairwise Sigmoid loss for Language-Image Pre-training (SigLIP). Unlike standard contrastive learning with softmax normalization, the sigmoid loss operates solely on image-text pairs and does not require a global view of the pairwise similarities for normalization. The sigmoid loss simultaneously allows further scaling up the batch size, while also performing better at smaller batch sizes. Combined with Locked-image Tuning, with only four TPUv4 chips, we train a SigLiT model that achieves 84.5% ImageNet zero-shot accuracy in two days. The disentanglement of the batch size from the loss further allows us to study the impact of examples vs pairs and negative to positive ratio. Finally, we push the batch size to the extreme, up to one million, and find that the benefits of growing batch size quickly diminish, with a more reasonable batch size of 32k being sufficient. We release our models at https://github.com/google-research/big_vision and hope our research motivates further explorations in improving the quality and efficiency of language-image pre-training.

ImageNet: A large-scalehierarchical image…ImageNet: A large-scale hierarchical image databaseRepresentation Learningwith Contrastive…Representation Learning with Contrastive Predictive CodingFlorence: A NewFoundation Model for…Florence: A New Foundation Model for Computer VisionLiT: Zero-Shot Transferwith Locked-image Text…LiT: Zero-Shot Transfer with Locked-image Text TuningLAION-5B: An openlarge-scale dataset for…LAION-5B: An open large-scale dataset for training next generation image-text modelsCoCa: ContrastiveCaptioners are…CoCa: Contrastive Captioners are Image-Text Foundation ModelsSimple Open-VocabularyObject DetectionSimple Open-Vocabulary Object DetectionRobust fine-tuning ofzero-shot modelsRobust fine-tuning of zero-shot modelsPaLI: A Jointly-ScaledMultilingual…PaLI: A Jointly-Scaled Multilingual Language-Image ModelLLaMA: Open andEfficient Foundation…LLaMA: Open and Efficient Foundation Language ModelsEVA-CLIP: ImprovedTraining Techniques for…EVA-CLIP: Improved Training Techniques for CLIP at ScaleCombined Scaling forZero-shot Transfer…Combined Scaling for Zero-shot Transfer LearningImage Captioners AreScalable Vision Learner…Image Captioners Are Scalable Vision Learners TooNLLB-CLIP - trainperformant multilingual…NLLB-CLIP - train performant multilingual image retrieval model on a budgetInsect-Foundation: AFoundation Model and…Insect-Foundation: A Foundation Model and Large-Scale 1M Dataset for Visual Insect UnderstandingBad Students Make GreatTeachers: Active…Bad Students Make Great Teachers: Active Learning Accelerates Large-Scale Visual UnderstandingMultimodal FoundationModels: From Specialist…Multimodal Foundation Models: From Specialists to General-Purpose AssistantsV-IRL: Grounding VirtualIntelligence in Real…V-IRL: Grounding Virtual Intelligence in Real LifeCAT-Seg: CostAggregation for…CAT-Seg: Cost Aggregation for Open-Vocabulary Semantic SegmentationRegionGPT: TowardsRegion Understanding…RegionGPT: Towards Region Understanding Vision Language ModelSaviorRec:Semantic-Behavior…SaviorRec: Semantic-Behavior Alignment for Cold-Start RecommendationCobra: Extending Mambato Multi-Modal Large…Cobra: Extending Mamba to Multi-Modal Large Language Model for Efficient InferenceKimi K2.5: VisualAgentic IntelligenceKimi K2.5: Visual Agentic IntelligenceOpenDriveVLA: TowardsEnd-to-end Autonomous…OpenDriveVLA: Towards End-to-end Autonomous Driving with Large Vision Language Action ModelSigmoid Loss forLanguage Image…Sigmoid Loss for Language Image Pre-Training過去の参考文献中心の論文この論文を引用する論文古い新しい

ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。