SimVLM: Simple Visual Language Model Pretraining with Weak Supervision

With recent progress in joint modeling of visual and textual representations, Vision-Language Pretraining (VLP) has achieved impressive performance on many multimodal downstream tasks. However, the requirement for expensive annotations including clean image captions and regional labels limits the scalability of existing approaches, and complicates the pretraining procedure with the introduction of multiple dataset-specific objectives. In this work, we relax these constraints and present a minimalist pretraining framework, named Simple Visual Language Model (SimVLM). Unlike prior work, SimVLM reduces the training complexity by exploiting large-scale weak supervision, and is trained end-to-end with a single prefix language modeling objective. Without utilizing extra data or task-specific customization, the resulting model significantly outperforms previous pretraining methods and achieves new state-of-the-art results on a wide range of discriminative and generative vision-language benchmarks, including VQA (+3.74% vqa-score), NLVR2 (+1.17% accuracy), SNLI-VE (+1.37% accuracy) and image captioning tasks (+10.1% average CIDEr score). Furthermore, we demonstrate that SimVLM acquires strong generalization and transfer ability, enabling zero-shot behavior including open-ended visual question answering and cross-modality transfer.

VisualBERT: A Simple andPerformant Baseline for…VisualBERT: A Simple and Performant Baseline for Vision and LanguageLXMERT: LearningCross-Modality Encoder…LXMERT: Learning Cross-Modality Encoder Representations from TransformersVisual Entailment: ANovel Task for…Visual Entailment: A Novel Task for Fine-Grained Image UnderstandingLarge-Scale AdversarialTraining for…Large-Scale Adversarial Training for Vision-and-Language Representation LearningExploring the Limits ofTransfer Learning with…Exploring the Limits of Transfer Learning with a Unified Text-to-Text TransformerUNIMO: TowardsUnified-Modal…UNIMO: Towards Unified-Modal Understanding and Generation via Cross-Modal Contrastive LearningViLT:Vision-and-Language…ViLT: Vision-and-Language Transformer Without Convolution or Region SupervisionUnifyingVision-and-Language…Unifying Vision-and-Language Tasks via Text GenerationScaling Up Visual andVision-Language…Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionLearning TransferableVisual Models From…Learning Transferable Visual Models From Natural Language SupervisionERNIE-ViL: KnowledgeEnhanced Vision-Languag…ERNIE-ViL: Knowledge Enhanced Vision-Language Representations through Scene GraphsHow Much Can CLIPBenefit…How Much Can CLIP Benefit Vision-and-Language Tasks?mPLUG: Effective andEfficient…mPLUG: Effective and Efficient Vision-Language Learning by Cross-modal Skip-connectionsMulti-Grained VisionLanguage Pre-Training…Multi-Grained Vision Language Pre-Training: Aligning Texts with Visual ConceptsCoCa: ContrastiveCaptioners are…CoCa: Contrastive Captioners are Image-Text Foundation ModelsUNIMO-2: End-to-EndUnified Vision-Language…UNIMO-2: End-to-End Unified Vision-Language Grounded LearningAll You May Need for VQAare Image CaptionsAll You May Need for VQA are Image CaptionsBridgeTower: BuildingBridges between Encoder…BridgeTower: Building Bridges between Encoders in Vision-Language Representation LearningLarge-scale Multi-ModalPre-trained Models: A…Large-scale Multi-Modal Pre-trained Models: A Comprehensive SurveyMasked Vision andLanguage Modeling for…Masked Vision and Language Modeling for Multi-modal Representation LearningMasked Autoencoding DoesNot Help Natural…Masked Autoencoding Does Not Help Natural Language Supervision at ScaleMixGen: A NewMulti-Modal Data…MixGen: A New Multi-Modal Data AugmentationTowards Models that CanSee and ReadTowards Models that Can See and ReadTag2Text: GuidingVision-Language Model…Tag2Text: Guiding Vision-Language Model via Image TaggingSimVLM: Simple VisualLanguage Model…SimVLM: Simple Visual Language Model Pretraining with Weak Supervision過去の参考文献中心の論文この論文を引用する論文古い新しい

ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。