Interpreting CLIP's Image Representation via Text-Based Decomposition

We investigate the CLIP image encoder by analyzing how individual model components affect the final representation. We decompose the image representation as a sum across individual image patches, model layers, and attention heads, and use CLIP's text representation to interpret the summands. Interpreting the attention heads, we characterize each head's role by automatically finding text representations that span its output space, which reveals property-specific roles for many heads (e.g. location or shape). Next, interpreting the image patches, we uncover an emergent spatial localization within CLIP. Finally, we use this understanding to remove spurious features from CLIP and to create a strong zero-shot image segmenter. Our results indicate that a scalable understanding of transformer models is attainable and can be used to repair and improve models.

ImageNet Auto-Annotationwith Segmentation…ImageNet Auto-Annotation with Segmentation PropagationUnderstanding Deep ImageRepresentations by…Understanding Deep Image Representations by Inverting ThemInverting ConvolutionalNetworks with…Inverting Convolutional Networks with Convolutional NetworksIntrusion Detection inComputer Networks using…Intrusion Detection in Computer Networks using Latent Space Representation and Machine LearningAxiomatic Attributionfor Deep NetworksAxiomatic Attribution for Deep NetworksDistributionally RobustNeural Networks for…Distributionally Robust Neural Networks for Group Shifts: On the Importance of Regularization for Worst-Case GeneralizationAnalyzing Multi-HeadSelf-Attention…Analyzing Multi-Head Self-Attention: Specialized Heads Do the Heavy Lifting, the Rest Can Be PrunedUnderstanding the Roleof Individual Units in…Understanding the Role of Individual Units in a Deep Neural NetworkZero-Shot Text-to-ImageGenerationZero-Shot Text-to-Image GenerationExplaining in Style:Training a GAN to…Explaining in Style: Training a GAN to explain a classifier in StyleSpaceNatural LanguageDescriptions of Deep…Natural Language Descriptions of Deep Visual FeaturesDeep Saliency Prior forReducing Visual…Deep Saliency Prior for Reducing Visual DistractionInterpreting CLIP withSparse Linear Concept…Interpreting CLIP with Sparse Linear Concept Embeddings (SpLiCE)Explainable andInterpretable Multimoda…Explainable and Interpretable Multimodal Large Language Models: A Comprehensive SurveyTranscoders FindInterpretable LLM…Transcoders Find Interpretable LLM Feature CircuitsUnderstanding MultimodalLLMs: the Mechanistic…Understanding Multimodal LLMs: the Mechanistic Interpretability of Llava in Visual Question AnsweringBlack-Box Access isInsufficient for…Black-Box Access is Insufficient for Rigorous AI AuditsMMNeuron: DiscoveringNeuron-Level…MMNeuron: Discovering Neuron-Level Domain-Specific Interpretation in Multimodal Large Language ModelInterpreting and EditingVision-Language…Interpreting and Editing Vision-Language Representations to Mitigate HallucinationsSparse Feature Circuits:Discovering and Editing…Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language ModelsTowards PrincipledEvaluations of Sparse…Towards Principled Evaluations of Sparse Autoencoders for Interpretability and ControlConceptAttention:Diffusion Transformers…ConceptAttention: Diffusion Transformers Learn Highly Interpretable FeaturesDebiasing CLIP:Interpreting and…Debiasing CLIP: Interpreting and Correcting Bias in Attention HeadsHow VisualRepresentations Map to…How Visual Representations Map to Language Feature Space in Multimodal LLMsInterpreting CLIP'sImage Representation vi…Interpreting CLIP's Image Representation via Text-Based Decomposition過去の参考文献中心の論文この論文を引用する論文古い新しい

ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。