DenseCLIP: Language-Guided Dense Prediction with Context-Aware Prompting

Recent progress has shown that large-scale pre-training using contrastive image-text pairs can be a promising alternative for high-quality visual representation learning from natural language supervision. Benefiting from a broader source of supervision, this new paradigm exhibits impressive transferability to downstream classification tasks and datasets. However, the problem of transferring the knowledge learned from image-text pairs to more complex dense prediction tasks has barely been visited. In this work, we present a new framework for dense prediction by implicitly and explicitly leveraging the pre-trained knowledge from CLIP. Specifically, we convert the original image-text matching problem in CLIP to a pixel-text matching problem and use the pixel-text score maps to guide the learning of dense prediction models. By further using the contextual information from the image to prompt the language model, we are able to facilitate our model to better exploit the pre-trained knowledge. Our method is model-agnostic, which can be applied to arbitrary dense prediction systems and various pre-trained visual backbones including both CLIP models and ImageNet pre-trained models. Extensive experiments demonstrate the superior performance of our methods on semantic segmentation, object detection, and instance segmentation tasks. Code is available at https://github.com/raoyongming/DenseCLIP

Faster R-CNN: TowardsReal-Time Object…Faster R-CNN: Towards Real-Time Object Detection with Region Proposal NetworksDeep Residual Learningfor Image RecognitionDeep Residual Learning for Image RecognitionPyramid Scene ParsingNetworkPyramid Scene Parsing NetworkFeature Pyramid Networksfor Object DetectionFeature Pyramid Networks for Object DetectionSemantic Understandingof Scenes Through the…Semantic Understanding of Scenes Through the ADE20K DatasetPanoptic Feature PyramidNetworksPanoptic Feature Pyramid NetworksPyramid VisionTransformer: A Versatil…Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without ConvolutionsSwin Transformer:Hierarchical Vision…Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsTip-Adapter:Training-free…Tip-Adapter: Training-free CLIP-Adapter for Better Vision-Language ModelingMasked Autoencoders AreScalable Vision LearnersMasked Autoencoders Are Scalable Vision LearnersCLIP-Adapter: BetterVision-Language Models…CLIP-Adapter: Better Vision-Language Models with Feature AdaptersCPT: Colorful PromptTuning for Pre-trained…CPT: Colorful Prompt Tuning for Pre-trained Vision-Language ModelsZero-Shot TemporalAction Detection via…Zero-Shot Temporal Action Detection via Vision-Language PromptingCan Language UnderstandDepth?Can Language Understand Depth?Effective Adaptation inMulti-Task Co-Training…Effective Adaptation in Multi-Task Co-Training for Unified Autonomous DrivingPrompt-aligned Gradientfor Prompt TuningPrompt-aligned Gradient for Prompt TuningCALIP: Zero-ShotEnhancement of CLIP wit…CALIP: Zero-Shot Enhancement of CLIP with Parameter-Free AttentionCLAMP: Prompt-basedContrastive Learning fo…CLAMP: Prompt-based Contrastive Learning for Connecting Language and Animal PosePLOT: Prompt Learningwith Optimal Transport…PLOT: Prompt Learning with Optimal Transport for Vision-Language ModelsCLIP-ReID: ExploitingVision-Language Model…CLIP-ReID: Exploiting Vision-Language Model for Image Re-Identification without Concrete Text LabelsGeneralizing MultipleObject Tracking to…Generalizing Multiple Object Tracking to Unseen Domains by Introducing Natural Language RepresentationConcept-Guided PromptLearning for…Concept-Guided Prompt Learning for Generalization in Vision-Language ModelsVadCLIP: AdaptingVision-Language Models…VadCLIP: Adapting Vision-Language Models for Weakly Supervised Video Anomaly DetectionToken-Level ContrastiveLearning with…Token-Level Contrastive Learning with Modality-Aware Prompting for Multimodal Intent RecognitionDenseCLIP:Language-Guided Dense…DenseCLIP: Language-Guided Dense Prediction with Context-Aware PromptingEarlier referencesFocus paperCiting papersOlderNewer

Click a node to pin it, click the empty canvas to go back to this paper, or hover to preview. Open a node’s page from its title.