Do Convnets Learn Correspondence?

Convolutional neural nets (convnets) trained from massive labeled datasets have substantially improved the state-of-the-art in image classification and object detection. However, visual understanding requires establishing correspondence on a finer level than object category. Given their large pooling regions and training from whole-image labels, it is not clear that convnets derive their success from an accurate correspondence model which could be used for precise localization. In this paper, we study the effectiveness of convnet activation features for tasks requiring correspondence. We present evidence that convnet features localize at a much finer scale than their receptive field sizes, that they can be used to perform intraclass alignment as well as conventional hand-engineered features, and that they outperform conventional features in keypoint prediction on objects from PASCAL VOC 2011.

Gradient-based learningapplied to document…Gradient-based learning applied to document recognitionImageNet: A large-scalehierarchical image…ImageNet: A large-scale hierarchical image databaseArticulated part-basedmodel for joint object…Articulated part-based model for joint object detection and pose estimationImageNet Classificationwith Deep Convolutional…ImageNet Classification with Deep Convolutional Neural NetworksArticulated PoseEstimation Using…Articulated Pose Estimation Using Discriminative Armlet ClassifiersRich Feature Hierarchiesfor Accurate Object…Rich Feature Hierarchies for Accurate Object Detection and Semantic SegmentationDescriptor Matching withConvolutional Neural…Descriptor Matching with Convolutional Neural Networks: a Comparison to SIFTCaffe: ConvolutionalArchitecture for Fast…Caffe: Convolutional Architecture for Fast Feature EmbeddingDeCAF: A DeepConvolutional Activatio…DeCAF: A Deep Convolutional Activation Feature for Generic Visual RecognitionThe Pascal Visual ObjectClasses Challenge: A…The Pascal Visual Object Classes Challenge: A RetrospectiveDeepPose: Human PoseEstimation via Deep…DeepPose: Human Pose Estimation via Deep Neural NetworksUsing k-Poselets forDetecting People and…Using k-Poselets for Detecting People and Localizing Their KeypointsViewpoints and KeypointsViewpoints and KeypointsFully ConvolutionalNetworks for Semantic…Fully Convolutional Networks for Semantic SegmentationDeformable Part Modelsare Convolutional Neura…Deformable Part Models are Convolutional Neural NetworksFlowWeb: Joint image setalignment by weaving…FlowWeb: Joint image set alignment by weaving consistent, pixel-wise correspondencesProposal FlowProposal FlowLearning DenseCorrespondence via…Learning Dense Correspondence via 3D-Guided Cycle ConsistencyFCSS: FullyConvolutional…FCSS: Fully Convolutional Self-Similarity for Dense Semantic CorrespondenceDeep Semantic FeatureMatchingDeep Semantic Feature MatchingConvolutional NeuralNetwork Architecture fo…Convolutional Neural Network Architecture for Geometric MatchingSCNet: Learning SemanticCorrespondenceSCNet: Learning Semantic CorrespondenceWeakly SupervisedManifold Learning for…Weakly Supervised Manifold Learning for Dense Semantic Object CorrespondenceMulti-Image SemanticMatching by Mining…Multi-Image Semantic Matching by Mining Consistent FeaturesDo Convnets LearnCorrespondence?Do Convnets Learn Correspondence?Earlier referencesFocus paperCiting papersOlderNewer

Click a node to pin it, click the empty canvas to go back to this paper, or hover to preview. Open a node’s page from its title.