DINOv2: Learning Robust Visual Features without Supervision

The recent breakthroughs in natural language processing for model pretraining on large quantities of data have opened the way for similar foundation models in computer vision. These models could greatly simplify the use of images in any system by producing all-purpose visual features, i.e., features that work across image distributions and tasks without finetuning. This work shows that existing pretraining methods, especially self-supervised methods, can produce such features if trained on enough curated data from diverse sources. We revisit existing approaches and combine different techniques to scale our pretraining in terms of data and model size. Most of the technical contributions aim at accelerating and stabilizing the training at scale. In terms of data, we propose an automatic pipeline to build a dedicated, diverse, and curated image dataset instead of uncurated data, as typically done in the self-supervised literature. In terms of models, we train a ViT model (Dosovitskiy et al., 2020) with 1B parameters and distill it into a series of smaller models that surpass the best available all-purpose features, OpenCLIP (Ilharco et al., 2021) on most of the benchmarks at image and pixel levels.

openalex_id:w3182707920openalex_id:w3182707920Bootstrap Your OwnLatent - A New Approach…Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningA Simple Framework forContrastive Learning of…A Simple Framework for Contrastive Learning of Visual RepresentationsUnsupervised Learning ofVisual Features by…Unsupervised Learning of Visual Features by Contrasting Cluster AssignmentsEmerging Properties inSelf-Supervised Vision…Emerging Properties in Self-Supervised Vision TransformersDivide and Contrast:Self-supervised Learnin…Divide and Contrast: Self-supervised Learning from Uncurated DataSelf-supervisedPretraining of Visual…Self-supervised Pretraining of Visual Features in the WildAn Empirical Study ofTraining Self-Supervise…An Empirical Study of Training Self-Supervised Vision TransformersiBOT: Image BERTPre-Training with Onlin…iBOT: Image BERT Pre-Training with Online TokenizerBEiT: BERT Pre-Trainingof Image TransformersBEiT: BERT Pre-Training of Image TransformersEfficientSelf-supervised Vision…Efficient Self-supervised Vision Transformers for Representation LearningLAION-5B: An openlarge-scale dataset for…LAION-5B: An open large-scale dataset for training next generation image-text modelsMonoVAN: VisualAttention for…MonoVAN: Visual Attention for Self-Supervised Monocular Depth EstimationEnhancing diagnosticdeep learning via…Enhancing diagnostic deep learning via self-supervised pretraining on large-scale, unlabeled non-medical imagesMultimodal Whole SlideFoundation Model for…Multimodal Whole Slide Foundation Model for PathologyV-STRONG: VisualSelf-Supervised…V-STRONG: Visual Self-Supervised Traversability Learning for Off-road NavigationRudolfV: A FoundationModel by Pathologists…RudolfV: A Foundation Model by Pathologists for PathologistsA Survey of AI-GeneratedVideo EvaluationA Survey of AI-Generated Video EvaluationMotionBooth:Motion-Aware Customized…MotionBooth: Motion-Aware Customized Text-to-Video GenerationSSR-Encoder: EncodingSelective Subject…SSR-Encoder: Encoding Selective Subject Representation for Subject-Driven GenerationCobra: Extending Mambato Multi-Modal Large…Cobra: Extending Mamba to Multi-Modal Large Language Model for Efficient InferencePhantom:Subject-Consistent Vide…Phantom: Subject-Consistent Video Generation via Cross-Modal AlignmentMOOZY: A Patient-FirstFoundation Model for…MOOZY: A Patient-First Foundation Model for Computational PathologyImage Generators areGeneralist Vision…Image Generators are Generalist Vision LearnersDINOv2: Learning RobustVisual Features without…DINOv2: Learning Robust Visual Features without Supervision過去の参考文献中心の論文この論文を引用する論文古い新しい

ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。