DINOv3

Self-supervised learning holds the promise of eliminating the need for manual data annotation, enabling models to scale effortlessly to massive datasets and larger architectures. By not being tailored to specific tasks or domains, this training paradigm has the potential to learn visual representations from diverse sources, ranging from natural to aerial images -- using a single algorithm. This technical report introduces DINOv3, a major milestone toward realizing this vision by leveraging simple yet effective strategies. First, we leverage the benefit of scaling both dataset and model size by careful data preparation, design, and optimization. Second, we introduce a new method called Gram anchoring, which effectively addresses the known yet unsolved issue of dense feature maps degrading during long training schedules. Finally, we apply post-hoc strategies that further enhance our models' flexibility with respect to resolution, model size, and alignment with text. As a result, we present a versatile vision foundation model that outperforms the specialized state of the art across a broad range of settings, without fine-tuning. DINOv3 produces high-quality dense features that achieve outstanding performance on various vision tasks, significantly surpassing previous self- and weakly-supervised foundation models. We also share the DINOv3 suite of vision models, designed to advance the state of the art on a wide spectrum of tasks and data by providing scalable solutions for diverse resource constraints and deployment scenarios.

Are Large-scale DatasetsNecessary for…Are Large-scale Datasets Necessary for Self-Supervised Pre-training?iBOT: Image BERTPre-Training with Onlin…iBOT: Image BERT Pre-Training with Online TokenizerSelf-supervisedPretraining of Visual…Self-supervised Pretraining of Visual Features in the WildAn Empirical Study ofTraining Self-Supervise…An Empirical Study of Training Self-Supervised Vision TransformersBEiT: BERT Pre-Trainingof Image TransformersBEiT: BERT Pre-Training of Image TransformersSelf-Supervised Learningfrom Images with a…Self-Supervised Learning from Images with a Joint-Embedding Predictive ArchitectureRevisiting FeaturePrediction for Learning…Revisiting Feature Prediction for Learning Visual Representations from VideoEVA-CLIP-18B: ScalingCLIP to 18 Billion…EVA-CLIP-18B: Scaling CLIP to 18 Billion ParametersSigLIP 2: MultilingualVision-Language Encoder…SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense FeaturesPerception Encoder: Thebest visual embeddings…Perception Encoder: The best visual embeddings are not at the output of the networkV-JEPA 2:Self-Supervised Video…V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and PlanningScaling Language-FreeVisual Representation…Scaling Language-Free Visual Representation LearningVision-CentricActivation and…Vision-Centric Activation and Coordination for Multimodal Large Language ModelsIn Pursuit of PixelSupervision for Visual…In Pursuit of Pixel Supervision for Visual Pre-trainingUnleashing the IntrinsicVisual Representation…Unleashing the Intrinsic Visual Representation Capability of Multimodal Large Language ModelsScaling ImageGeo-Localization to…Scaling Image Geo-Localization to Continent LevelUnlocking 3D AffordanceSegmentation with 2D…Unlocking 3D Affordance Segmentation with 2D Semantic KnowledgeImage Generators areGeneralist Vision…Image Generators are Generalist Vision LearnersMindVLA-U1: VLA Beats VAwith Unified Streaming…MindVLA-U1: VLA Beats VA with Unified Streaming Architecture for Autonomous DrivingCoME-VL: ScalingComplementary…CoME-VL: Scaling Complementary Multi-Encoder Vision-Language LearningVision-aligned LatentReasoning for…Vision-aligned Latent Reasoning for Multi-modal Large Language ModelImplicit NeuralRepresentation…Implicit Neural Representation Facilitates Unified Universal Vision EncodingLook Before Acting:Enhancing Vision…Look Before Acting: Enhancing Vision Foundation Representations for Vision-Language-Action ModelsSpaRRTa: A SyntheticBenchmark for Evaluatin…SpaRRTa: A Synthetic Benchmark for Evaluating Spatial Intelligence in Visual Foundation ModelsDINOv3DINOv3Earlier referencesFocus paperCiting papersOlderNewer

Click a node to pin it, click the empty canvas to go back to this paper, or hover to preview. Open a node’s page from its title.