Depth Anything 3: Recovering the Visual Space from Any Views

We present Depth Anything 3 (DA3), a model that predicts spatially consistent geometry from an arbitrary number of visual inputs, with or without known camera poses. In pursuit of minimal modeling, DA3 yields two key insights: a single plain transformer (e.g., vanilla DINO encoder) is sufficient as a backbone without architectural specialization, and a singular depth-ray prediction target obviates the need for complex multi-task learning. Through our teacher-student training paradigm, the model achieves a level of detail and generalization on par with Depth Anything 2 (DA2). We establish a new visual geometry benchmark covering camera pose estimation, any-view geometry and visual rendering. On this benchmark, DA3 sets a new state-of-the-art across all tasks, surpassing prior SOTA VGGT by an average of 44.3% in camera pose accuracy and 25.1% in geometric accuracy. Moreover, it outperforms DA2 in monocular depth estimation. All models are trained exclusively on public academic datasets.

The Replica Dataset: ADigital Replica of…The Replica Dataset: A Digital Replica of Indoor SpacesDeepV2D: Video to Depthwith Differentiable…DeepV2D: Video to Depth with Differentiable Structure from MotionSplatt3R: Zero-shotGaussian Splatting from…Splatt3R: Zero-shot Gaussian Splatting from Uncalibrated Image PairsDINOv2: Learning RobustVisual Features without…DINOv2: Learning Robust Visual Features without Supervisionπ3: ScalablePermutation-Equivariant…π3: Scalable Permutation-Equivariant Visual Geometry LearningMapAnything: UniversalFeed-Forward Metric 3D…MapAnything: Universal Feed-Forward Metric 3D ReconstructionStreaming 4D VisualGeometry TransformerStreaming 4D Visual Geometry TransformerDepth Pro: SharpMonocular Metric Depth…Depth Pro: Sharp Monocular Metric Depth in Less Than a SecondVGGT-SLAM: Dense RGBSLAM Optimized on the…VGGT-SLAM: Dense RGB SLAM Optimized on the SL(4) ManifoldUniDepthV2: UniversalMonocular Metric Depth…UniDepthV2: Universal Monocular Metric Depth Estimation Made SimplerStructured 3D Latentsfor Scalable and…Structured 3D Latents for Scalable and Versatile 3D GenerationOmniWorld: AMulti-Domain and…OmniWorld: A Multi-Domain and Multi-Modal Dataset for 4D World ModelingMapAnything: UniversalFeed-Forward Metric 3D…MapAnything: Universal Feed-Forward Metric 3D ReconstructionKV-Tracker: Real-TimePose Tracking with…KV-Tracker: Real-Time Pose Tracking with TransformersGeometric ContextTransformer for…Geometric Context Transformer for Streaming 3D ReconstructionR3: 3D Reconstructionvia Relative RegressionR3: 3D Reconstruction via Relative RegressionOmniStream: MasteringPerception…OmniStream: Mastering Perception, Reconstruction and Action in Continuous StreamsNeoVerse: Enhancing 4DWorld Model with…NeoVerse: Enhancing 4D World Model with in-the-wild Monocular VideosImage Generators areGeneralist Vision…Image Generators are Generalist Vision LearnersEgoGrasp: World-SpaceHand-Object Interaction…EgoGrasp: World-Space Hand-Object Interaction Estimation from Egocentric VideosAirZoo: A UnifiedLarge-Scale Dataset for…AirZoo: A Unified Large-Scale Dataset for Grounding Aerial Geometric 3D VisionSelf-Improving 4DPerception via…Self-Improving 4D Perception via Self-DistillationRobo3R: EnhancingRobotic Manipulation…Robo3R: Enhancing Robotic Manipulation with Accurate Feed-Forward 3D ReconstructionAny3D-VLA: Enhancing VLARobustness via Diverse…Any3D-VLA: Enhancing VLA Robustness via Diverse Point CloudsDepth Anything 3:Recovering the Visual…Depth Anything 3: Recovering the Visual Space from Any Views過去の参考文献中心の論文この論文を引用する論文古い新しい

ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。