VGGT-SLAM: Dense RGB SLAM Optimized on the SL(4) Manifold

We present VGGT-SLAM, a dense RGB SLAM system constructed by incrementally and globally aligning submaps created from the feed-forward scene reconstruction approach VGGT using only uncalibrated monocular cameras. While related works align submaps using similarity transforms (i.e., translation, rotation, and scale), we show that such approaches are inadequate in the case of uncalibrated cameras. In particular, we revisit the idea of reconstruction ambiguity, where given a set of uncalibrated cameras with no assumption on the camera motion or scene structure, the scene can only be reconstructed up to a 15-degrees-of-freedom projective transformation of the true geometry. This inspires us to recover a consistent scene reconstruction across submaps by optimizing over the SL(4) manifold, thus estimating 15-degrees-of-freedom homography transforms between sequential submaps while accounting for potential loop closure constraints. As verified by extensive experiments, we demonstrate that VGGT-SLAM achieves improved map quality using long video sequences that are infeasible for VGGT due to its high GPU requirements.

Quaternion kinematicsfor the error-state…Quaternion kinematics for the error-state Kalman filterA GeneralOptimization-based…A General Optimization-based Framework for Global Pose Estimation with Multiple SensorsSplatt3R: Zero-shotGaussian Splatting from…Splatt3R: Zero-shot Gaussian Splatting from Uncalibrated Image PairsDINOv2: Learning RobustVisual Features without…DINOv2: Learning Robust Visual Features without SupervisionPreF3R: Pose-FreeFeed-Forward 3D Gaussia…PreF3R: Pose-Free Feed-Forward 3D Gaussian Splatting from Variable-length Image SequenceMASt3R-SfM: AFully-Integrated…MASt3R-SfM: A Fully-Integrated Solution for Unconstrained Structure-from-Motion3D Reconstruction withSpatial Memory3D Reconstruction with Spatial MemoryVGGT: Visual GeometryGrounded TransformerVGGT: Visual Geometry Grounded TransformerMASt3R-SLAM: Real-TimeDense SLAM with 3D…MASt3R-SLAM: Real-Time Dense SLAM with 3D Reconstruction PriorsReloc3r: Large-ScaleTraining of Relative…Reloc3r: Large-Scale Training of Relative Camera Pose Regression for Generalizable, Fast, and Accurate Visual LocalizationPow3R: EmpoweringUnconstrained 3D…Pow3R: Empowering Unconstrained 3D Reconstruction with Camera and Scene PriorsContinuous 3D PerceptionModel with Persistent…Continuous 3D Perception Model with Persistent StateVGGT-Long: Chunk it,Loop it, Align it -…VGGT-Long: Chunk it, Loop it, Align it - Pushing VGGT's Limits on Kilometer-scale Long RGB SequencesTTT3R: 3D Reconstructionas Test-Time TrainingTTT3R: 3D Reconstruction as Test-Time TrainingDepth Anything 3:Recovering the Visual…Depth Anything 3: Recovering the Visual Space from Any ViewsViPE: Video Pose Enginefor 3D Geometric…ViPE: Video Pose Engine for 3D Geometric PerceptionAMB3R: AccurateFeed-forward…AMB3R: Accurate Feed-forward Metric-scale 3D Reconstruction with BackendAdvances in Feed-Forward3D Reconstruction and…Advances in Feed-Forward 3D Reconstruction and View Synthesis: A SurveySAIL-Recon: Large SfM byAugmenting Scene…SAIL-Recon: Large SfM by Augmenting Scene Regression with LocalizationScal3R: ScalableTest-Time Training for…Scal3R: Scalable Test-Time Training for Large-Scale 3D ReconstructionLoGeR: Long-ContextGeometric Reconstructio…LoGeR: Long-Context Geometric Reconstruction with Hybrid MemoryVGG-T3: OfflineFeed-Forward 3D…VGG-T3: Offline Feed-Forward 3D Reconstruction at ScaleR3: 3D Reconstructionvia Relative RegressionR3: 3D Reconstruction via Relative RegressionViSTA-SLAM: Visual SLAMwith Symmetric Two-View…ViSTA-SLAM: Visual SLAM with Symmetric Two-View AssociationVGGT-SLAM: Dense RGBSLAM Optimized on the…VGGT-SLAM: Dense RGB SLAM Optimized on the SL(4) ManifoldEarlier referencesFocus paperCiting papersOlderNewer

Click a node to pin it, click the empty canvas to go back to this paper, or hover to preview. Open a node’s page from its title.