VGGT-SLAM
VGGT-SLAM 將 VGGT 產生的子地圖逐步對齊,指出在未校正相機下重建只確定到 15 自由度的射影變換,因此以 SL(4) 流形上的單應矩陣取代相似變換對齊子地圖,並加入以 SALAD 檢索的迴圈約束。作者明言重建不具公制尺度,影像須先去除鏡頭畸變,且當多張影像只看到單一平面(TUM 僅拍地板的片段)時單應估計會退化並使重建發散。NeurIPS 正式版補充的焦距統計顯示,VGGT 在同一場景內估計的焦距會明顯波動(桌面與路樁場景標準差 37.1 與 51.8 像素),這正是 SL(4) 明顯優於 Sim(3) 的情境;但在 7-Scenes 與 TUM 等一般場景,Sim(3) 版本表現相近,附錄中 w = 8 的 Sim(3) 版本在 TUM 平均 ATE 甚至更低(0.040 m 對 0.053 m)。每個子地圖的 VGGT 推論約 662 ms,SL(4) 對齊只多約 17 ms。建築室內若出現只拍到大面樓板或牆面的連續影格,可能觸發同類退化,需以實測驗證(推論)。
本頁內容
Aligns VGGT submaps with 15-DoF homographies on SL(4) plus loop closures to handle projective ambiguity of uncalibrated reconstruction.
技術屬性
欄位內容為文獻擷取紀錄的原文用語(英文),以原文為據;「未查證」表示本研究尚未讀到該資訊,不代表該方法不具備此能力。
| 感測輸入 | monocular camera (uncalibrated) |
|---|---|
| 原文測試平台 | 未記錄 |
| 狀態估計 | nonlinear factor-graph optimization on the SL(4) manifold estimating 15-DoF homographies between VGGT submaps |
| 資料關聯 | shared frames between submaps; homography estimation with 5-point RANSAC on VGGT points |
| 時間表示 | discrete poses |
| 去畸變 | 不適用 |
| 迴圈閉合 | SALAD image-descriptor retrieval + relative homography loop constraints |
| 全域最佳化 | SL(4) factor graph over odometry and loop-closure constraints |
| 地圖表示 | VGGT dense point-cloud submaps |
| 先驗資訊 | VGGT learned feed-forward reconstruction prior |
| 可輸出幾何 | dense point cloud and trajectory; reconstruction not in metric scale (Sec. 1) |
| 計算需求 | NVIDIA GeForce RTX 4090 (24 GB) with AMD Ryzen Threadripper 7960X; per submap on office_loop with w = 16: keyframe detection 176 ms, VGGT inference 662 ms, loop-closure detection 105 ms, relative transformation 28 ms for SL(4) versus 11 ms for Sim(3), back end 0.5 ms (Table 4, NeurIPS version); VGGT alone limited to about 60 images on 24 GB (Sec. 1) |
使用設備
原文使用的感測器、運算硬體與載具(equipment)。型號保留原文寫法,連結到設備頁中同一型號的歸併名稱;角色依原文用途分為方法輸入、資料集感測器、執行運算平台、參考或真值量測(reference or ground truth)與比較對象設備。
| 類別 | 型號(原文寫法) | 角色 | 資料集 | 原文規格 | 出處 |
|---|---|---|---|---|---|
| 相機 | 原文未報告 (monocular cameras for custom office-loop, tabletop and bollards scenes) | 方法輸入 | 未標示 | A single camera per scene, different scenes may use different cameras; intrinsics unknown to the method | (Maggio et al., 2025, Sec. 5.5; Appendix B.3; Appendix C.1) |
| 運算硬體 | NVIDIA GeForce RTX 4090 (24 GB) | 執行運算平台 | 未標示 | With AMD Ryzen Threadripper 7960X CPU; VGGT limited to about 60 images at once on this GPU | (Maggio et al., 2025, Sec. 1; Sec. 5.1; Table 4) |
論文圖片
只收錄原文以開放授權(open license)釋出的圖片,並依授權條件標示出處、圖號、授權與修改方式。

Fig. 1以 Sim(3) 與 SL(4) 對齊 6 個 VGGT 子地圖的比較(Clio 公寓與隔間場景),顯示射影歧義
出處:Maggio et al., 2025,Fig. 1。授權:CC BY 4.0 (arXiv v2)。原始圖檔。修改:縮小至寬度不超過 1400 px,並轉存為 WebP 格式。

Fig. 27-Scenes office 場景 8 個子地圖,以及 55 m 辦公走廊迴圈 22 個子地圖的重建與位姿
出處:Maggio et al., 2025,Fig. 2。授權:CC BY 4.0 (arXiv v2)。原始圖檔。修改:縮小至寬度不超過 1400 px,並轉存為 WebP 格式。

Fig. 6 (arXiv v2; Fig. 8 in NeurIPS version)戶外儲槽周邊黃色路樁場景:Sim(3) 對齊失敗產生重影,SL(4) 可修正
出處:Maggio et al., 2025,Fig. 6 (arXiv v2; Fig. 8 in NeurIPS version)。授權:CC BY 4.0 (arXiv v2)。原始圖檔。修改:轉存為 WebP 格式。

Fig. 9 (arXiv v2; Fig. 11 in NeurIPS version)TUM room 場景 6 個子地圖的稠密重建,相機位姿依子地圖著色
出處:Maggio et al., 2025,Fig. 9 (arXiv v2; Fig. 11 in NeurIPS version)。授權:CC BY 4.0 (arXiv v2)。原始圖檔。修改:縮小至寬度不超過 1400 px,並轉存為 WebP 格式。
作者報告的優勢與限制
優勢
- Best accuracy and Chamfer RMSE on 7-Scenes dense evaluation following the MASt3R-SLAM protocol, by a small margin (Chamfer 0.055 m vs 0.056 m for uncalibrated MASt3R-SLAM); baseline values in this table differ from those reported in (Murai et al., 2025) Table 3 (Sec. 5.3; Table 3)
- Handles long videos infeasible for VGGT alone (abstract)
- SL(4) alignment costs only about 17 ms more per submap than Sim(3), about 2.5% of VGGT inference time; back-end optimization about 0.5 ms (Sec. 5.4; Table 4)
- VGGT focal-length estimates vary within a scene (std 37.1 and 51.8 px in the tabletop and bollards scenes versus 7.3 and 9.0 px in office loop and 7-Scenes), matching where SL(4) clearly beats Sim(3) (Appendix B.3; Table 10)
- Qualitative 55 m office-corridor loop with 22 submaps closed into a consistent map (Sec. 5.5; Fig. 2)
限制
- Homography estimation degenerate for planar points; unstable on TUM planar floor scene (Sec. 6)
- Vulnerable to outliers from VGGT points (Sec. 6)
- Additional drift modes including scene perspective (Sec. 6)
- Not metric scale (Sec. 1)
- Images must be undistorted because lens distortion is not rectified by the homography (Sec. 6, NeurIPS version)
- In the appendix the Sim(3) variant with w = 8 attains a lower TUM average ATE (0.040 m) than the SL(4) w = 32 configuration highlighted in the main text (0.053 m); SL(4) with w = 1 is numerically unstable on TUM floor and 360 (Appendix B.1; Table 6)
- Calibrated baseline numbers are copied from MASt3R-SLAM and the evo trajectory alignment mode is not stated (Sec. 5.1)
營建工程相關證據
論文未涉及營建場域。作者觀察到當多張影像只看到平坦地板時,15 自由度單應估計出現非唯一解並使重建發散;建築室內若有只拍到大面樓板或牆面的連續影格,可能觸發同類退化,需另行驗證(推論)。
原文驗證環境:公開基準
報告的性能數據
以下是原文作者報告的性能數值(author-reported results),不是本研究重新量測的結果。每張圖只並列同一個比較組(comparison group,同一張表、同一組實驗設定)內的方法;不同比較組之間的數值不可直接比較,也不構成排名。
本方法共出現在 4 個比較組,合計 27 筆紀錄。
Maggio et al., 2025 · Table 2 本方法 10 筆
指標ATE RMSE [m]
表格設定(擷取紀錄原文):ATE RMSE on TUM RGB-D (RGB only) computed with evo (alignment not stated); uncalibrated rows; floor sequence degenerate for SL(4) homography (Maggio et al., 2025, Table 2)
ATE RMSE [m],TUM RGB-D · 360
只並列這張表在相同設定下報告的方法;以「本方法:」開頭者為本頁方法。失敗、未執行與未報告以標記呈現,不是 0。
按 Tab 進入圖表後,用上下方向鍵逐一瀏覽各類別,Esc 關閉提示框;也可開啟表格檢視閱讀全部數值。
這些是 Maggio et al., 2025 在此表設定下報告的數值(author-reported results),只能在同一個比較組內對照,不代表方法在其他資料或設定下的表現。
資料來源作者報告值(Maggio et al., 2025, Table 2)
| 方法(原文寫法) | 報告值 | 出處 |
|---|---|---|
| DROID-SLAM* | 0.202 m | (Maggio et al., 2025, Table 2) |
| MASt3R-SLAM* | 0.07 m | (Maggio et al., 2025, Table 2) |
| Ours (Sim(3), w = 32) | 0.123 m | (Maggio et al., 2025, Table 2) |
| Ours (SL(4), w = 32)本方法原文提出 | 0.071 m | (Maggio et al., 2025, Table 2) |
Maggio et al., 2025 · Table 1 本方法 8 筆
指標ATE RMSE [m]
表格設定(擷取紀錄原文,這些數值分屬表中不同部分):(Maggio et al., 2025, Table 1)
- ATE RMSE on 7-Scenes computed with evo (alignment not stated); calibrated intrinsics; value reported from MASt3R-SLAM
- ATE RMSE on 7-Scenes computed with evo (alignment not stated); uncalibrated; DROID-SLAM* intrinsics from an automatic calibration pipeline, run by the authors
- ATE RMSE on 7-Scenes computed with evo (alignment not stated); uncalibrated; value reported from MASt3R-SLAM
- ATE RMSE on 7-Scenes computed with evo (alignment not stated); uncalibrated; VGGT-SLAM average of five runs
ATE RMSE [m],7-Scenes · chess
只並列這張表在相同設定下報告的方法;以「本方法:」開頭者為本頁方法。失敗、未執行與未報告以標記呈現,不是 0。
按 Tab 進入圖表後,用上下方向鍵逐一瀏覽各類別,Esc 關閉提示框;也可開啟表格檢視閱讀全部數值。
這些是 Maggio et al., 2025 在此表設定下報告的數值(author-reported results),只能在同一個比較組內對照,不代表方法在其他資料或設定下的表現。
資料來源作者報告值(Maggio et al., 2025, Table 1)
| 方法(原文寫法) | 報告值 | 出處 |
|---|---|---|
| NICER-SLAM | 0.033 m | (Maggio et al., 2025, Table 1) |
| DROID-SLAM | 0.036 m | (Maggio et al., 2025, Table 1) |
| MASt3R-SLAM | 0.053 m | (Maggio et al., 2025, Table 1) |
| DROID-SLAM* | 0.047 m | (Maggio et al., 2025, Table 1) |
| MASt3R-SLAM* | 0.063 m | (Maggio et al., 2025, Table 1) |
| Ours (Sim(3), w = 32) | 0.037 m | (Maggio et al., 2025, Table 1) |
| Ours (SL(4), w = 32)本方法原文提出 | 0.036 m | (Maggio et al., 2025, Table 1) |
Maggio et al., 2025 · Table 4 本方法 5 筆
資料集與序列office_loop (authors' custom sequence) · office_loop
表格設定(擷取紀錄原文):Runtime per stage on the custom office_loop sequence with window size w = 16, averaged over five runs; stage times cover all frames of a submap (NeurIPS version only) (Maggio et al., 2025, Table 4)
Keyframe detection time [ms],office_loop (authors' custom sequence) · office_loop
只並列這張表在相同設定下報告的方法;以「本方法:」開頭者為本頁方法。失敗、未執行與未報告以標記呈現,不是 0。
按 Tab 進入圖表後,用上下方向鍵逐一瀏覽各類別,Esc 關閉提示框;也可開啟表格檢視閱讀全部數值。
這些是 Maggio et al., 2025 在此表設定下報告的數值(author-reported results),只能在同一個比較組內對照,不代表方法在其他資料或設定下的表現。
資料來源作者報告值(Maggio et al., 2025, Table 4)
| 方法(原文寫法) | 報告值 | 出處 |
|---|---|---|
| VGGT-SLAM w/ Sim(3)硬體:NVIDIA GeForce RTX 4090 (24 GB) with AMD Ryzen Threadripper 7960X | 176 ms | (Maggio et al., 2025, Table 4 (NeurIPS version, Sec. 5.4)) |
| VGGT-SLAM w/ SL(4)本方法原文提出硬體:NVIDIA GeForce RTX 4090 (24 GB) with AMD Ryzen Threadripper 7960X | 176 ms | (Maggio et al., 2025, Table 4 (NeurIPS version, Sec. 5.4)) |
Maggio et al., 2025 · Table 3 本方法 4 筆
資料集與序列7-Scenes · 7-Scenes (single aggregate column; aggregation over sequences not stated)
表格設定(擷取紀錄原文,這些數值分屬表中不同部分):(Maggio et al., 2025, Table 3)
- Dense reconstruction on 7-Scenes following the MASt3R-SLAM protocol, RMSE in metres; calibrated; @n means a keyframe every n images
- Dense reconstruction on 7-Scenes following the MASt3R-SLAM protocol, RMSE in metres; uncalibrated; @n means a keyframe every n images
ATE [m],7-Scenes · 7-Scenes (single aggregate column; aggregation over sequences not stated)
只並列這張表在相同設定下報告的方法;以「本方法:」開頭者為本頁方法。失敗、未執行與未報告以標記呈現,不是 0。
- 不適用
按 Tab 進入圖表後,用上下方向鍵逐一瀏覽各類別,Esc 關閉提示框;也可開啟表格檢視閱讀全部數值。
這些是 Maggio et al., 2025 在此表設定下報告的數值(author-reported results),只能在同一個比較組內對照,不代表方法在其他資料或設定下的表現。
資料來源作者報告值(Maggio et al., 2025, Table 3)
| 方法(原文寫法) | 報告值 | 出處 |
|---|---|---|
| DROID-SLAM | 0.049 m | (Maggio et al., 2025, Table 3) |
| MASt3R-SLAM | 0.047 m | (Maggio et al., 2025, Table 3) |
| Spann3R @20 | 無數值不適用註記(擷取紀錄):不適用 (N/A) | (Maggio et al., 2025, Table 3) |
| Spann3R @2 | 無數值不適用註記(擷取紀錄):不適用 (N/A) | (Maggio et al., 2025, Table 3) |
| MASt3R-SLAM* | 0.066 m | (Maggio et al., 2025, Table 3) |
| Ours (Sim(3), w = 32) | 0.067 m | (Maggio et al., 2025, Table 3) |
| Ours (SL(4), w = 32)本方法原文提出 | 0.067 m | (Maggio et al., 2025, Table 3) |
來源
Maggio et al., 2025
(2025)VGGT-SLAM: Dense RGB SLAM Optimized on the SL(4) ManifoldAdvances in Neural Information Processing Systems 38 (NeurIPS 2025), pp. 143976-144004
DOI 10.52202/085713-4324arXiv 2505.12549程式碼
同儕審查已出版已讀全文近十年查證後修正
相關版本
- 預印本:arXiv:2505.12549 https://arxiv.org/abs/2505.12549
- 後續版本預印本:VGGT-SLAM 2.0: Real-time Dense Feed-forward Scene Reconstruction (arXiv:2601.19887); the code_url repository HEAD now documents this version https://arxiv.org/abs/2601.19887
程式碼:https://github.com/MIT-SPARK/VGGT-SLAM(授權:BSD-2-Clause)。有公開程式碼不等於已被重現,也不代表目前版本與論文版本相同。