Loopy-SLAM
Loopy-SLAM 在 Point-SLAM 的神經點雲上加入子地圖、詞袋式全域地點辨識與穩健位姿圖最佳化,迴圈閉合後直接剛性平移子地圖中的點以修正地圖,毋須保存全部歷史影格。作者未研究光束調整精修,且實作尚非即時。
本頁內容
Adds submaps, BoW place recognition and robust pose-graph optimization to a neural point-cloud SLAM so maps can be corrected after loop closure.
技術屬性
欄位內容為文獻擷取紀錄的原文用語(英文),以原文為據;「未查證」表示本研究尚未讀到該資訊,不代表該方法不具備此能力。
| 感測輸入 | RGB-D |
|---|---|
| 原文測試平台 | 未記錄 |
| 狀態估計 | frame-to-model tracking on neural point submaps + robust pose graph optimization (line process) over global keyframes |
| 資料關聯 | Frame-to-model tracking by minimizing depth and colour re-rendering losses on the active submap (Point-SLAM style); loop edges from coarse-to-fine dense registration of submap surfaces: all depth frames of a submap are TSDF-fused and points sampled on the marching-cubes surface, FPFH features with RANSAC give the coarse alignment and ICP on full-resolution clouds refines it; loop edges pre-filtered by constraint translation magnitude and a fitness (overlap) score (Sec. 3.1, 3.2) |
| 時間表示 | discrete poses |
| 去畸變 | 不適用 |
| 迴圈閉合 | Global place recognition with a bag-of-visual-words database (DBoW3) queried when a submap is completed (top K = 4 on Replica, 1 on TUM-RGBD and ScanNet, dynamic similarity threshold); on ScanNet a PGO is triggered on average every 151 frames (Sec. 3.2, Sec. 4, App. E Table 8) |
| 全域最佳化 | Robust pose graph optimization of global-keyframe corrections with a dense surface-registration objective and a line process weighting loop edges, solved with Levenberg-Marquardt in two stages; submaps and frame poses then corrected rigidly by shifting points; feature fusion and colour and geometry feature refinement at the end of capture; bundle adjustment not studied (Sec. 3.2, App. D) |
| 地圖表示 | submaps of neural point clouds |
| 先驗資訊 | none |
| 可輸出幾何 | Mesh from TSDF fusion (1 cm voxels) of depth and colour rendered every fifth frame along the estimated trajectory, following Point-SLAM; rendered RGB-D images (Sec. 4 implementation details) |
| 計算需求 | PyTorch and Open3D through Python bindings, not optimized for real time; all experiments on NVIDIA GPUs with at most 12 GB memory; Replica office 0: tracking 0.85 s and mapping 9.85 s per frame (identical to the Point-SLAM row); on fr1 desk 7 PGOs of about 1 ms each, each needing about 8 global registrations of about 12 s on an Intel Core i7-11800H (Sec. 4.4, Table 5, App. B, App. E) |
使用設備
原文使用的感測器、運算硬體與載具(equipment)。型號保留原文寫法,連結到設備頁中同一型號的歸併名稱;角色依原文用途分為方法輸入、資料集感測器、執行運算平台、參考或真值量測(reference or ground truth)與比較對象設備。
| 類別 | 型號(原文寫法) | 角色 | 資料集 | 原文規格 | 出處 |
|---|---|---|---|---|---|
| RGB-D 相機 | RGBD camera (model not named) | 方法輸入 | Replica; TUM-RGBD; ScanNet | RGBD stream is the only input; TUM-RGBD and ScanNet are real-world data, the Replica trajectories come from a simulated RGBD sensor | (Liso et al., 2024, Abstract; Fig. 2; Sec. 4 Datasets) |
| 運算硬體 | Nvidia GPUs with a maximum memory of 12 GB (models not named) | 執行運算平台 | 未標示 | all results gathered on various GPUs of at most 12 GB | (Liso et al., 2024, Sec. 4.4; App. B) |
| 運算硬體 | 11th Gen Intel Core i7-11800H | 執行運算平台 | TUM-RGBD | CPU used for the 12 s per registration timing | (Liso et al., 2024, App. E) |
| 運算硬體 | AMD EPYC 7742 | 執行運算平台 | TUM-RGBD | processor used for the global-registration ablation timings | (Liso et al., 2024, App. E, Table 9) |
| 其他 | external motion capture system (not named) | 參考或真值量測 | TUM-RGBD | source of TUM-RGBD ground-truth poses | (Liso et al., 2024, Sec. 4 Datasets) |
作者報告的優勢與限制
優勢
- Replica: lowest average ATE RMSE (0.29 cm vs 0.35 cm GO-SLAM and 0.52 cm Point-SLAM) and best average depth L1 (0.35 cm) and F1 at 1 cm (90.77%) (Table 1, Table 11)
- TUM-RGBD: average ATE 3.85 cm vs 8.92 cm for Point-SLAM with about 14% more scene points (Table 2, Table 6)
- ScanNet: lowest ATE on the multi-room scene 54 (7.5 cm) and drift-reduced meshes compared with Point-SLAM and ESLAM (Table 3, Fig. 4)
- Map correction without storing all input frames (abstract; Sec. 3)
限制
- Not real-time; Python implementation (Limitations)
- Effect of bundle adjustment not studied (App. D)
- Place recognition could be improved with learned variants (Limitations)
- No relocalization (Limitations)
- Frame-to-model tracking alone is not the most robust choice; the authors suggest combining it with frame-to-frame tracking (Limitations)
- Behind GO-SLAM on the ScanNet average (7.7 vs 7.0 cm Avg.-9) and behind ORB-SLAM2 (1.98 cm average) on TUM-RGBD (Table 2, Table 3)
- Global registration takes about 12 s per registration with the default stopping criteria (Sec. 4.4, App. E Table 9)
- Evaluation caveats: ScanNet reference poses come from BundleFusion, and meshes are ICP-aligned to ground truth before precision and recall (Sec. 4 Datasets, App. C)
營建工程相關證據
論文未涉及營建場域;資料為 Replica、TUM-RGBD、ScanNet。
原文驗證環境:模擬、公開基準
報告的性能數據
以下是原文作者報告的性能數值(author-reported results),不是本研究重新量測的結果。每張圖只並列同一個比較組(comparison group,同一張表、同一組實驗設定)內的方法;不同比較組之間的數值不可直接比較,也不構成排名。
本方法共出現在 6 個比較組,合計 19 筆紀錄。以下列出本方法紀錄最多的 4 組,其餘 2 組列在最後,並連到性能比較頁。
Liso et al., 2024 · Table 2 本方法 6 筆
指標ATE RMSE
表格設定(擷取紀錄原文):ATE RMSE (cm) on TUM-RGBD, trajectories aligned with Horn's closed-form solution before ATE (App. C; whether scale was estimated is not stated); average of three runs unless noted; baseline numbers taken from the respective papers where available, otherwise reproduced; top block dense neural RGB-D methods, bottom block traditional dense and sparse SLAM; LC = loop closure. N/A cells (not reported) omitted. (Liso et al., 2024, Table 2)
ATE RMSE,TUM-RGBD · fr1/desk
只並列這張表在相同設定下報告的方法;以「本方法:」開頭者為本頁方法。失敗、未執行與未報告以標記呈現,不是 0。
按 Tab 進入圖表後,用上下方向鍵逐一瀏覽各類別,Esc 關閉提示框;也可開啟表格檢視閱讀全部數值。
這些是 Liso et al., 2024 在此表設定下報告的數值(author-reported results),只能在同一個比較組內對照,不代表方法在其他資料或設定下的表現。
資料來源作者報告值(Liso et al., 2024, Table 2)
| 方法(原文寫法) | 報告值 | 出處 |
|---|---|---|
| DI-Fusion [21] | 4.4 cm | (Liso et al., 2024, Table 2) |
| NICE-SLAM [77] | 4.26 cm | (Liso et al., 2024, Table 2) |
| Vox-Fusion [73] | 3.52 cm | (Liso et al., 2024, Table 2) |
| MIPS-Fusion [57] (LC) | 3 cm | (Liso et al., 2024, Table 2) |
| Point-SLAM [45] | 4.34 cm | (Liso et al., 2024, Table 2) |
| ESLAM [28] | 2.47 cm | (Liso et al., 2024, Table 2) |
| Co-SLAM [61] | 2.4 cm | (Liso et al., 2024, Table 2) |
| GO-SLAM [76] (LC) | 1.5 cm | (Liso et al., 2024, Table 2) |
| Loopy-SLAM (Ours, LC)本方法原文提出 | 3.79 cm | (Liso et al., 2024, Table 2) |
| BAD-SLAM [48] (LC) | 1.7 cm | (Liso et al., 2024, Table 2) |
| Kintinuous [69] (LC) | 3.7 cm | (Liso et al., 2024, Table 2) |
| ORB-SLAM2 [34] (LC) | 1.6 cm | (Liso et al., 2024, Table 2) |
| ElasticFusion [71] (LC) | 2.53 cm | (Liso et al., 2024, Table 2) |
| BundleFusion [13] (LC) | 1.6 cm | (Liso et al., 2024, Table 2) |
| Cao et al. [7] (LC) | 1.5 cm | (Liso et al., 2024, Table 2) |
| Yan et al. [72] (LC) | 1.6 cm | (Liso et al., 2024, Table 2) |
Liso et al., 2024 · Table 11 本方法 4 筆
資料集與序列Replica · Avg. of 8 scenes
表格設定(擷取紀錄原文):Replica mesh reconstruction averaged over 8 scenes: meshes from marching cubes, ICP-aligned to ground truth before precision and recall; precision, recall and F1 at a 1 cm threshold; depth L1 renders depth from 1000 random viewpoints on reconstructed and ground-truth meshes. GO-SLAM depth L1 3.38 is the value GO-SLAM reports with ground-truth poses; the starred 4.68 is the authors' reproduction from random poses. Truncated to scene averages. (Liso et al., 2024, Table 11)
Depth L1 on mesh at random poses,Replica · Avg. of 8 scenes
只並列這張表在相同設定下報告的方法;以「本方法:」開頭者為本頁方法。失敗、未執行與未報告以標記呈現,不是 0。
按 Tab 進入圖表後,用上下方向鍵逐一瀏覽各類別,Esc 關閉提示框;也可開啟表格檢視閱讀全部數值。
這些是 Liso et al., 2024 在此表設定下報告的數值(author-reported results),只能在同一個比較組內對照,不代表方法在其他資料或設定下的表現。
資料來源作者報告值(Liso et al., 2024, Table 11)
| 方法(原文寫法) | 報告值 | 出處 |
|---|---|---|
| NICE-SLAM [77] | 2.97 cm | (Liso et al., 2024, Table 11; Fig. 3a) |
| Vox-Fusion [73] | 2.46 cm | (Liso et al., 2024, Table 11; Fig. 3a) |
| ESLAM [28] | 1.18 cm | (Liso et al., 2024, Table 11; Fig. 3a) |
| Co-SLAM [61] | 1.51 cm | (Liso et al., 2024, Table 11; Fig. 3a) |
| GO-SLAM [76] (reproduced, random poses) | 4.68 cm | (Liso et al., 2024, Table 11; Fig. 3a) |
| Point-SLAM [45] | 0.44 cm | (Liso et al., 2024, Table 11; Fig. 3a) |
| Loopy-SLAM (Ours)本方法原文提出 | 0.35 cm | (Liso et al., 2024, Table 11; Fig. 3a) |
Liso et al., 2024 · Table 3 本方法 3 筆
指標ATE RMSE
表格設定(擷取紀錄原文):ATE RMSE (cm) on ScanNet after Horn closed-form alignment (App. C); reference poses come from BundleFusion; Avg.-6 and Avg.-9 average over 6 and 9 scenes. Truncated: only scene 54 (the only multi-room and largest scene) and the two averages are kept; '-' cells omitted. (Liso et al., 2024, Table 3)
ATE RMSE,ScanNet · Avg.-6 (scenes 00, 59, 106, 169, 181, 207)
只並列這張表在相同設定下報告的方法;以「本方法:」開頭者為本頁方法。失敗、未執行與未報告以標記呈現,不是 0。
按 Tab 進入圖表後,用上下方向鍵逐一瀏覽各類別,Esc 關閉提示框;也可開啟表格檢視閱讀全部數值。
這些是 Liso et al., 2024 在此表設定下報告的數值(author-reported results),只能在同一個比較組內對照,不代表方法在其他資料或設定下的表現。
資料來源作者報告值(Liso et al., 2024, Table 3)
| 方法(原文寫法) | 報告值 | 出處 |
|---|---|---|
| Vox-Fusion [73] | 18.5 cm | (Liso et al., 2024, Table 3) |
| Co-SLAM [61] | 8.8 cm | (Liso et al., 2024, Table 3) |
| MIPS-Fusion [57] | 10 cm | (Liso et al., 2024, Table 3) |
| NICE-SLAM [77] | 10.7 cm | (Liso et al., 2024, Table 3) |
| ESLAM [28] | 7.4 cm | (Liso et al., 2024, Table 3) |
| Point-SLAM [45] | 12.2 cm | (Liso et al., 2024, Table 3) |
| GO-SLAM [76] | 6.9 cm | (Liso et al., 2024, Table 3) |
| Loopy-SLAM (Ours)本方法原文提出 | 7.7 cm | (Liso et al., 2024, Table 3) |
Liso et al., 2024 · Table 5 本方法 3 筆
資料集與序列Replica · office 0
表格設定(擷取紀錄原文):Runtime and memory on Replica office 0. Loopy-SLAM tracking and mapping times are identical to the Point-SLAM row (the text says they are equivalent excluding loop closure). GO-SLAM reports a single 0.125 s value spanning both per-frame columns. Per-iteration columns not extracted. (Liso et al., 2024, Table 5)
Tracking/Frame,Replica · office 0
只並列這張表在相同設定下報告的方法;以「本方法:」開頭者為本頁方法。失敗、未執行與未報告以標記呈現,不是 0。
按 Tab 進入圖表後,用上下方向鍵逐一瀏覽各類別,Esc 關閉提示框;也可開啟表格檢視閱讀全部數值。
這些是 Liso et al., 2024 在此表設定下報告的數值(author-reported results),只能在同一個比較組內對照,不代表方法在其他資料或設定下的表現。
資料來源作者報告值(Liso et al., 2024, Table 5)
| 方法(原文寫法) | 報告值 | 出處 |
|---|---|---|
| NICE-SLAM [77] | 1.32 s | (Liso et al., 2024, Table 5) |
| Vox-Fusion [73] | 0.36 s | (Liso et al., 2024, Table 5) |
| Point-SLAM [45] | 0.85 s | (Liso et al., 2024, Table 5) |
| ESLAM [28] | 0.12 s | (Liso et al., 2024, Table 5) |
| Loopy-SLAM (Ours)本方法原文提出硬體:NVIDIA GPU with at most 12 GB memory (model not stated) | 0.85 s | (Liso et al., 2024, Table 5) |
其他比較組
來源
Liso et al., 2024
(2024)Loopy-SLAM: Dense Neural SLAM with Loop Closures2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 20363-20373
DOI 10.1109/cvpr52733.2024.01925arXiv 2402.09944程式碼
同儕審查已出版已讀全文近十年查證後修正
相關版本
- 預印本:arXiv:2402.09944 https://arxiv.org/abs/2402.09944
- 專案頁面:Loopy-SLAM project page repository (no source code) https://github.com/notchla/Loopy-SLAM
程式碼:https://github.com/eriksandstroem/Loopy-SLAM(授權:Apache-2.0)。有公開程式碼不等於已被重現,也不代表目前版本與論文版本相同。