Hierarchical feature grids with pre-trained decoders enable local, scalable neural implicit RGB-D SLAM, without loop closure.

技術屬性

欄位內容為文獻擷取紀錄的原文用語(英文),以原文為據;「未查證」表示本研究尚未讀到該資訊,不代表該方法不具備此能力。

NICE-SLAM 的技術屬性
感測輸入RGB-D
原文測試平台未記錄
狀態估計alternating gradient-based optimization in parallel threads: staged mapping (mid-level grid, then mid and fine grids with the depth L1 loss) followed by a local bundle adjustment that jointly optimizes all feature grids, the colour decoder and the poses of K selected keyframes (Eq. 10); tracking optimizes only the current camera pose with a variance-weighted depth loss plus a photometric loss (Eq. 11 and 12)
資料關聯direct depth (L1) and photometric re-rendering losses
時間表示discrete poses
去畸變不適用
迴圈閉合none (authors list loop closure as future work)
全域最佳化none
地圖表示hierarchical coarse/mid/fine feature grids with pre-trained occupancy decoders plus a colour grid
先驗資訊coarse, mid and fine occupancy decoders pre-trained as part of ConvONet on its Synthetic Indoor Scene Dataset (room_grid64 setting, point-cloud encoder) and kept fixed during SLAM; the colour decoder is optimized online
可輸出幾何mesh via marching cubes: fine-level decoder occupancy for observed points; for unseen points inside partially observed coarse voxels the coarse decoder predicts occupancy (shown in cyan); other points set to zero occupancy
計算需求desktop PC with a 3.80 GHz Intel i7-10700K CPU and an NVIDIA RTX 3090 GPU (Sec. 4.1); Table 4 reports 47 ms tracking and 130 ms mapping at Mt = 200 and M = 1000 pixel samples (per iteration or per frame not stated) and 104.16 x10^3 FLOPs per point query; map memory 12.02 MB on Replica (Table 1). Independent measurements: (Sandström et al., 2023) Table 6 reports 1.32 s tracking and 10.92 s mapping per frame on Replica office 0 (RTX 2080 Ti); (Huang et al., 2024c) Table 1 reports 2.331 tracking FPS and more than 10 min operation time on Replica RGB-D (RTX 4090)

使用設備

原文使用的感測器、運算硬體與載具(equipment)。型號保留原文寫法,連結到設備頁中同一型號的歸併名稱;角色依原文用途分為方法輸入、資料集感測器、執行運算平台、參考或真值量測(reference or ground truth)與比較對象設備。

原文使用的設備
類別型號(原文寫法)角色資料集原文規格出處
RGB-D 相機原文未報告方法輸入self-captured multi-room apartmentsensor of the self-captured sequence in a large multi-room apartment is not named(Zhu et al., 2022a, Sec. 4.1; Sec. 4.2 Evaluation on a Larger Scene)
運算硬體Intel i7-10700K CPU (3.80 GHz)執行運算平台未標示desktop PC used for all NICE-SLAM runs(Zhu et al., 2022a, Sec. 4.1)
運算硬體NVIDIA RTX 3090執行運算平台未標示GPU of the desktop PC used for all NICE-SLAM runs(Zhu et al., 2022a, Sec. 4.1)

作者報告的優勢與限制

優勢

限制

營建工程相關證據

論文未涉及營建場域;資料為 Replica、ScanNet、TUM RGB-D、Co-Fusion 與自錄公寓。

原文驗證環境:模擬、公開基準、受控實驗

報告的性能數據

以下是原文作者報告的性能數值(author-reported results),不是本研究重新量測的結果。每張圖只並列同一個比較組(comparison group,同一張表、同一組實驗設定)內的方法;不同比較組之間的數值不可直接比較,也不構成排名。

本方法共出現在 73 個比較組,合計 429 筆紀錄。以下列出本方法紀錄最多的 4 組,其餘 69 組列在最後,並連到性能比較頁。

Rosinol et al., 2023 · Table I 本方法 36 筆

表格設定(擷取紀錄原文):Replica rendered sequences (2000 frames per scene from iMAP); iMAP* and NICE-SLAM (GT depth) use rendered ground-truth depth; TSDF-Fusion, sigma-Fusion and NeRF-SLAM use DROID-SLAM poses and depths; Depth L1 is a proxy for geometric accuracy; values from the IROS 2023 version of record (the NeRF-SLAM row differs from arXiv v1) (Rosinol et al., 2023, Table I)

Depth L1 [cm],Replica · room-0

只並列這張表在相同設定下報告的方法;以「本方法:」開頭者為本頁方法。失敗、未執行與未報告以標記呈現,不是 0。

按 Tab 進入圖表後,用上下方向鍵逐一瀏覽各類別,Esc 關閉提示框;也可開啟表格檢視閱讀全部數值。

這些是 Rosinol et al., 2023 在此表設定下報告的數值(author-reported results),只能在同一個比較組內對照,不代表方法在其他資料或設定下的表現。

統計量:原文未報告;對齊方式:不適用;單位:cm;場景:synthetic renders of scanned indoor scenes

資料來源作者報告值(Rosinol et al., 2023, Table I)

數值與出處
方法(原文寫法)報告值出處
iMAP* [27] (GT depth)5.7 cm(Rosinol et al., 2023, Table I)
Nice-SLAM [28] (GT depth)本方法2.53 cm(Rosinol et al., 2023, Table I)
TSDF-Fusion Res. = 256 (our depth)23.51 cm(Rosinol et al., 2023, Table I)
sigma-Fusion [15] Res. = 256 (our depth)21.92 cm(Rosinol et al., 2023, Table I)
Nice-SLAM [28] (no depth)本方法11.12 cm(Rosinol et al., 2023, Table I)
Ours (our depth)原文提出8.11 cm(Rosinol et al., 2023, Table I)

Yang et al., 2022 · Table 2 本方法 27 筆

表格設定(擷取紀錄原文):Replica mesh reconstruction; iMAP values from its paper, NICE-SLAM values from its supplementary without mesh culling (Yang et al., 2022, Table 2)

Acc. [cm],Replica · Room-0

只並列這張表在相同設定下報告的方法;以「本方法:」開頭者為本頁方法。失敗、未執行與未報告以標記呈現,不是 0。

按 Tab 進入圖表後,用上下方向鍵逐一瀏覽各類別,Esc 關閉提示框;也可開啟表格檢視閱讀全部數值。

這些是 Yang et al., 2022 在此表設定下報告的數值(author-reported results),只能在同一個比較組內對照,不代表方法在其他資料或設定下的表現。

統計量:平均值(mean);對齊方式:不適用;單位:cm;場景:synthetic indoor scenes

資料來源作者報告值(Yang et al., 2022, Table 2)

數值與出處
方法(原文寫法)報告值出處
iMap [ 31 ]3.58 cm(Yang et al., 2022, Table 2)
NICE-SLAM [ 40 ]本方法3.53 cm(Yang et al., 2022, Table 2)
Ours原文提出2.41 cm(Yang et al., 2022, Table 2)

Keetha et al., 2024 · Table 1 本方法 22 筆

指標ATE RMSE [cm]

表格設定(擷取紀錄原文):Online camera-pose estimation, ATE RMSE [cm]; Baseline numbers taken from Point-SLAM; SplaTAM averaged over 3 seeds (Keetha et al., 2024, Table 1)

ATE RMSE [cm],TUM-RGBD · Avg.

只並列這張表在相同設定下報告的方法;以「本方法:」開頭者為本頁方法。失敗、未執行與未報告以標記呈現,不是 0。

按 Tab 進入圖表後,用上下方向鍵逐一瀏覽各類別,Esc 關閉提示框;也可開啟表格檢視閱讀全部數值。

這些是 Keetha et al., 2024 在此表設定下報告的數值(author-reported results),只能在同一個比較組內對照,不代表方法在其他資料或設定下的表現。

統計量:均方根誤差(RMSE);對齊方式:原文未報告;單位:cm;場景:real RGB-D sequences from old low-quality cameras (sparse depth, strong motion blur)

資料來源作者報告值(Keetha et al., 2024, Table 1)

數值與出處
方法(原文寫法)報告值出處
Vox-Fusion11.31 cm(Keetha et al., 2024, Table 1)
NICE-SLAM本方法15.87 cm(Keetha et al., 2024, Table 1)
Point-SLAM8.92 cm(Keetha et al., 2024, Table 1)
SplaTAM原文提出5.48 cm(Keetha et al., 2024, Table 1)
Kintinuous4.84 cm(Keetha et al., 2024, Table 1)
ElasticFusion6.91 cm(Keetha et al., 2024, Table 1)
ORB-SLAM21.98 cm(Keetha et al., 2024, Table 1)

Isaacson et al., 2023 · Table III 本方法 16 筆

表格設定(擷取紀錄原文):Map accuracy and completion (m, mean nearest-point distances) and precision and recall at a 0.1 m threshold; meshes from each method sampled to point clouds, all clouds voxel-downsampled to 5 cm (1 cm for MCR), ground truth cropped to geometry observed by the sensor. SHINE Mapping used ground-truth poses. '-' = invalid configuration (accuracy and precision are '-' for every method on Canteen and Garden; NICE-SLAM '-' on Quad), x = failed (NICE-SLAM on Canteen and Garden, all metrics). (Isaacson et al., 2023, Table III)

Accuracy (mean distance estimated to ground truth),Fusion Portable · MCR Slow 01

只並列這張表在相同設定下報告的方法;以「本方法:」開頭者為本頁方法。失敗、未執行與未報告以標記呈現,不是 0。

按 Tab 進入圖表後,用上下方向鍵逐一瀏覽各類別,Esc 關閉提示框;也可開啟表格檢視閱讀全部數值。

這些是 Isaacson et al., 2023 在此表設定下報告的數值(author-reported results),只能在同一個比較組內對照,不代表方法在其他資料或設定下的表現。

統計量:平均值(mean);對齊方式:原文未報告;單位:m;場景:indoor lab, quadruped

資料來源作者報告值(Isaacson et al., 2023, Table III)

數值與出處
方法(原文寫法)報告值出處
NICE-SLAM本方法0.621 m(Isaacson et al., 2023, Table III)
SHINE (ground-truth poses)0.164 m(Isaacson et al., 2023, Table III)
LONER w./ L_CLONeR0.11 m(Isaacson et al., 2023, Table III)
LONER w./ L_URF0.153 m(Isaacson et al., 2023, Table III)
LONER原文提出0.186 m(Isaacson et al., 2023, Table III)

其他比較組

列出其餘 69 個比較組

來源

  • Zhu et al., 2022a

    Zihan Zhu, Songyou Peng, Viktor Larsson, Weiwei Xu, Hujun Bao, Zhaopeng Cui, Martin R. Oswald, Marc Pollefeys(2022)NICE-SLAM: Neural Implicit Scalable Encoding for SLAM2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 12776-12786

    同儕審查已出版已讀全文近十年

回到方法圖鑑

選擇開啟Esc關閉