RGB-only neural implicit SLAM that tracks and maps with one hierarchical SDF and colour grid, disambiguated by monocular depth and normal priors, optical flow and a warping loss, with a locally adaptive SDF-to-density transform; no loop closure and not real-time.

技術屬性

欄位內容為文獻擷取紀錄的原文用語(英文),以原文為據;「未查證」表示本研究尚未讀到該資訊,不代表該方法不具備此能力。

NICER-SLAM 的技術屬性
感測輸入monocular RGB camera
原文測試平台simulation (Replica)、原文未報告 (7-Scenes and the self-captured outdoor Azure Kinect dataset; carrier not described)
狀態估計end-to-end optimization through differentiable volume rendering: tracking optimizes the current pose with an RGB rendering loss (100 iterations, 1024 pixels) with the map fixed; mapping runs a 3-stage optimization with RGB, warping, optical-flow, monocular depth, monocular normal and Eikonal losses, ending with local bundle adjustment over 16 selected frames of which half are frozen (Sec. 3.3; iteration counts from arXiv v1 Sec. 3.4)
資料關聯direct photometric rendering loss plus dense correspondence cues: RGB warping between keyframes and optical flow from GMFlow (Sec. 3.3)
時間表示discrete poses
去畸變不適用
迴圈閉合none (stated as a limitation, Sec. 5)
全域最佳化none; local BA over selected mapping frames only (Sec. 3.3-3.4)
地圖表示hierarchical neural implicit SDF: coarse 32^3 dense feature grid plus 8-level fine residual grids (32-128) and a 16-level colour grid (16-2048) with small MLP decoders; VolSDF-style SDF-to-density with a locally adaptive beta from per-voxel sample counts (Sec. 3.1-3.2)
先驗資訊monocular depth and normal predictions from an off-the-shelf predictor (Omnidata in arXiv v1), optical flow from GMFlow; COLMAP used to obtain intrinsics for 7-Scenes and the self-captured outdoor dataset (Sec. 3.3, Sec. 4)
可輸出幾何camera trajectory and a triangle mesh extracted by marching cubes at 512^3; rendered novel views (Sec. 3.4)
計算需求not reported in the 3DV main text; arXiv v1 Sec. 3.4 reports a single NVIDIA A100 with on average 496 ms per mapping iteration and 147 ms per tracking iteration (100 iterations each); not real-time (Sec. 5)

使用設備

原文使用的感測器、運算硬體與載具(equipment)。型號保留原文寫法,連結到設備頁中同一型號的歸併名稱;角色依原文用途分為方法輸入、資料集感測器、執行運算平台、參考或真值量測(reference or ground truth)與比較對象設備。

原文使用的設備
類別型號(原文寫法)角色資料集原文規格出處
RGB-D 相機Azure Kinect方法輸入SCO (self-captured outdoor)used to capture the self-captured outdoor (SCO) dataset of 6 scenes with 800 to 2700 frames; only RGB images are input, the depth is shown for visualization and is unreliable outdoors(Zhu et al., 2024, Sec. 4 Datasets; Sec. 4.1; Fig. 6)
運算硬體A100執行運算平台未標示single GPU; 496 ms per mapping iteration, 147 ms per tracking iteration(Zhu et al., 2024, arXiv v1 Sec. 3.4 (implementation details are not in the 3DV main text))

作者報告的優勢與限制

優勢

限制

營建工程相關證據

論文未涉及營建場域;定量評估只在合成 Replica 資料,3DV 版另以 Azure Kinect 自行拍攝 6 個戶外場景,但只做定性比較,且作者指出戶外深度量測不可靠。其以單眼深度與法向量先驗補足 RGB 幾何的做法,說明僅靠影像取得公分級網格仍依賴學習先驗,且未做迴圈閉合、速度遠非即時;對以一般相機記錄工地的情境可作為上限參考,但無法支持施工驗收等級的幾何主張(推論)。

原文驗證環境:公開基準、模擬

報告的性能數據

性能數據仍在分批查證,目前尚未收錄此方法的報告值。

來源

  • Zhu et al., 2024

    Zihan Zhu, Songyou Peng, Viktor Larsson, Zhaopeng Cui, Martin R. Oswald, Andreas Geiger, Marc Pollefeys(2024)NICER-SLAM: Neural Implicit Scene Encoding for RGB SLAM2024 International Conference on 3D Vision (3DV), pp. 42-52

    同儕審查已出版已讀全文近十年查證後修正

回到方法圖鑑

選擇開啟Esc關閉