Extends DROID-SLAM with flow-based loop closing and online full bundle adjustment, and continuously re-fits a hash-encoded neural SDF to the globally optimized keyframe poses and depths, giving globally consistent meshes from monocular, stereo or RGB-D video.

技術屬性

欄位內容為文獻擷取紀錄的原文用語(英文),以原文為據;「未查證」表示本研究尚未讀到該資訊,不代表該方法不具備此能力。

GO-SLAM 的技術屬性
感測輸入monocular camera、stereo camera、RGB-D camera
原文測試平台UAV (EuRoC MAV dataset)、simulation (Replica synthetic scenes)、原文未報告 (capture platforms of TUM RGB-D, ETH3D-SLAM and ScanNet are not described in the paper)
狀態估計DROID-SLAM tracking extended with online global optimization: RAFT-based recurrent update operator predicts dense flow and confidence, and a differentiable dense bundle adjustment (DBA) layer solves poses and per-pixel inverse depths by damped Gauss-Newton; front end optimizes a local keyframe window with loop edges, back end runs online full BA over all keyframes in a separate thread (Sec. 3.1)
資料關聯dense learned optical flow with per-pixel confidence between keyframe pairs; keyframe-graph edges chosen by co-visibility measured as mean rigid flow (threshold tau_co = 25) with neighbourhood suppression (Sec. 3.1, Sec. 4.1)
時間表示discrete poses (keyframes)
去畸變不適用
迴圈閉合flow-based: edges sampled from the unexplored part of the co-visibility matrix between local-window and historical keyframes; a loop is accepted after three consecutive candidates with mean flow below tau_co, then optimized by DBA (Sec. 3.1)
全域最佳化online full bundle adjustment over all keyframes in a back-end thread, running concurrently with tracking and loop closing (Sec. 3.1)
地圖表示neural implicit SDF and colour field with multi-resolution hash encoding (16 levels, Instant-NGP style) and shallow MLPs, rendered by NeuS-style unbiased volume rendering; mapping re-trains on selected keyframes (latest two, top-10 by pose change, 10 stratified) (Sec. 3.2, Supp. A)
先驗資訊pretrained DROID-SLAM weights for tracking; rendering networks trained from scratch per scene (Sec. 4.1)
可輸出幾何keyframe trajectory, per-keyframe depth and a triangle mesh extracted by marching cubes from the SDF (Sec. 4.1)
計算需求Intel Core i9-10920X 3.5 GHz CPU and NVIDIA RTX 3090 GPU; about 8 FPS on Replica RGB-D with 15.63 GB GPU memory in Table 9, while Sec. 4.4 states a maximum of 18 GB (Sec. 4.1, Sec. 4.4, Table 9)

使用設備

原文使用的感測器、運算硬體與載具(equipment)。型號保留原文寫法,連結到設備頁中同一型號的歸併名稱;角色依原文用途分為方法輸入、資料集感測器、執行運算平台、參考或真值量測(reference or ground truth)與比較對象設備。

原文使用的設備
類別型號(原文寫法)角色資料集原文規格出處
運算硬體Intel Core i9-10920X執行運算平台未標示3.5 GHz CPU(Zhang et al., 2023b, Sec. 4.1)
運算硬體NVIDIA RTX 3090執行運算平台未標示GPU; 15.63 GB peak memory used on Replica RGB-D (Table 9)(Zhang et al., 2023b, Sec. 4.1; Table 9)

作者報告的優勢與限制

優勢

限制

營建工程相關證據

論文未涉及營建場域;驗證資料為 TUM RGB-D、EuRoC、ETH3D-SLAM、ScanNet 與 Replica 等室內或飛行器資料集,幾何精度只在 Replica 合成場景以公分級 Accuracy 與 Completion 評估,未使用全測站或 TLS 等獨立參考量測。其長序列迴圈閉合與全域平差可說明神經隱式地圖要維持全域一致須仰賴傳統 SLAM 式的後端,但單眼 ScanNet 部分場景仍有數十公分的 ATE(Table 3),距營建量測所需精度仍有差距(推論)。

原文驗證環境:公開基準、模擬

報告的性能數據

性能數據仍在分批查證,目前尚未收錄此方法的報告值。

來源

  • Zhang et al., 2023b

    Youmin Zhang, Fabio Tosi, Stefano Mattoccia, Matteo Poggi(2023)GO-SLAM: Global Optimization for Consistent 3D Instant Reconstruction2023 IEEE/CVF International Conference on Computer Vision (ICCV), pp. 3704-3714

    同儕審查已出版已讀全文近十年查證後修正

回到方法圖鑑

選擇開啟Esc關閉