Real-time monocular dense SLAM that optimizes key-frame poses and learned 32-D depth codes in a GTSAM/iSAM2 factor graph with photometric, reprojection and geometric factors, plus local and bag-of-words global loop closure.

技術屬性

欄位內容為文獻擷取紀錄的原文用語(英文),以原文為據;「未查證」表示本研究尚未讀到該資訊,不代表該方法不具備此能力。

DeepFactors 的技術屬性
感測輸入monocular camera
原文測試平台原文未報告 (ScanNet, ICL-NUIM and TUM RGB-D sequences and a live camera; capture platforms not described)
狀態估計batch MAP factor graph over key-frame poses and compact depth codes solved incrementally with iSAM2 in GTSAM; camera tracking by GPU direct whole-image SE(3) Lucas-Kanade against the closest key-frame; tracking and mapping interleaved; one-way frames refine the latest key-frame and are then marginalized (Sec. V)
資料關聯three pairwise factor types: dense photometric error, BRISK keypoint reprojection error with Cauchy cost, and sparse geometric depth-consistency error with Huber cost (Sec. III)
時間表示discrete poses (key-frames)
去畸變不適用
迴圈閉合local loops by a pose-based criterion within the last 10 key-frames; global loops by bag-of-words (DBoW2) candidates verified by tracking inliers and pose distance, closed with reprojection factors only (Sec. V-D)
全域最佳化incremental batch optimization of all key-frames (iSAM2) with zero-code prior factors (Sec. V-B)
地圖表示key-frame depth maps decoded linearly from a 32-dimensional learned code conditioned on the image (CodeSLAM-style variational auto-encoder) (Sec. III-IV, Sec. VI-A)
先驗資訊network trained on about 1.4M ScanNet images with merged rendered and sensor depth, image size 256x192, code size 32; an explicit code-prediction path gives the initial depth of each key-frame (Sec. IV, Sec. VI-A)
可輸出幾何key-frame poses and dense key-frame depth maps fused for visualization as reconstructions (Figs. 1 and 7)
計算需求single NVIDIA GTX 1080 at 256x192; about 340 ms per new key-frame for network and Jacobian (16 ms forward pass), tracking at about 250 Hz; overall speed depends on connectivity, enabled factors and loop closures (Sec. VII)

使用設備

原文使用的感測器、運算硬體與載具(equipment)。型號保留原文寫法,連結到設備頁中同一型號的歸併名稱;角色依原文用途分為方法輸入、資料集感測器、執行運算平台、參考或真值量測(reference or ground truth)與比較對象設備。

原文使用的設備
類別型號(原文寫法)角色資料集原文規格出處
運算硬體NVIDIA GTX 1080執行運算平台未標示single GPU running the network, CUDA kernels and visualization at 256x192(Czarnowski et al., 2020, Sec. VII)

作者報告的優勢與限制

優勢

限制

營建工程相關證據

論文未涉及營建場域;評估使用 ScanNet、ICL-NUIM 與 TUM RGB-D 室內資料,深度精度以落在真值 10% 以內的像素比例衡量,平均約 27%,且需以最佳尺度校正單眼結果。其價值在於示範學習式深度先驗可與因子圖、iSAM2 與迴圈閉合等傳統 SLAM 後端結合(推論),但解析度與精度都不足以直接產生施工量測用點雲。

原文驗證環境:公開基準、模擬

報告的性能數據

性能數據仍在分批查證,目前尚未收錄此方法的報告值。

來源

  • Czarnowski et al., 2020

    Jan Czarnowski, Tristan Laidlow, Ronald Clark, Andrew J. Davison(2020)DeepFactors: Real-Time Probabilistic Dense Monocular SLAMIEEE Robotics and Automation Letters, 5(2):721-728

    同儕審查已出版已讀全文近十年查證後修正

回到方法圖鑑

選擇開啟Esc關閉