Monocular direct sparse odometry (DSO) boosted by a self-supervised network that supplies metric depth (with a virtual stereo term), per-pixel photometric uncertainty as residual weights, and relative-pose priors for tracking and bundle adjustment; no loop closure.

技術屬性

欄位內容為文獻擷取紀錄的原文用語(英文),以原文為據;「未查證」表示本研究尚未讀到該資訊,不代表該方法不具備此能力。

D3VO 的技術屬性
感測輸入monocular camera at run time (stereo videos only for self-supervised network training)
原文測試平台vehicle (KITTI)、UAV (EuRoC MAV)
狀態估計DSO-style windowed sparse photometric bundle adjustment (Gauss-Newton) with three learned inputs: DepthNet depth initializes points at metric scale and adds a virtual stereo term, learned photometric uncertainty sets residual weights, and PoseNet relative poses act as front-end tracking prior factors and as a pose energy term in the back end (Sec. 3.2, Supp. C)
資料關聯direct photometric residuals on sparse points with an 8-pixel neighbourhood pattern and Huber norm, weighted by the learned uncertainty (Sec. 3.2)
時間表示discrete poses (keyframes)
去畸變不適用
迴圈閉合none
全域最佳化none; windowed photometric bundle adjustment with marginalization as in DSO (Sec. 3.2)
地圖表示sparse point set hosted in keyframes (DSO) with depths initialized from the network (Sec. 3.2, Fig. 2)
先驗資訊self-supervised DepthNet (ResNet-18 encoder, ImageNet initialization) and PoseNet trained on stereo videos of the same datasets (KITTI Eigen split; EuRoC sequences not used for testing), predicting depth, photometric uncertainty, relative pose and affine brightness parameters (Sec. 3.1, Sec. 4, Supp. A-B)
可輸出幾何camera trajectory and sparse point cloud (Fig. 2)
計算需求networks trained with PyTorch on a single Titan X Pascal GPU at 512x256; VO runtime not reported (Supp. A)

使用設備

原文使用的感測器、運算硬體與載具(equipment)。型號保留原文寫法,連結到設備頁中同一型號的歸併名稱;角色依原文用途分為方法輸入、資料集感測器、執行運算平台、參考或真值量測(reference or ground truth)與比較對象設備。

原文使用的設備
類別型號(原文寫法)角色資料集原文規格出處
運算硬體Titan X Pascal執行運算平台未標示single GPU used to train DepthNet and PoseNet; VO runtime not reported(Yang et al., 2020a, Supp. A)

作者報告的優勢與限制

優勢

限制

營建工程相關證據

論文未涉及營建場域;只在 KITTI 車載與 EuRoC 飛行器資料上評估,且深度網路以同資料集的立體影片訓練,換到工地場景是否仍能提供可靠公制尺度未經驗證(推論)。系統只輸出稀疏點與軌跡,沒有迴圈閉合,主要價值在於示範學習式深度、不確定度與位姿先驗如何整合進直接法里程計。

原文驗證環境:公開基準

報告的性能數據

性能數據仍在分批查證,目前尚未收錄此方法的報告值。

來源

  • Yang et al., 2020a

    Nan Yang, Lukas von Stumberg, Rui Wang, Daniel Cremers(2020)D3VO: Deep Depth, Deep Pose and Deep Uncertainty for Monocular Visual Odometry2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 1278-1289

    同儕審查已出版已讀全文近十年查證後修正

回到方法圖鑑

選擇開啟Esc關閉