[{"data":1,"prerenderedAt":87},["ShallowReactive",2],{"method-d3vo2020":3},{"method":4,"reference":57,"equipment":79,"figures":86,"results":46},{"id":5,"label":6,"shortName":7,"title":8,"year":9,"era":10,"cluster":11,"scope":12,"keyIdeaZh":13,"keyIdeaEn":14,"fulltextStatus":15,"publicationStatus":16,"recommendation":17,"constructionRelevance":18,"validationEnvironment":19,"strengths":21,"limitations":26,"sensors":31,"platform":33,"estimator":36,"association":37,"timeModel":38,"deskew":39,"loopClosure":40,"globalOptimization":41,"mapRepresentation":42,"prior":43,"outputGeometry":44,"compute":45,"codeUrl":46,"codeLicense":47,"relatedVersions":48},"d3vo2020","Yang et al., 2020a","D3VO","D3VO: Deep Depth, Deep Pose and Deep Uncertainty for Monocular Visual Odometry",2020,"recent","C09","odometry_with_local_mapping","D3VO 在直接稀疏里程計（DSO）中三個層次加入深度網路：自監督的 DepthNet 預測深度使新點一開始就有公制尺度，並形成虛擬立體項；網路同時預測光度不確定度，用來取代傳統的殘差權重；PoseNet 預測的相對位姿則在前端追蹤當作先驗因子，在後端光度平差中當作位姿能量項。網路只以立體影片自監督訓練，並預測仿射亮度參數以處理曝光變化。系統仍是無迴圈閉合的單眼里程計，輸出軌跡與稀疏點雲。","Monocular direct sparse odometry (DSO) boosted by a self-supervised network that supplies metric depth (with a virtual stereo term), per-pixel photometric uncertainty as residual weights, and relative-pose priors for tracking and bundle adjustment; no loop closure.","full_text_reviewed","peer_reviewed_published","background","論文未涉及營建場域；只在 KITTI 車載與 EuRoC 飛行器資料上評估，且深度網路以同資料集的立體影片訓練，換到工地場景是否仍能提供可靠公制尺度未經驗證（推論）。系統只輸出稀疏點與軌跡，沒有迴圈閉合，主要價值在於示範學習式深度、不確定度與位姿先驗如何整合進直接法里程計。",[20],"public_benchmark",[22,23,24,25],"Mean t_rel 0.82% on the paper's KITTI test split, below stereo DSO (0.89) and stereo ORB-SLAM2 (0.91) (Table 4)","EuRoC mean ATE 0.10 with a single camera, comparable to monocular VIO such as VI-DSO (0.11) (Table 6)","Deep pose integration strongly improves the hardest EuRoC sequences (V1_03 from 0.63 to 0.13 when adding Dp to Dd) (Table 6)","Self-supervised depth network with brightness parameters and uncertainty outperforms Monodepth2 on KITTI and EuRoC depth (Tables 1-2)",[27,28,29,30],"Networks are trained on sequences of the same datasets used for testing; cross-scene generalization of monocular depth remains a challenge (Sec. 4.1, Tables 2-3)","Without SE(3) alignment, ATE on KITTI can be large (for example 26.9 m on sequence 01) because early pose errors propagate (Supp. E, Supp. Table 4)","No loop closure or global optimization (Sec. 3.2) (inference from the method description)","Runtime of the VO system is not reported (inference)",[32],"monocular camera at run time (stereo videos only for self-supervised network training)",[34,35],"vehicle (KITTI)","UAV (EuRoC MAV)","DSO-style windowed sparse photometric bundle adjustment (Gauss-Newton) with three learned inputs: DepthNet depth initializes points at metric scale and adds a virtual stereo term, learned photometric uncertainty sets residual weights, and PoseNet relative poses act as front-end tracking prior factors and as a pose energy term in the back end (Sec. 3.2, Supp. C)","direct photometric residuals on sparse points with an 8-pixel neighbourhood pattern and Huber norm, weighted by the learned uncertainty (Sec. 3.2)","discrete poses (keyframes)","not_applicable","none","none; windowed photometric bundle adjustment with marginalization as in DSO (Sec. 3.2)","sparse point set hosted in keyframes (DSO) with depths initialized from the network (Sec. 3.2, Fig. 2)","self-supervised DepthNet (ResNet-18 encoder, ImageNet initialization) and PoseNet trained on stereo videos of the same datasets (KITTI Eigen split; EuRoC sequences not used for testing), predicting depth, photometric uncertainty, relative pose and affine brightness parameters (Sec. 3.1, Sec. 4, Supp. A-B)","camera trajectory and sparse point cloud (Fig. 2)","networks trained with PyTorch on a single Titan X Pascal GPU at 512x256; VO runtime not reported (Supp. A)",null,"not_applicable (no official code release found)",[49,53],{"relation":50,"title":51,"doi_or_url":52},"preprint","arXiv:2003.01060v2","https:\u002F\u002Farxiv.org\u002Fabs\u002F2003.01060",{"relation":54,"title":55,"doi_or_url":56},"supplementary","CVPR 2020 supplementary material (CVF open access)","https:\u002F\u002Fopenaccess.thecvf.com\u002Fcontent_CVPR_2020\u002Fsupplemental\u002FYang_D3VO_Deep_Depth_CVPR_2020_supplemental.pdf",{"id":5,"kind":58,"shortName":7,"title":8,"authors":59,"year":9,"venue":64,"venueType":65,"publisher":66,"volumeIssuePages":67,"doi":68,"arxivId":69,"url":70,"firstPublicDate":71,"publicationStatus":16,"metadataStatus":72,"fulltextStatus":15,"era":10,"classicReason":39,"codeUrl":46,"cluster":11,"topics":73,"mdpi":74,"verification":75,"label":6,"fulltextRoute":76,"versionRead":77,"addedByCensus":78},"method",[60,61,62,63],"Nan Yang","Lukas von Stumberg","Rui Wang","Daniel Cremers","2020 IEEE\u002FCVF Conference on Computer Vision and Pattern Recognition (CVPR)","conference","IEEE","pp. 1278-1289","10.1109\u002Fcvpr42600.2020.00136","2003.01060","https:\u002F\u002Fapi.crossref.org\u002Fworks\u002F10.1109\u002FCVPR42600.2020.00136","2020-03-02","metadata_verified",[11],false,"corrected","arXiv","arXiv v2 (2003.01060v2, 28 Mar 2020, labelled CVPR 2020) plus the CVPR 2020 supplementary material from CVF open access; cross-checked against the CVPR 2020 CVF open-access paper and supplementary (all extracted values match)",true,[80],{"category":81,"model":82,"canonical":82,"role":83,"dataset":46,"specs":84,"locator":85},"compute","Titan X Pascal","compute for runtime","single GPU used to train DepthNet and PoseNet; VO runtime not reported","Supp. A",[],1790510662261]