[{"data":1,"prerenderedAt":89},["ShallowReactive",2],{"method-dpvslam2024":3},{"method":4,"reference":59,"equipment":80,"figures":88,"results":85},{"id":5,"label":6,"shortName":7,"title":8,"year":9,"era":10,"cluster":11,"scope":12,"keyIdeaZh":13,"keyIdeaEn":14,"fulltextStatus":15,"publicationStatus":16,"recommendation":17,"constructionRelevance":18,"validationEnvironment":19,"strengths":22,"limitations":27,"sensors":32,"platform":34,"estimator":39,"association":40,"timeModel":41,"deskew":42,"loopClosure":43,"globalOptimization":44,"mapRepresentation":45,"prior":46,"outputGeometry":47,"compute":48,"codeUrl":49,"codeLicense":50,"relatedVersions":51},"dpvslam2024","Lipson et al., 2024","DPV-SLAM","Deep Patch Visual SLAM",2024,"recent","C09","full_slam_with_global_correction","DPV-SLAM 在稀疏影像區塊（patch）視覺里程計 DPVO 上加入兩種迴圈閉合，讓深度學習式單眼 SLAM 可在單張 GPU 上以穩定的影格速率運作。近距迴圈閉合依相機位置偵測重訪，只保留舊影格的區塊特徵並建立指向近期影格的單向邊，再以自製的 CUDA 區塊稀疏光束法平差做全域最佳化；DPV-SLAM++ 另以 DBoW2 影像檢索、特徵匹配與 RANSAC 加 Umeyama 估計 Sim(3) 漂移，於 CPU 執行位姿圖最佳化以修正尺度漂移。輸出僅為相機軌跡與稀疏點。","Extends the DPVO sparse-patch learned visual odometry to monocular SLAM on one GPU with proximity-based loop edges solved by block-sparse global BA and, in DPV-SLAM++, image-retrieval loop closure with Sim(3) pose-graph optimization; output is a trajectory and sparse points.","full_text_reviewed","peer_reviewed_published","supplementary","論文未涉及營建場域；驗證資料為 KITTI、EuRoC、TUM RGB-D 與 TartanAir。它只輸出軌跡與稀疏點，無法直接產生可量測的稠密點雲；KITTI 上即使 DPV-SLAM++ 平均 ATE 仍約 25.76 m，顯示單眼尺度漂移在大範圍戶外仍是主要問題。可作為影像式重建或其他稠密建圖模組的位姿前端參考（推論）。",[20,21],"public_benchmark","simulation",[23,24,25,26],"EuRoC average ATE 0.023 m versus 0.022 m for DROID-SLAM while running at 50 versus 20 FPS with 5.0 versus 20 GB GPU memory (ECCV version Table 3; arXiv v1 printed 0.024 m)","Proximity loop closure reduces DPVO's EuRoC error from 0.105 to 0.023 m (about 4.5x, ECCV version Sec. 4) with a small speed and memory cost (Table 3)","DPV-SLAM++ reaches 25.76 m average ATE on KITTI 00-10 where DROID-SLAM fails on two sequences (Table 2(b))","Runs with a single GPU at a relatively consistent frame rate; no catastrophic failures reported across the four datasets (Sec. 1, Sec. 6)",[28,29,30,31],"Requires a GPU and provides only a sparse 3D reconstruction (Sec. 5)","Global BA cost grows quadratically with pose count, so it is limited to 1000 keyframes (Sec. 5, Appendix A)","Scale drift in outdoor environments; proximity-only loop detection fails under strong scale drift, as the KITTI results show (Sec. 4, Sec. 5)","Image retrieval can give false positives, and the classical loop closure adds about 2 GB GPU memory (Sec. 5)",[33],"monocular camera",[35,36,37,38],"vehicle (KITTI)","UAV (EuRoC MAV dataset)","handheld (TUM RGB-D freiburg1)","simulation (TartanAir test set)","DPVO recurrent update operator predicts sparse patch-flow residuals and confidences; poses and patch inverse depths solved by bundle adjustment on the patch graph; a CUDA block-sparse BA performs global optimization with loop factors; DPV-SLAM++ adds a CPU Sim(3) pose-graph optimization solved by Levenberg-Marquardt (Sec. 3)","sparse, randomly selected p x p patches tracked by learned optical flow from correlation features; loop candidates by camera proximity (DPV-SLAM) and additionally by DBoW2 ORB image retrieval with off-the-shelf keypoint matching, structure-only BA and RANSAC plus Umeyama Sim(3) alignment (DPV-SLAM++) (Sec. 3.1-3.3)","discrete poses (keyframes)","not_applicable","proximity loop closure: uni-directional long-range edges from stored patches of old frames to recent frames, followed by global BA; optional classical loop closure (DPV-SLAM++) with image retrieval requiring consecutive detections and Sim(3) drift estimation (Sec. 3.2-3.3)","global bundle adjustment over the patch graph mixed with odometry factors, limited to 1000 keyframes because cost grows quadratically; plus Sim(3) pose-graph optimization on the CPU for DPV-SLAM++ (Sec. 3.2-3.3, Sec. 5)","patch graph: sparse image patches with inverse depth attached to frames; only sparse 3D reconstruction (Sec. 3.1, Sec. 5)","network trained only on synthetic data (Sec. 2); DPV-SLAM++ uses pretrained off-the-shelf keypoint detectors and matchers during loop closure (Sec. 3.3)","camera trajectory and sparse 3D points of tracked patches (Sec. 5)","single GPU (RTX 3090 for timing); 50 FPS and 5.0 GB (DPV-SLAM) or 7.0 GB (DPV-SLAM++) on EuRoC, 39 FPS on KITTI, 27 FPS on TartanAir; classical loop closure adds about 2 GB GPU memory (Tables 1-3, Sec. 4, Sec. 5)","https:\u002F\u002Fgithub.com\u002Fprinceton-vl\u002FDPVO","MIT (LICENSE file checked)",[52,56],{"relation":53,"title":54,"doi_or_url":55},"preprint","arXiv:2408.01654v1","https:\u002F\u002Farxiv.org\u002Fabs\u002F2408.01654",{"relation":57,"title":58,"doi_or_url":49},"code_release","princeton-vl\u002FDPVO (includes DPV-SLAM)",{"id":5,"kind":60,"shortName":7,"title":8,"authors":61,"year":9,"venue":65,"venueType":66,"publisher":67,"volumeIssuePages":68,"doi":69,"arxivId":70,"url":71,"firstPublicDate":72,"publicationStatus":16,"metadataStatus":73,"fulltextStatus":15,"era":10,"classicReason":42,"codeUrl":49,"cluster":11,"topics":74,"mdpi":75,"verification":76,"label":6,"fulltextRoute":77,"versionRead":78,"addedByCensus":79},"method",[62,63,64],"Lahav Lipson","Zachary Teed","Jia Deng","Computer Vision - ECCV 2024 (Lecture Notes in Computer Science)","conference","Springer Nature Switzerland","pp. 424-440","10.1007\u002F978-3-031-72627-9_24","2408.01654","https:\u002F\u002Fapi.crossref.org\u002Fworks\u002F10.1007\u002F978-3-031-72627-9_24","2024-08-03","metadata_verified",[11],false,"corrected","arXiv","arXiv v1 (2408.01654v1, 3 Aug 2024) read in full, plus the ECCV 2024 LNCS version of record (Springer PDF via NTU institutional access) compared for the abstract, Sec. 4 text and Tables 1-4; extracted values follow the version of record where the two differ",true,[81],{"category":82,"model":83,"canonical":83,"role":84,"dataset":85,"specs":86,"locator":87},"compute","RTX-3090","compute for runtime",null,"GPU used for all timing experiments","Sec. 4",[],1790510658024]