[{"data":1,"prerenderedAt":93},["ShallowReactive",2],{"method-dpvo2023":3},{"method":4,"reference":63,"equipment":84,"figures":92,"results":89},{"id":5,"label":6,"shortName":7,"title":8,"year":9,"era":10,"cluster":11,"scope":12,"keyIdeaZh":13,"keyIdeaEn":14,"fulltextStatus":15,"publicationStatus":16,"recommendation":17,"constructionRelevance":18,"validationEnvironment":19,"strengths":22,"limitations":27,"sensors":33,"platform":35,"estimator":39,"association":40,"timeModel":41,"deskew":42,"loopClosure":43,"globalOptimization":44,"mapRepresentation":45,"prior":46,"outputGeometry":47,"compute":48,"codeUrl":49,"codeLicense":50,"relatedVersions":51},"dpvo2023","Teed et al., 2023","DPVO","Deep Patch Visual Odometry",2023,"recent","C09","odometry","DPVO 是深度學習式單眼視覺里程計，把 DROID-SLAM 的稠密光流改為只追蹤稀疏影像區塊（patch）。每張影格隨機取樣區塊，循環更新網路依相關特徵、時間向卷積與訊息傳遞預測區塊軌跡修正量與信心權重，再由可微分光束法平差在滑動視窗內更新位姿與區塊逆深度。整個網路只以合成資料 TartanAir 訓練，在 RTX-3090 上平均每秒 60 影格、約 4.9 GB 記憶體，但不含迴圈閉合，輸出僅為軌跡與稀疏點。","Learned monocular visual odometry that tracks sparse random image patches with a recurrent update operator and differentiable bundle adjustment in a sliding window, matching or beating dense-flow DROID-VO at a fraction of the memory; no loop closure and only sparse points.","full_text_reviewed","peer_reviewed_published","background","論文未涉及營建場域；評估資料為 TartanAir、EuRoC、TUM RGB-D 與 ICL-NUIM，均未含施工現場。DPVO 只輸出軌跡與稀疏點，且無迴圈閉合，無法單獨產生可量測的點雲；其價值在於作為低記憶體的學習式位姿前端，已被 DPV-SLAM 等後續系統沿用（推論）。",[20,21],"public_benchmark","simulation",[23,24,25,26],"EuRoC average ATE 0.105 m versus 0.186 m for DROID-VO (Table 2)","TartanAir test average ATE 0.21 versus 0.33 for DROID-SLAM with global optimization (Table 1)","Averages 60 FPS with 4.9 GB versus 8.7 GB for DROID-VO, and the frame rate stays nearly constant regardless of motion (Sec. 1, Figs. 8 and 10)","Learning-based and did not fail on any TUM RGB-D fr1 sequence, unlike ORB-SLAM3 and DSO (Table 3)",[28,29,30,31,32],"Odometry only: no loop closure or global correction, so drift accumulates (Sec. 2; DPV-SLAM Sec. 1)","Only sparse 3D points are produced (Fig. 1, Appendix E) (inference for dense-mapping use)","Treating all frames as keyframes is slower than necessary for slow or still camera motion (Appendix B)","Initialization needs camera motion of at least 8 pixels average flow (Sec. 3.3)","Reported by the DPV-SLAM authors: DPVO suffers the same performance issues as DROID-SLAM on outdoor data (DPV-SLAM Sec. 2)",[34],"monocular camera",[36,37,38],"UAV (EuRoC MAV dataset)","simulation (TartanAir, ICL-NUIM)","not_reported (TUM RGB-D platform not described; the paper notes erratic motion and blur)","recurrent update operator (correlation, 1D temporal convolution, softmax aggregation, transition block, factor head) predicts 2D patch-trajectory revisions and confidences; a differentiable bundle adjustment layer applies two Gauss-Newton iterations with the Schur complement to camera poses and patch inverse depths; poses of all but the last 10 keyframes are fixed (Sec. 3.1, Sec. 3.3, Appendix F)","sparse image patches at random locations (96 per frame by default, 48 in the fast setting) tracked by learned correlation features against frames within distance r in a bipartite patch graph (Sec. 3, Sec. 4)","discrete poses (keyframes; relative poses stored for removed keyframes)","not_applicable","none (visual odometry only; DPV-SLAM later adds loop closure)","none; sliding-window optimization over the last 10 keyframes (Sec. 3.3)","patch graph of sparse fronto-parallel patches with inverse depth; sparse 3D reconstruction (Sec. 3, Fig. 1, Appendix E)","network trained entirely on synthetic TartanAir data (Sec. 3.2, Appendix D)","camera trajectory and sparse 3D points of tracked patches (Fig. 1, Fig. B)","RTX-3090: default averages 60 FPS with 4.9 GB, fast setting 120 FPS with 2.5 GB, frame rate above 48 FPS for 95% of frames; training 3.5 days on one RTX-3090 (Sec. 1, Sec. 3.2, Figs. 8 and 10)","https:\u002F\u002Fgithub.com\u002Fprinceton-vl\u002FDPVO","MIT (LICENSE file checked)",[52,56,59],{"relation":53,"title":54,"doi_or_url":55},"preprint","arXiv:2208.04726v2","https:\u002F\u002Farxiv.org\u002Fabs\u002F2208.04726",{"relation":57,"title":58,"doi_or_url":49},"code_release","princeton-vl\u002FDPVO",{"relation":60,"title":61,"doi_or_url":62},"extension","Deep Patch Visual SLAM (DPV-SLAM), ECCV 2024","https:\u002F\u002Fdoi.org\u002F10.1007\u002F978-3-031-72627-9_24",{"id":5,"kind":64,"shortName":7,"title":8,"authors":65,"year":9,"venue":69,"venueType":70,"publisher":71,"volumeIssuePages":72,"doi":73,"arxivId":74,"url":75,"firstPublicDate":76,"publicationStatus":16,"metadataStatus":77,"fulltextStatus":15,"era":10,"classicReason":42,"codeUrl":49,"cluster":11,"topics":78,"mdpi":79,"verification":80,"label":6,"fulltextRoute":81,"versionRead":82,"addedByCensus":83},"method",[66,67,68],"Zachary Teed","Lahav Lipson","Jia Deng","Advances in Neural Information Processing Systems 36 (NeurIPS 2023)","conference","Neural Information Processing Systems Foundation, Inc. (NeurIPS)","pp. 39033-39051","10.52202\u002F075280-1696","2208.04726","https:\u002F\u002Fapi.crossref.org\u002Fworks\u002F10.52202\u002F075280-1696","2022-08-08","metadata_verified",[11],false,"corrected","arXiv","arXiv v2 (2208.04726v2, 23 May 2023) including Appendices A-H read in full; cross-checked against the NeurIPS 2023 proceedings version (all extracted values match)",true,[85],{"category":86,"model":87,"canonical":87,"role":88,"dataset":89,"specs":90,"locator":91},"compute","RTX-3090","compute for runtime",null,"GPU for runtime measurements and for training (single GPU, 3.5 days)","Sec. 1, Sec. 3.2",[],1790510662270]