[{"data":1,"prerenderedAt":87},["ShallowReactive",2],{"method-deepfactors2020":3},{"method":4,"reference":56,"equipment":78,"figures":86,"results":83},{"id":5,"label":6,"shortName":7,"title":8,"year":9,"era":10,"cluster":11,"scope":12,"keyIdeaZh":13,"keyIdeaEn":14,"fulltextStatus":15,"publicationStatus":16,"recommendation":17,"constructionRelevance":18,"validationEnvironment":19,"strengths":22,"limitations":27,"sensors":32,"platform":34,"estimator":36,"association":37,"timeModel":38,"deskew":39,"loopClosure":40,"globalOptimization":41,"mapRepresentation":42,"prior":43,"outputGeometry":44,"compute":45,"codeUrl":46,"codeLicense":47,"relatedVersions":48},"deepfactors2020","Czarnowski et al., 2020","DeepFactors","DeepFactors: Real-Time Probabilistic Dense Monocular SLAM",2020,"recent","C09","full_slam_with_global_correction","DeepFactors 把 CodeSLAM 的學習式精簡深度編碼放進標準因子圖：每個關鍵影格的深度由 32 維編碼經以影像為條件的線性解碼器產生，位姿與編碼一起以 GTSAM 的 iSAM2 做批次最大後驗估計。關鍵影格間同時使用稠密光度誤差、BRISK 特徵重投影誤差與稀疏幾何深度一致性誤差三種因子，並有局部與以詞袋檢索的全域迴圈閉合。它在單張 GTX 1080 上可即時運作，但輸出深度解析度只有 256x192，且小編碼難以表示完全平坦的表面。","Real-time monocular dense SLAM that optimizes key-frame poses and learned 32-D depth codes in a GTSAM\u002FiSAM2 factor graph with photometric, reprojection and geometric factors, plus local and bag-of-words global loop closure.","full_text_reviewed","peer_reviewed_published","supplementary","論文未涉及營建場域；評估使用 ScanNet、ICL-NUIM 與 TUM RGB-D 室內資料，深度精度以落在真值 10% 以內的像素比例衡量，平均約 27%，且需以最佳尺度校正單眼結果。其價值在於示範學習式深度先驗可與因子圖、iSAM2 與迴圈閉合等傳統 SLAM 後端結合（推論），但解析度與精度都不足以直接產生施工量測用點雲。",[20,21],"public_benchmark","simulation",[23,24,25,26],"Combining photometric, reprojection and geometric factors gives the best ATE and depth accuracy in the ScanNet ablation (Table I)","Average 27.10% of depth pixels within 10% of truth on ICL-NUIM and TUM versus 19.77% for CNN-SLAM (Table II)","Outperforms CodeSLAM on all and CNN-SLAM on four of five TUM fr1 sequences while running in real time (Table III)","Standard factor-graph formulation allows other sensors to be added (Sec. VIII)",[28,29,30,31],"Small code size (32) cannot represent fully flat depth well, hurting the flat-wall TUM seq2 sequence (Sec. VI-C)","Network Jacobian computation takes about 340 ms per key-frame; speed depends on enabled factors, and the geometric factor is disabled for fast exploration (Sec. VII)","Monocular: trajectories and depth are scaled with the optimal scale from the TUM scripts for evaluation (Sec. VI-C)","Low working resolution of 256x192 (Sec. VI-A, Sec. VII)",[33],"monocular camera",[35],"not_reported (ScanNet, ICL-NUIM and TUM RGB-D sequences and a live camera; capture platforms not described)","batch MAP factor graph over key-frame poses and compact depth codes solved incrementally with iSAM2 in GTSAM; camera tracking by GPU direct whole-image SE(3) Lucas-Kanade against the closest key-frame; tracking and mapping interleaved; one-way frames refine the latest key-frame and are then marginalized (Sec. V)","three pairwise factor types: dense photometric error, BRISK keypoint reprojection error with Cauchy cost, and sparse geometric depth-consistency error with Huber cost (Sec. III)","discrete poses (key-frames)","not_applicable","local loops by a pose-based criterion within the last 10 key-frames; global loops by bag-of-words (DBoW2) candidates verified by tracking inliers and pose distance, closed with reprojection factors only (Sec. V-D)","incremental batch optimization of all key-frames (iSAM2) with zero-code prior factors (Sec. V-B)","key-frame depth maps decoded linearly from a 32-dimensional learned code conditioned on the image (CodeSLAM-style variational auto-encoder) (Sec. III-IV, Sec. VI-A)","network trained on about 1.4M ScanNet images with merged rendered and sensor depth, image size 256x192, code size 32; an explicit code-prediction path gives the initial depth of each key-frame (Sec. IV, Sec. VI-A)","key-frame poses and dense key-frame depth maps fused for visualization as reconstructions (Figs. 1 and 7)","single NVIDIA GTX 1080 at 256x192; about 340 ms per new key-frame for network and Jacobian (16 ms forward pass), tracking at about 250 Hz; overall speed depends on connectivity, enabled factors and loop closures (Sec. VII)","https:\u002F\u002Fgithub.com\u002Fjczarnowski\u002FDeepFactors","custom Imperial College licence for non-commercial, internal or academic research use (LICENSE file checked; GitHub reports NOASSERTION)",[49,53],{"relation":50,"title":51,"doi_or_url":52},"preprint","arXiv:2001.05049v1 (accepted manuscript of the RA-L paper)","https:\u002F\u002Farxiv.org\u002Fabs\u002F2001.05049",{"relation":54,"title":55,"doi_or_url":46},"code_release","jczarnowski\u002FDeepFactors",{"id":5,"kind":57,"shortName":7,"title":8,"authors":58,"year":9,"venue":63,"venueType":64,"publisher":65,"volumeIssuePages":66,"doi":67,"arxivId":68,"url":69,"firstPublicDate":70,"publicationStatus":16,"metadataStatus":71,"fulltextStatus":15,"era":10,"classicReason":39,"codeUrl":46,"cluster":11,"topics":72,"mdpi":73,"verification":74,"label":6,"fulltextRoute":75,"versionRead":76,"addedByCensus":77},"method",[59,60,61,62],"Jan Czarnowski","Tristan Laidlow","Ronald Clark","Andrew J. Davison","IEEE Robotics and Automation Letters","journal","IEEE","5(2):721-728","10.1109\u002Flra.2020.2965415","2001.05049","https:\u002F\u002Fapi.crossref.org\u002Fworks\u002F10.1109\u002FLRA.2020.2965415","2020-01-09","metadata_verified",[11],false,"corrected","arXiv","arXiv v1 (2001.05049v1, 14 Jan 2020), the accepted RA-L manuscript, read in full; the RA-L version of record on IEEE Xplore was also read and its tables match",true,[79],{"category":80,"model":81,"canonical":81,"role":82,"dataset":83,"specs":84,"locator":85},"compute","NVIDIA GTX 1080","compute for runtime",null,"single GPU running the network, CUDA kernels and visualization at 256x192","Sec. VII",[],1790510664760]