[{"data":1,"prerenderedAt":93},["ShallowReactive",2],{"method-cnnslam2017":3},{"method":4,"reference":54,"equipment":76,"figures":92,"results":47},{"id":5,"label":6,"shortName":7,"title":8,"year":9,"era":10,"cluster":11,"scope":12,"keyIdeaZh":13,"keyIdeaEn":14,"fulltextStatus":15,"publicationStatus":16,"recommendation":17,"constructionRelevance":18,"validationEnvironment":19,"strengths":22,"limitations":27,"sensors":32,"platform":34,"estimator":37,"association":38,"timeModel":39,"deskew":40,"loopClosure":41,"globalOptimization":42,"mapRepresentation":43,"prior":44,"outputGeometry":45,"compute":46,"codeUrl":47,"codeLicense":48,"relatedVersions":49},"cnnslam2017","Tateno et al., 2017","CNN-SLAM","CNN-SLAM: Real-Time Dense Monocular SLAM with Learned Depth Prediction",2017,"recent","C09","full_slam_with_global_correction","CNN-SLAM 以 LSD-SLAM 的直接法關鍵影格架構為基礎，只在建立關鍵影格時用卷積網路預測稠密深度，並依目前相機與訓練相機的焦距比例調整尺度，再以後續影格的小基線立體匹配依不確定度加權修正深度。低紋理區保留網路預測、高梯度區由立體量測主導，因此單眼 SLAM 可取得絕對尺度，在純旋轉運動下也能重建。關鍵影格以位姿圖最佳化，深度圖可再融合成含語意標籤的三維模型。","Monocular direct SLAM (LSD-SLAM style) whose key-frame depth is initialized by CNN prediction, rescaled by focal length and refined by uncertainty-weighted small-baseline stereo, recovering absolute scale, handling pure rotation and fusing semantic labels; key-frames are pose-graph optimized.","full_text_reviewed","peer_reviewed_published","background","論文未涉及營建場域；定量評估只用合成 ICL-NUIM 與 TUM 辦公室序列，平均 ATE 約 0.25 m，且深度誤差在 10% 內的比例僅約兩成，與施工量測所需的公分級精度差距很大。其以學習式深度先驗恢復單眼絕對尺度的想法，是後續單眼神經 SLAM 的重要源頭（推論）。",[20,21],"public_benchmark","simulation",[23,24,25,26],"Lowest average ATE (0.246 m) across 9 ICL-NUIM and TUM sequences, below LSD-SLAM bootstrapped with ground-truth depth (0.562 m) and ORB-SLAM (0.643 m) (Table 1)","Highest average share of correct depth (22.464%) versus 3.032% for bootstrapped LSD-SLAM and 18.452% for raw CNN depth fusion (Table 1)","Reconstructs scenes under mostly pure rotation (TUM fr1\u002Frpy) where LSD-SLAM is noisy and ORB-SLAM fails to initialize (Sec. 4.2, Fig. 5)","First joint 3D and semantic reconstruction from a monocular camera, per the authors (Sec. 4.3)",[28,29,30,31],"Bootstrapped LSD-SLAM has lower ATE on 4 of 9 sequences (for example TUM seq3: 0.037 versus 0.214 m) (Table 1)","Only 12-37% of depth values fall within 10% of ground truth even after refinement (Table 1)","Absolute scale depends on a CNN trained on NYU Depth v2 and on a focal-length correction; performance with other cameras and scene types is shown only on ICL-NUIM and TUM (Sec. 3.3, Sec. 4) (inference)","Future work: closing the loop between geometric refinement and depth prediction (Sec. 5)",[33],"monocular camera",[35,36],"simulation (ICL-NUIM)","not_reported (TUM RGB-D sequences recorded with a Kinect; own office sequence setup not described)","LSD-SLAM-style direct key-frame tracking: weighted Gauss-Newton minimization of Huber-weighted photometric residuals on high-gradient pixels against the nearest key-frame; key-frame depth from a CNN (ResNet-50 fully convolutional network of Laina et al.) scaled by the focal-length ratio and refined by uncertainty-weighted fusion of small-baseline stereo depth; key-frame pose-graph optimization (Sec. 3.1-3.4)","direct photometric alignment; per-frame depth from 5-pixel epipolar matching for refinement (Sec. 3.1, Sec. 3.4)","discrete poses (key-frames)","not_applicable","pose-graph edges added between a new key-frame and existing key-frames with a similar field of view (small relative pose); no appearance-based place recognition described (Sec. 3.3)","pose-graph optimization of key-frame poses at each new key-frame (Sec. 3.3)","dense per-key-frame depth and uncertainty maps fused into a global 3D model with optional semantic labels using the fusion scheme of its reference [27] (Sec. 3.5)","CNN depth prediction and 4-class semantic segmentation trained on NYU Depth v2 (indoor, Kinect), ResNet-50 initialized on ImageNet; depth scaled by the ratio of current to training focal length (Sec. 3.2-3.3, Sec. 4)","camera trajectory, dense key-frame depth maps and a fused, optionally semantically labelled, 3D reconstruction (Sec. 3.5, Figs. 1 and 6)","Intel Xeon 2.4 GHz CPU with 16 GB RAM and Nvidia Quadro K5200 (8 GB); CNN runs on the GPU and the other stages on two CPU threads; described as real-time but no frame rate is reported (Sec. 4)",null,"not_applicable (no official code release found)",[50],{"relation":51,"title":52,"doi_or_url":53},"preprint","arXiv:1704.03489v1","https:\u002F\u002Farxiv.org\u002Fabs\u002F1704.03489",{"id":5,"kind":55,"shortName":7,"title":8,"authors":56,"year":9,"venue":61,"venueType":62,"publisher":63,"volumeIssuePages":64,"doi":65,"arxivId":66,"url":67,"firstPublicDate":68,"publicationStatus":16,"metadataStatus":69,"fulltextStatus":15,"era":10,"classicReason":40,"codeUrl":47,"cluster":11,"topics":70,"mdpi":71,"verification":72,"label":6,"fulltextRoute":73,"versionRead":74,"addedByCensus":75},"method",[57,58,59,60],"Keisuke Tateno","Federico Tombari","Iro Laina","Nassir Navab","2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR)","conference","IEEE","pp. 6565-6574","10.1109\u002Fcvpr.2017.695","1704.03489","https:\u002F\u002Fapi.crossref.org\u002Fworks\u002F10.1109\u002FCVPR.2017.695","2017-04-11","metadata_verified",[11],false,"confirmed","arXiv","arXiv v1 (1704.03489v1, 11 Apr 2017), the only arXiv version, labelled CVPR 2017 by the authors; cross-checked against the CVPR 2017 CVF open-access paper (all extracted values match)",true,[77,83,86],{"category":78,"model":79,"canonical":79,"role":80,"dataset":47,"specs":81,"locator":82},"compute","Intel Xeon CPU at 2.4GHz","compute for runtime","desktop PC with 16GB of RAM","Sec. 4",{"category":78,"model":84,"canonical":84,"role":80,"dataset":47,"specs":85,"locator":82},"Nvidia Quadro K5200","8GB VRAM; runs CNN depth prediction and semantic segmentation",{"category":87,"model":88,"canonical":88,"role":89,"dataset":90,"specs":91,"locator":82},"rgbd","Kinect","dataset sensor","TUM RGB-D; NYU Depth v2","TUM RGB-D is 'acquired with a Kinect sensor'; NYU Depth v2 ground truth from a Microsoft Kinect camera",[],1790510663873]