[{"data":1,"prerenderedAt":80},["ShallowReactive",2],{"method-deepvo2017":3},{"method":4,"reference":50,"equipment":72,"figures":79,"results":43},{"id":5,"label":6,"shortName":7,"title":8,"year":9,"era":10,"cluster":11,"scope":12,"keyIdeaZh":13,"keyIdeaEn":14,"fulltextStatus":15,"publicationStatus":16,"recommendation":17,"constructionRelevance":18,"validationEnvironment":19,"strengths":21,"limitations":25,"sensors":30,"platform":32,"estimator":34,"association":35,"timeModel":36,"deskew":37,"loopClosure":38,"globalOptimization":38,"mapRepresentation":39,"prior":40,"outputGeometry":41,"compute":42,"codeUrl":43,"codeLicense":44,"relatedVersions":45},"deepvo2017","Wang et al., 2017","DeepVO","DeepVO: Towards end-to-end visual odometry with deep Recurrent Convolutional Neural Networks",2017,"recent","C09","odometry","DeepVO 是早期的端到端單眼視覺里程計：把相鄰兩張 RGB 影像疊合後送入以 FlowNet 預訓練權重初始化的卷積網路擷取運動特徵，再以兩層 LSTM 建模時間序列，直接迴歸每一時刻的六自由度位姿。方法不需特徵擷取、匹配、光束法平差，甚至不需相機校正，絕對尺度由訓練資料隱式學得。論文只在 KITTI 上驗證，平移漂移優於單眼 LIBVISO2，但仍明顯不如立體 LIBVISO2，且不產生地圖。","End-to-end monocular visual odometry that regresses 6-DoF poses directly from stacked consecutive RGB frames with a FlowNet-initialized CNN and a two-layer LSTM, learning absolute scale from KITTI training data; outputs poses only.","full_text_reviewed","peer_reviewed_published","background","論文未涉及營建場域；只以 KITTI 車載資料訓練與測試，輸出僅為軌跡而無地圖。其絕對尺度由訓練資料學得，換到不同相機或工地場景時是否仍成立並未驗證（推論），因此主要作為學習式里程計的歷史節點，而非可直接產生施工點雲的方法。",[20],"public_benchmark",[22,23,24],"Mean translational drift 5.96% on KITTI test sequences versus 17.48% for monocular VISO2 (Table II)","Recovers absolute scale without camera-height priors or post alignment to ground truth (Sec. IV-B)","Needs no hand-designed VO modules or camera calibration (Sec. I, Sec. V)",[26,27,28,29],"Less accurate than stereo VISO2 (mean translational drift 1.89%) (Table II)","Translational error grows at high speeds because training data above 50 km\u002Fh are scarce; sequence 12 shows large errors (Sec. IV-B)","Prone to overfitting, especially in orientation, which degrades generalization (Sec. IV-A)","Authors stress it is a complement, not a replacement, for geometry-based VO (Sec. V)",[31],"monocular camera",[33],"vehicle (KITTI)","end-to-end regression: a 9-layer CNN initialized from pretrained FlowNet extracts features from two stacked consecutive RGB frames, and two stacked LSTM layers (1000 hidden units each) output a 6-DoF pose per time step; trained with MSE on positions and Euler angles (orientation weight 100) (Sec. III)","none explicit; motion is learned implicitly from stacked image pairs without feature matching (Sec. III)","discrete poses (one per frame)","not_applicable","none","none (poses only)","supervised training on KITTI sequences with ground-truth poses; CNN initialized from a pretrained FlowNet model; absolute scale is learned from the training data (Sec. IV-A)","6-DoF camera trajectory only; no map or point cloud","implemented in Theano and trained on an NVIDIA Tesla K40 GPU; inference runtime not reported (Sec. IV-A)",null,"not_applicable (no official code found; the project page links the paper, a video and a dataset only)",[46],{"relation":47,"title":48,"doi_or_url":49},"postprint","arXiv:1709.08429v1 (ICRA 2017 paper posted after the conference)","https:\u002F\u002Farxiv.org\u002Fabs\u002F1709.08429",{"id":5,"kind":51,"shortName":7,"title":8,"authors":52,"year":9,"venue":57,"venueType":58,"publisher":59,"volumeIssuePages":60,"doi":61,"arxivId":62,"url":63,"firstPublicDate":64,"publicationStatus":16,"metadataStatus":65,"fulltextStatus":15,"era":10,"classicReason":37,"codeUrl":43,"cluster":11,"topics":66,"mdpi":67,"verification":68,"label":6,"fulltextRoute":69,"versionRead":70,"addedByCensus":71},"method",[53,54,55,56],"Sen Wang","Ronald Clark","Hongkai Wen","Niki Trigoni","2017 IEEE International Conference on Robotics and Automation (ICRA)","conference","IEEE","pp. 2043-2050","10.1109\u002Ficra.2017.7989236","1709.08429","https:\u002F\u002Fapi.crossref.org\u002Fworks\u002F10.1109\u002FICRA.2017.7989236","2017-05","metadata_verified",[11],false,"confirmed","arXiv","arXiv v1 (1709.08429v1, 25 Sep 2017), the ICRA 2017 paper as posted by the authors, read in full; the ICRA version of record on IEEE Xplore was also read and its Table II matches",true,[73],{"category":74,"model":75,"canonical":75,"role":76,"dataset":43,"specs":77,"locator":78},"compute","NVIDIA Tesla K40","compute for runtime","GPU used for training in Theano; inference runtime not reported","Sec. IV-A",[],1790510662611]