[{"data":1,"prerenderedAt":190},["ShallowReactive",2],{"method-asadi2019imagebimslam":3},{"method":4,"reference":53,"equipment":74,"figures":95,"results":96},{"id":5,"label":6,"shortName":7,"title":8,"year":9,"era":10,"cluster":11,"scope":12,"keyIdeaZh":13,"keyIdeaEn":14,"fulltextStatus":15,"publicationStatus":16,"recommendation":17,"constructionRelevance":18,"validationEnvironment":19,"strengths":22,"limitations":26,"sensors":32,"platform":34,"estimator":36,"association":37,"timeModel":38,"deskew":39,"loopClosure":40,"globalOptimization":41,"mapRepresentation":42,"prior":43,"outputGeometry":44,"compute":45,"codeUrl":46,"codeLicense":47,"relatedVersions":48},"asadi2019imagebimslam","Asadi et al., 2019","Asadi et al. 2019 SLAM image-to-BIM registration","Real-Time Image Localization and Registration with BIM Using Perspective Alignment for Indoor Monitoring of Construction",2019,"recent","C11b","localization_in_prior_map_or_bim","作者提出把影片關鍵影格即時對位到設計 BIM 的方法。定位部分以 ORB-SLAM2 為基礎，改用由影片產生的自訂詞袋字典提升缺乏特徵室內場景的追蹤；第一次蒐集時以 MVE 稠密點雲和 BIM 的手動對應角點求相似轉換，建立真實尺度的全域地圖，之後各次蒐集只在該地圖中重新定位。每個關鍵影格先由 SLAM 位姿產生對應的 BIM 視圖，再以梯度下降最小化影像與 BIM 視圖之間消失點距離與消失線夾角，得到精修位姿。作者在走廊與室內施工工地兩段影片驗證，並報告在 Jetson TX1 上的每影格計算時間。","Real-time registration of video keyframes to an as-planned BIM: an ORB-SLAM2-based monocular SLAM with a custom vocabulary relocalizes in a BIM-scaled global map, and each keyframe pose is refined by aligning vanishing points and lines with the rendered BIM view; tested in a hallway and an indoor construction site on a Jetson TX1.","full_text_reviewed","peer_reviewed_published","supplementary","其中一個驗證場景是施工公司提供 BIM 的室內施工工地（120 秒影片、59 個關鍵影格），用來說明影像對位 BIM 以支援進度監測；精度只以消失點與消失線的對齊誤差描述，沒有獨立量測的相機位姿參考（Experimental Setup and Results、Discussion）。",[20,21],"completed_building","real_construction_site",[23,24,25],"Augmented SLAM with custom vocabulary tracked two 90 deg turns in a featureless hallway where ORB-SLAM2 produced two separate hallways (Fig. 7; Sec. Initial Frame Registration)","Angular errors were below 3 deg before and below 0.01 deg after iteration in both scenes (Figs. 9-10 text)","Real-time registration for the hallway video (0.65 s per keyframe versus 0.74 keyframes per second) (Discussion)",[27,28,29,30,31],"On the cluttered construction site the 640 x 360 perspective estimation was inaccurate for about 15 keyframes; higher resolution was more accurate but far too slow (Discussion; Figs. 13-14)","Construction-site processing (1.9 s per keyframe) exceeded the keyframe rate, so the UGV had to move slower (Discussion)","Fine-pose alignment cannot correct the camera position along the direction normal to the image plane, which still relies on SLAM (Conclusion)","Curved walls or arches are expected to produce higher error because the method relies on straight edges (Conclusion)","First-frame registration and the first global map require a manual step (Method)",[33],"monocular camera (webcam, 1920 x 1080 at 30 fps, fixed focal length; model not reported)",[35],"UGV with a webcam from Asadi et al. 2018d (the AutCon 96:470-482 system); platform model and drive type not restated in this paper","Augmented monocular SLAM built on ORB-SLAM2 with a custom DBoW2 vocabulary generated from the video; first session builds a global map that is scaled by a manual similarity transform to BIM, later sessions relocalize in this map with the stored scale; keyframe poses then refined by gradient-descent alignment of vanishing points and vanishing lines between the keyframe and the rendered BIM view (max 500 iterations, thresholds 1 pixel and 1 deg)","ORB features for tracking and relocalization; Canny edges and Hedau et al. vanishing point voting for keyframe perspective; BIM vanishing points computed directly from model geometry","discrete keyframes","not_applicable (camera only)","ORB-SLAM2 place recognition and relocalization; the perspective step is presented as reducing drift before loop closure","none beyond ORB-SLAM2; refined keyframe poses overwrite the SLAM poses at the end, described as similar to global bundle adjustment after loop closure","sparse ORB-SLAM2 map in real-world scale; dense MVE point cloud used once for manual alignment to BIM","as-planned BIM (visible model lines, rendered views), manual first-frame registration via corresponding corners between dense point cloud and BIM","keyframe camera poses in the BIM coordinate system and the matching BIM views; no point-cloud accuracy product","NVIDIA Jetson TX1 on the UGV: 657 ms (hallway) and 1,904 ms (construction site) average per keyframe; desktop Intel i7 3.4 GHz 6-core below 0.2 s per site keyframe",null,"not_applicable",[49],{"relation":50,"title":51,"doi_or_url":52},"conference_precursor","Real-time image-to-BIM registration using perspective alignment for automated construction monitoring, Construction Research Congress 2018, pp. 388-397 (cited as Asadi and Han 2018; not read)","https:\u002F\u002Fdoi.org\u002F10.1061\u002F9780784481264.038",{"id":5,"kind":54,"shortName":7,"title":8,"authors":55,"year":9,"venue":60,"venueType":61,"publisher":62,"volumeIssuePages":63,"doi":64,"arxivId":46,"url":65,"firstPublicDate":66,"publicationStatus":16,"metadataStatus":67,"fulltextStatus":15,"era":10,"classicReason":47,"codeUrl":46,"cluster":11,"topics":68,"mdpi":69,"verification":70,"label":6,"fulltextRoute":71,"versionRead":72,"addedByCensus":73},"method",[56,57,58,59],"Khashayar Asadi","Hariharan Ramshankar","Mojtaba Noghabaei","Kevin Han","Journal of Computing in Civil Engineering","journal","American Society of Civil Engineers (ASCE)","33(5):04019031","10.1061\u002F(asce)cp.1943-5487.0000847","https:\u002F\u002Fascelibrary.org\u002Fdoi\u002Ffull\u002F10.1061\u002F(ASCE)CP.1943-5487.0000847","2019-06-13","metadata_verified",[11],false,"corrected","NTU institutional (Chrome)","version of record, J. Comput. Civ. Eng. 33(5):04019031, ASCE Library HTML full text",true,[75,81,85,91],{"category":76,"model":77,"canonical":77,"role":78,"dataset":46,"specs":79,"locator":80},"camera","webcam (monocular; model not reported)","method input","1,920 x 1,080 video at 30 fps; fixed focal length; intrinsics from MATLAB calibration; keyframes downsampled to 640 x 360 for real time","Sec. Experimental Setup and Results (Initial Setup)",{"category":82,"model":83,"canonical":83,"role":78,"dataset":46,"specs":84,"locator":80},"platform","UGV from Asadi et al. 2018d (model not restated)","moved along a path while recording video",{"category":86,"model":87,"canonical":87,"role":88,"dataset":46,"specs":89,"locator":90},"compute","NVIDIA Jetson TX1","compute for runtime","runs the proposed SLAM and perspective processing on the UGV","Sec. Initial Frame Registration and Camera Localization; Sec. Computation Time",{"category":86,"model":92,"canonical":92,"role":88,"dataset":46,"specs":93,"locator":94},"desktop with Intel i7 processor","3.4 GHz, 6 cores","Sec. Discussion",[],{"totalRows":97,"groupCount":98,"groups":99,"others":189},4,3,[100,139,165],{"slug":101,"group":102,"sourceId":5,"sourceLabel":6,"table":103,"selfRows":104,"metrics":105,"seqs":110,"entrants":119,"cells":122,"outcomes":130,"locators":131,"hardware":133,"wordings":135,"notes":136},"asadi2019imagebimslam-text-computation-time","asadi2019imagebimslam:Text Computation Time","Text Computation Time",2,[106],{"label":107,"unit":108,"statistic":109,"alignment":47},"average computation time per keyframe","ms","mean",[111,115],{"dataset":112,"sequence":113,"environment":114},"authors' hallway video","hallway (52 keyframes)","completed building hallway with featureless walls",{"dataset":116,"sequence":117,"environment":118},"authors' construction-site video","indoor construction site (59 keyframes)","indoor construction site (cluttered)",[120],{"name":121,"methodId":5,"linkable":73,"proposed":73,"self":73},"proposed augmented SLAM + perspective alignment",[123,127],[124,124,124,125,126,124,124,126,124],0,657,-1,[124,124,128,129,126,124,124,126,128],1,1904,[],[132],"Sec. Computation Time (Fig. 12 values stated in text)",[134],"NVIDIA Jetson TX1 on the UGV",[],[137,138],"Average computation per keyframe at 640 x 360 (rough pose from the augmented SLAM, VP\u002FVL estimation, iterative fine-pose alignment) on the UGV's Jetson TX1; hallway video 70 s with 52 keyframes","Average computation per keyframe at 640 x 360 on the Jetson TX1; construction-site video 120 s with 59 keyframes; many iterations reached the cap eta = 500",{"slug":140,"group":141,"sourceId":5,"sourceLabel":6,"table":142,"selfRows":128,"metrics":143,"seqs":148,"entrants":151,"cells":154,"outcomes":157,"locators":159,"hardware":160,"wordings":162,"notes":163},"asadi2019imagebimslam-text-discussion","asadi2019imagebimslam:Text Discussion","Text Discussion",[144],{"label":145,"unit":146,"statistic":147,"alignment":47},"processing time per keyframe on desktop (stated as less than 0.2 s)","s","not_reported",[149],{"dataset":116,"sequence":150,"environment":118},"indoor construction site",[152],{"name":153,"methodId":5,"linkable":73,"proposed":73,"self":73},"proposed method on desktop",[155],[124,124,124,156,124,124,124,126,124],0.2,[158],"other: upper bound stated in text ('less than 0.2 s')",[94],[161],"desktop Intel i7 (3.4 GHz, 6 cores)",[],[164],"Same construction-site keyframes processed on a desktop computer",{"slug":166,"group":167,"sourceId":5,"sourceLabel":6,"table":168,"selfRows":128,"metrics":169,"seqs":173,"entrants":176,"cells":179,"outcomes":182,"locators":183,"hardware":185,"wordings":186,"notes":187},"asadi2019imagebimslam-text-practical-implications","asadi2019imagebimslam:Text Practical Implications","Text Practical Implications",[170],{"label":171,"unit":172,"statistic":147,"alignment":47},"distance error of 18 pixels (Fig. 13) in the 17th keyframe","pixel",[174],{"dataset":116,"sequence":175,"environment":118},"indoor construction site, keyframe 17",[177],{"name":178,"methodId":5,"linkable":73,"proposed":73,"self":73},"proposed method (VP estimation at 640 x 360; keyframe 17 of the construction-site video)",[180],[124,124,124,181,126,124,126,126,124],18,[],[184],"Sec. Practical Implications (refers to Fig. 13)",[],[],[188],"Distance error of 18 pixels for keyframe 17 of the construction-site video, cited in the text with reference to Fig. 13 (per-keyframe mean square error between estimated and ground-truth vanishing points at 640 x 360) to explain the minor misalignment shown in Fig. 15(b and d)",[],1790510656272]