[{"data":1,"prerenderedAt":358},["ShallowReactive",2],{"method-monoslam2007":3},{"method":4,"reference":54,"equipment":76,"figures":113,"results":114},{"id":5,"label":6,"shortName":7,"title":8,"year":9,"era":10,"cluster":11,"scope":12,"keyIdeaZh":13,"keyIdeaEn":14,"fulltextStatus":15,"publicationStatus":16,"recommendation":17,"constructionRelevance":18,"validationEnvironment":19,"strengths":22,"limitations":27,"sensors":35,"platform":38,"estimator":41,"association":42,"timeModel":43,"deskew":44,"loopClosure":45,"globalOptimization":46,"mapRepresentation":47,"prior":48,"outputGeometry":49,"compute":50,"codeUrl":51,"codeLicense":52,"relatedVersions":53},"monoslam2007","Davison et al., 2007","MonoSLAM","MonoSLAM: Real-Time Single Camera SLAM",2007,"classic","C08","full_slam_with_global_correction","MonoSLAM 以單一延伸卡爾曼濾波器（Extended Kalman Filter, EKF）同時估計相機位姿與稀疏自然地標，並保留兩者之間的完整共變異數（covariance），讓單眼相機即可即時建立持續存在的機率式地圖。作者引入依預測不確定度進行的主動量測（active measurement）與平滑運動模型，以降低影像處理成本。其地圖為稀疏地標而非稠密點雲，作者也將適用範圍描述為室內、房間尺度。","MonoSLAM estimates camera pose and a sparse set of natural landmarks within one EKF with full covariance, using uncertainty-guided active feature search to run monocular SLAM in real time at room scale.","full_text_reviewed","peer_reviewed_published","background","論文未報告營建工地、既有建築量測或基礎設施測試；示範為手持擴增實境與人形機器人。對本文主要為視覺 SLAM 濾波式架構的歷史背景。",[20,21],"controlled_experiment","independent_reference",[23,24,25,26],"Real-time 30 Hz monocular operation on commodity hardware (abstract)","Persistent probabilistic map with explicit uncertainty of camera and features (Sec. 3.1)","Tabletop ground-truth test: localization accurate to a few centimetres with jitter of about 1 to 2 cm (Sec. 6.1)","Small and large loops are routinely closed; the HRP-2 loop on a 0.75 m radius circle is closed with drift corrected (Sec. 4, 5.3, Fig. 10)",[28,29,30,31,32,33,34],"Authors describe operation in room-sized indoor domains and list larger environments (indoors and outdoors), more dynamic motions and changing lighting as future work (Sec. 1; Sec. 7)","(inference) Sparse landmark map does not by itself provide a dense point cloud for engineering geometry","Cannot cope with views without useful features, such as a blank wall or ceiling (Sec. 4)","Cannot cope with very sudden jerky movement under the chosen motion noise (Sec. 4)","Feature initialisation would perform poorly for motion along the optic axis (Sec. 3.6)","Mismatches between clutter and landmarks can cause catastrophic failure, and active search is not suited to relocalising a lost camera (Sec. 3.5, 3.7)","Some waypoint errors exceed the jitter (x estimate -0.93 m for -1.00 m) and shrink only about 1 cm per loop (Sec. 6.1)",[36,37],"monocular wide-angle camera (field of view near 100 degrees, 30 Hz)","3-axis gyro fused as an internal angular-velocity measurement in the HRP-2 humanoid experiment only",[39,40],"handheld","legged","extended Kalman filter over joint camera and landmark state with full covariance (Sec. 3.1)","Shi-Tomasi salient 11x11 pixel patches stored as locally planar templates, warped to the predicted view and matched by normalised cross-correlation only inside 3-sigma innovation-covariance ellipses (typically 15 to 20 pixels across); per frame the 10 to 12 features with the highest innovation covariance are measured; new features start as 3D rays with 100 depth particles between 0.5 and 5 m and become 3D points when depth std over depth falls below 0.3; features failing more than 50% of attempted measurements are deleted","Discrete-time EKF with a constant velocity, constant angular velocity model driven by zero-mean Gaussian acceleration impulses (std 10 m\u002Fs^2 linear and 6 rad\u002Fs^2 angular in the hand-held setup); 13-parameter camera state (position, orientation quaternion, velocity, angular velocity)","not_applicable","implicit loop closure inside the single EKF: re-observing early-mapped features after exploration corrects accumulated drift, and active feature selection favours such re-observation; no separate place-recognition module (Sec. 4 AR results; Sec. 5 humanoid circular walk, Fig. 10 'loop closed and drift corrected')","none (single EKF maintains joint covariance; no separate optimization back-end)","Single state vector and full covariance over the camera and about 100 sparse 3D point features, each stored with an oriented planar patch template; optional surface-normal estimates kept in separate two-parameter EKFs per feature","A known initialisation target (typically four features of known position and appearance, such as the corners of a black rectangle) fixes the world frame and metric scale, with the camera started at an approximately known pose; on HRP-2 natural and artificial features at measured positions on a wall replace the target. The camera is pre-calibrated (e.g. fku = fkv = 195 px, (u0, v0) = (162, 125), K1 = 6e-6 at 320x240, one-parameter radial model)","camera trajectory and sparse 3D landmark positions with uncertainty (Sec. 3.1)","Typically 19 ms per frame on a 1.6 GHz Pentium M at 30 Hz (image loading 2 ms, correlation searches 3 ms, Kalman update 5 ms, feature initialisation search 4 ms, graphical rendering 5 ms); the O(N^2) filter bounds the map to about 100 features at 30 Hz","http:\u002F\u002Fwww.doc.ic.ac.uk\u002F~ajd\u002FScene\u002F","LGPL (SceneLib 1.0; stated in paper Sec. 6.3 and on the SceneLib homepage; licence file not opened)",[],{"id":5,"kind":55,"shortName":7,"title":8,"authors":56,"year":9,"venue":61,"venueType":62,"publisher":63,"volumeIssuePages":64,"doi":65,"arxivId":66,"url":67,"firstPublicDate":68,"publicationStatus":16,"metadataStatus":69,"fulltextStatus":15,"era":10,"classicReason":70,"codeUrl":51,"cluster":11,"topics":71,"mdpi":72,"verification":73,"label":6,"fulltextRoute":74,"versionRead":75,"addedByCensus":72},"method",[57,58,59,60],"Andrew J. Davison","Ian D. Reid","Nicholas D. Molton","Olivier Stasse","IEEE Transactions on Pattern Analysis and Machine Intelligence","journal","IEEE","29(6):1052-1067","10.1109\u002Ftpami.2007.1049",null,"https:\u002F\u002Fwww.doc.ic.ac.uk\u002F~ajd\u002FPublications\u002Fdavison_etal_pami2007.pdf","2007-04","metadata_verified","necessary technical node: first real-time monocular SLAM with a single EKF over camera and sparse landmarks, the filtering baseline against which keyframe\u002FBA systems (PTAM, ORB-SLAM) are later contrasted.",[11],false,"corrected","author copy","Author-hosted PDF in the IEEE Computer Society typeset layout (16 pages numbered 1 to 16, header TPAMI Vol. 29 No. 6, June 2007); IEEE Xplore paginated version (pp. 1052-1067) not compared",[77,83,87,92,97,103,108],{"category":78,"model":79,"canonical":79,"role":80,"dataset":66,"specs":81,"locator":82},"camera","low-cost IEEE 1394 webcam with a wide-angle lens (model not stated)","method input","30 Hz; field of view nearly 100 degrees; calibrated at 320x240 with fku = fkv = 195 px; monochrome images used","Sec. 3.2, 3.5, 4",{"category":78,"model":84,"canonical":84,"role":80,"dataset":66,"specs":85,"locator":86},"HRP-2 additional wide-angle camera (model not stated)","field of view around 90 degrees; one-parameter radial distortion model","Sec. 5.1",{"category":88,"model":89,"canonical":89,"role":80,"dataset":66,"specs":90,"locator":91},"imu","HRP-2 3-axis chest gyro (model not stated)","reports angular velocity at 200 Hz, sampled at 30 Hz, std 0.01 rad\u002Fs per axis","Sec. 5.2",{"category":93,"model":94,"canonical":94,"role":80,"dataset":66,"specs":95,"locator":96},"platform","HRP-2 humanoid robot","walked a 0.75 m radius circle in about 30 s, SLAM on board with a wireless Ethernet link","Sec. 5, 5.3; Fig. 9",{"category":98,"model":99,"canonical":99,"role":100,"dataset":66,"specs":101,"locator":102},"compute","1.6 GHz Pentium M","compute for runtime","typical 19 ms processing per frame at 30 Hz","Sec. 6.2",{"category":104,"model":105,"canonical":105,"role":80,"dataset":66,"specs":106,"locator":107},"other","initialisation target: black rectangle with four known corner features","defines world frame and metric scale at start-up","Sec. 3.3; Fig. 2a",{"category":104,"model":109,"canonical":109,"role":110,"dataset":66,"specs":111,"locator":112},"plumb-line of known length over a precisely measured rectangular desktop track","reference or ground truth","ground-truth camera coordinates at four waypoints with an assessed 1 cm precision","Sec. 6.1; Fig. 11",[],{"totalRows":115,"groupCount":116,"groups":117,"others":357},19,3,[118,188,227],{"slug":119,"group":120,"sourceId":5,"sourceLabel":6,"table":121,"selfRows":122,"metrics":123,"seqs":139,"entrants":150,"cells":153,"outcomes":181,"locators":182,"hardware":184,"wordings":185,"notes":186},"monoslam2007-table-in-sec-6-1","monoslam2007:Table in Sec. 6.1","Table in Sec. 6.1",12,[124,129,131,133,135,137],{"label":125,"unit":126,"statistic":127,"alignment":128},"estimated camera x coordinate (std 0.01 m)","m","mean","none",{"label":130,"unit":126,"statistic":127,"alignment":128},"estimated camera y coordinate (std 0.01 m)",{"label":132,"unit":126,"statistic":127,"alignment":128},"estimated camera z coordinate (std 0.01 m)",{"label":134,"unit":126,"statistic":127,"alignment":128},"estimated camera x coordinate (std 0.03 m)",{"label":136,"unit":126,"statistic":127,"alignment":128},"estimated camera y coordinate (std 0.02 m)",{"label":138,"unit":126,"statistic":127,"alignment":128},"estimated camera z coordinate (std 0.02 m)",[140,144,146,148],{"dataset":141,"sequence":142,"environment":143},"authors' desktop ground-truth track","waypoint 1, ground truth (0.00, 0.00, -0.62) m","indoor cluttered desktop, hand-held camera",{"dataset":141,"sequence":145,"environment":143},"waypoint 2, ground truth (-1.00, 0.00, -0.62) m",{"dataset":141,"sequence":147,"environment":143},"waypoint 3, ground truth (-1.00, 0.50, -0.62) m",{"dataset":141,"sequence":149,"environment":143},"waypoint 4, ground truth (0.00, 0.50, -0.62) m",[151],{"name":7,"methodId":5,"linkable":152,"proposed":152,"self":152},true,[154,157,160,163,165,168,171,173,175,177,178,180],[155,155,155,155,156,155,156,156,155],0,-1,[155,158,155,159,156,155,156,156,155],1,0.01,[155,161,155,162,156,155,156,156,155],2,0.64,[155,116,158,164,156,155,156,156,155],-0.93,[155,166,158,167,156,155,156,156,155],4,0.06,[155,169,158,170,156,155,156,156,155],5,0.63,[155,116,161,172,156,155,156,156,155],-0.98,[155,166,161,174,156,155,156,156,155],0.46,[155,169,161,176,156,155,156,156,155],0.66,[155,155,116,159,156,155,156,156,155],[155,166,116,179,156,155,156,156,155],0.47,[155,169,116,162,156,155,156,156,155],[],[183],"Sec. 6.1 table",[],[],[187],"Ground-truth characterisation on a desktop track: mean MonoSLAM camera position over several looped revisits (std in brackets) at four waypoints, hand-held wide-angle camera at 30 Hz; minus signs read from the rendered page; estimated z printed as positive while ground-truth z is -0.62 m (recorded as printed); no post-hoc trajectory alignment: the world frame is fixed at start-up by the standard initialisation target placed at one corner of the track (Sec. 6.1)",{"slug":189,"group":190,"sourceId":5,"sourceLabel":6,"table":191,"selfRows":192,"metrics":193,"seqs":208,"entrants":212,"cells":214,"outcomes":221,"locators":222,"hardware":223,"wordings":224,"notes":225},"monoslam2007-text-sec-6-2","monoslam2007:Text Sec.6.2","Text Sec.6.2",6,[194,198,200,202,204,206],{"label":195,"unit":196,"statistic":197,"alignment":44},"Total per frame","ms","not_reported",{"label":199,"unit":196,"statistic":197,"alignment":44},"Image loading and administration",{"label":201,"unit":196,"statistic":197,"alignment":44},"Image correlation searches",{"label":203,"unit":196,"statistic":197,"alignment":44},"Kalman Filter update",{"label":205,"unit":196,"statistic":197,"alignment":44},"Feature initialization search",{"label":207,"unit":196,"statistic":197,"alignment":44},"Graphical rendering",[209],{"dataset":197,"sequence":210,"environment":211},"typical frame","not_reported (typical per-frame breakdown at 30 Hz; the sequence is not specified)",[213],{"name":7,"methodId":5,"linkable":152,"proposed":152,"self":152},[215,216,217,218,219,220],[155,155,155,115,156,155,155,156,155],[155,158,155,161,156,155,155,156,155],[155,161,155,116,156,155,155,156,155],[155,116,155,169,156,155,155,156,155],[155,166,155,166,156,155,155,156,155],[155,169,155,169,156,155,155,156,155],[],[102],[99],[],[226],"Typical breakdown of per-frame processing time at 30 Hz (33 ms budget) on a 1.6 GHz Pentium M",{"slug":228,"group":229,"sourceId":230,"sourceLabel":231,"table":232,"selfRows":158,"metrics":233,"seqs":236,"entrants":240,"cells":301,"outcomes":341,"locators":352,"hardware":353,"wordings":354,"notes":355},"ghadimzadeh2025slamnde-table-2","ghadimzadeh2025slamnde:Table 2","ghadimzadeh2025slamnde","Ghadimzadeh Alamdari et al., 2025","Table 2",[234],{"label":235,"unit":128,"statistic":197,"alignment":128},"Result (run outcome)",[237],{"dataset":238,"sequence":197,"environment":239},"Luleå SubT tunnel dataset (Koval et al. 2022)","underground tunnel",[241,243,246,248,250,253,256,259,262,265,268,270,273,275,277,279,282,284,287,289,291,294,296,299],{"name":242,"methodId":5,"linkable":152,"proposed":72,"self":152},"Mono-SLAM",{"name":244,"methodId":245,"linkable":152,"proposed":72,"self":72},"PTAM","ptam2007",{"name":247,"methodId":66,"linkable":72,"proposed":72,"self":72},"S-PTAM",{"name":249,"methodId":66,"linkable":72,"proposed":72,"self":72},"OV2SLAM",{"name":251,"methodId":252,"linkable":152,"proposed":72,"self":72},"ORB-SLAM (footnote 1)","orbslam2015",{"name":254,"methodId":255,"linkable":152,"proposed":72,"self":72},"DTAM","dtam2011",{"name":257,"methodId":258,"linkable":152,"proposed":72,"self":72},"LSD-SLAM","lsdslam2014",{"name":260,"methodId":261,"linkable":152,"proposed":72,"self":72},"SVO","svo2017",{"name":263,"methodId":264,"linkable":152,"proposed":72,"self":72},"DSO","dso2018",{"name":266,"methodId":267,"linkable":152,"proposed":72,"self":72},"Kinetic Fusion","kinectfusion2011",{"name":269,"methodId":66,"linkable":72,"proposed":72,"self":72},"Dense visual SLAM",{"name":271,"methodId":272,"linkable":152,"proposed":72,"self":72},"Elastic Fusion SLAM","elasticfusion2015",{"name":274,"methodId":66,"linkable":72,"proposed":72,"self":72},"Realtime onboard VI estimation",{"name":276,"methodId":66,"linkable":72,"proposed":72,"self":72},"Multi-sensor fusion",{"name":278,"methodId":66,"linkable":72,"proposed":72,"self":72},"SOFT-SLAM",{"name":280,"methodId":281,"linkable":152,"proposed":72,"self":72},"MSCKF","mourikis2007msckf",{"name":283,"methodId":66,"linkable":72,"proposed":72,"self":72},"ROVIO",{"name":285,"methodId":286,"linkable":152,"proposed":72,"self":72},"OKVIS","okvis2015",{"name":288,"methodId":66,"linkable":72,"proposed":72,"self":72},"VIORB",{"name":290,"methodId":66,"linkable":72,"proposed":72,"self":72},"S-MSCKF",{"name":292,"methodId":293,"linkable":152,"proposed":72,"self":72},"VINS-Mono","vinsmono2018",{"name":295,"methodId":66,"linkable":72,"proposed":72,"self":72},"STCM-SLAM",{"name":297,"methodId":298,"linkable":152,"proposed":72,"self":72},"Kimera","kimera2020",{"name":300,"methodId":66,"linkable":72,"proposed":72,"self":72},"Yolo-SLAM",[302,303,304,305,306,307,308,309,311,313,315,317,319,320,322,324,326,328,330,332,333,335,337,339],[155,155,155,66,155,155,156,156,155],[158,155,155,66,158,155,156,156,155],[161,155,155,66,161,155,156,156,155],[116,155,155,66,161,155,156,156,155],[166,155,155,66,116,155,156,156,155],[169,155,155,66,166,155,156,156,155],[192,155,155,66,169,155,156,156,155],[310,155,155,66,192,155,156,156,155],7,[312,155,155,66,161,155,156,156,155],8,[314,155,155,66,166,155,156,156,155],9,[316,155,155,66,155,155,156,156,155],10,[318,155,155,66,310,155,156,156,155],11,[122,155,155,66,166,155,156,156,155],[321,155,155,66,166,155,156,156,155],13,[323,155,155,66,166,155,156,156,155],14,[325,155,155,66,192,155,156,156,155],15,[327,155,155,66,161,155,156,156,155],16,[329,155,155,66,192,155,156,156,155],17,[331,155,155,66,310,155,156,156,155],18,[115,155,155,66,166,155,156,156,155],[334,155,155,66,312,155,156,156,155],20,[336,155,155,66,166,155,156,156,155],21,[338,155,155,66,155,155,156,156,155],22,[340,155,155,66,314,155,156,156,155],23,[342,343,344,345,346,347,348,349,350,351],"failed (feature detection and tracking)","failed (initialization for ground floor)","not_run (authors could not run the code)","success (footnote 1: authors could not run ORB-SLAM 3, so the original ORB-SLAM was used)","not_run (no publicly available repository)","failed (feature tracking)","failed (tracking)","not_run (inconsistent repository)","success","other: Result cell reads 'SLAM for dynamic environments'; no run outcome stated",[232],[],[],[356],"Run outcome ('Result' column) of each reviewed vision-based method on the Luleå tunnel test dataset; '+' marks methods not integrated with ROS; the '*' (incompatible with VLP-16) symbol is printed on almost every row",[],1790510659121]