[{"data":1,"prerenderedAt":422},["ShallowReactive",2],{"method-kimera2020":3},{"method":4,"reference":63,"equipment":83,"figures":114,"results":115},{"id":5,"label":6,"shortName":7,"title":8,"year":9,"era":10,"cluster":11,"scope":12,"keyIdeaZh":13,"keyIdeaEn":14,"fulltextStatus":15,"publicationStatus":16,"recommendation":17,"constructionRelevance":18,"validationEnvironment":19,"strengths":22,"limitations":27,"sensors":33,"platform":37,"estimator":39,"association":40,"timeModel":41,"deskew":42,"loopClosure":43,"globalOptimization":44,"mapRepresentation":45,"prior":46,"outputGeometry":47,"compute":48,"codeUrl":49,"codeLicense":50,"relatedVersions":51},"kimera2020","Rosinol et al., 2020","Kimera","Kimera: an Open-Source Library for Real-Time Metric-Semantic Localization and Mapping",2020,"recent","C08","full_slam_with_global_correction","Kimera 是模組化的開源度量語意（metric-semantic）視覺慣性 SLAM 函式庫，包含以 GTSAM iSAM2 固定延遲平滑器實作的 VIO、以 PCM 剔除錯誤迴圈的強健位姿圖最佳化、低延遲 3D 網格生成器，以及用雙目稠密匹配（SGM）與 Voxblox TSDF 產生全域語意網格的模組。各模組可獨立或組合執行，並在 CPU 上即時運作。作者以 EuRoC 地面真值點雲評估網格的精度與完整度，但評估前先以 ICP 將估計點雲對齊至真值。","Kimera combines a GTSAM-based VIO, robust PCM-filtered pose-graph optimisation, a fast mesher and a Voxblox TSDF semantic mesh built from dense stereo, evaluating mesh accuracy and completeness against EuRoC ground-truth clouds after ICP alignment.","full_text_reviewed","peer_reviewed_published","main_body","論文未報告營建測試；評估使用 EuRoC 與照片級模擬器。作者特別提到不同樓層相同房間造成的感知混淆（perceptual aliasing），與多樓層建築掃描直接相關。其幾何評估流程（取樣網格、ICP 對齊、精度與完整度）可作為工程點雲評估參考，但 ICP 對齊會隱藏絕對位置誤差（推論）；且作者報告的全域網格平均誤差為 0.35 至 0.48 m（ICP 門檻 1.0 m；Table IV 標題稱為完整度），遠大於一般工程量測容許差（推論）。作者也指出稠密立體匹配難以處理無紋理牆面，是模擬場景中幾何與語意誤差最大的來源（Sec. III-C、Table V）。",[20,21],"public_benchmark","simulation",[23,24,25,26],"Modules can run in isolation or together, falling back to VIO or full SLAM (abstract)","On EuRoC V1 and V2 the global TSDF mesh had lower error than the multi-frame mesh on 4 of 6 sequences (slightly higher on V1_02 and V2_01), while the multi-frame mesher requires two orders of magnitude less time (Sec. III-B; Table IV)","With PCM, V1_01 ATE stays between 0.045 and 0.05 m for loop thresholds from 10 to 0.001, while PGO without PCM reaches up to 1.74 m (Table III)","VIO drift below 0.2% (4 cm over a 32 m simulated trajectory) (Sec. III-C)",[28,29,30,31,32],"Loop closures may still contain outliers from perceptual aliasing, e.g. identical rooms on different floors (Sec. II-B1)","(inference) Mesh accuracy is measured after ICP registration to ground truth, so global placement error is not included in the reported metric","Dense stereo has difficulty resolving texture-less walls, causing the largest drop in geometric and semantic accuracy (RMSE 0.215 m vs 0.131 m with ground-truth depth; mIoU 57.23% vs 80.03%) (Sec. III-C; Table V; Fig. 4)","Kimera-RPGO runtime depends on pose-graph size (Sec. III-D)","(observation) Table IV caption calls the metric completeness while the text calls the same values average error, so the direction of the reported distance is ambiguous",[34,35,36],"monocular camera","stereo","IMU",[38,21],"UAV","keyframe-based MAP visual-inertial estimator run as full or fixed-lag smoothing (fixed-lag typically used to bound estimation time), with on-manifold IMU preintegration and structureless vision factors solved by iSAM2 in GTSAM and marginalisation of states leaving the horizon; Kimera-RPGO keeps odometry and loop edges separately, selects the largest consistent loop set with a modified incremental PCM (odometry chi-squared check plus pairwise consistency, fast maximum clique) and optimises the pose graph with Gauss-Newton in GTSAM","Shi-Tomasi corners tracked by Lucas-Kanade, left-right stereo matching, 5-point mono and 3-point stereo RANSAC verification (optional 2-point and 1-point variants using IMU rotation) at keyframes; structureless vision factors triangulated by DLT with degenerate and high-reprojection-error points removed; DBoW2 bag-of-words loop candidates verified with the same mono and stereo checks","discrete keyframe states; IMU-rate estimates (Sec. II)","not_applicable","DBoW2 putative loops, geometric verification, outlier rejection with a modified PCM before GTSAM PGO (Sec. II-B)","robust pose-graph optimisation (Kimera-RPGO) (Sec. II-B2)","per-frame mesh from 2D Delaunay triangulation of tracked features back-projected with VIO landmark estimates; multi-frame mesh over the VIO horizon, coupled back to VIO through regularity factors when planar surfaces are detected in the mesh; global Voxblox TSDF built at keyframes from semi-global-matching dense stereo with bundled raycasting (fast option), meshed by marching cubes, with Bayesian per-voxel semantic label updates within the truncation distance","2D semantic segmentation of images for labelling (Sec. II-D2)","trajectory, low-latency local mesh, and global semantically annotated mesh from TSDF (Sec. II)","CPU only (model not reported), four threads; IMU preintegration about 40 us (IMU-rate estimates above 200 Hz), feature tracking 4.5 ms per frame, keyframe front-end 45 ms, per-frame mesh under 5 ms, multi-frame mesh 15 ms, VIO back-end under 40 ms, Kimera-RPGO 55 ms on average on EuRoC, Kimera-Semantics about 0.1 s per keyframe for a 720x480 depth image","https:\u002F\u002Fgithub.com\u002FMIT-SPARK\u002FKimera","BSD (README states BSD License, LICENSE.BSD)",[52,56,60],{"relation":53,"title":54,"doi_or_url":55},"preprint","Kimera (arXiv)","https:\u002F\u002Farxiv.org\u002Fabs\u002F1910.02490",{"relation":57,"title":58,"doi_or_url":59},"journal_extension","Kimera: From SLAM to spatial perception with 3D dynamic scene graphs (IJRR 40(12-14):1510-1546, 2021)","10.1177\u002F02783649211056674",{"relation":61,"title":62,"doi_or_url":49},"code_release","Kimera \u002F Kimera-VIO",{"id":5,"kind":64,"shortName":7,"title":8,"authors":65,"year":9,"venue":70,"venueType":71,"publisher":72,"volumeIssuePages":73,"doi":74,"arxivId":75,"url":55,"firstPublicDate":76,"publicationStatus":16,"metadataStatus":77,"fulltextStatus":15,"era":10,"classicReason":42,"codeUrl":49,"cluster":11,"topics":78,"mdpi":79,"verification":80,"label":6,"fulltextRoute":81,"versionRead":82,"addedByCensus":79},"software",[66,67,68,69],"Antoni Rosinol","Marcus Abate","Yun Chang","Luca Carlone","2020 IEEE International Conference on Robotics and Automation (ICRA)","conference","IEEE","pp. 1689-1696","10.1109\u002Ficra40945.2020.9196885","1910.02490","2019-10-06","metadata_verified",[11],false,"corrected","arXiv","arXiv 1910.02490v3 (2020-03-04, 8 pages, marked accepted for ICRA 2020); ICRA 2020 IEEE version of record not compared",[84,91,96,102,107],{"category":85,"model":86,"canonical":86,"role":87,"dataset":88,"specs":89,"locator":90},"stereo_camera","EuRoC MAV stereo camera (model not reported in this paper)","dataset sensor","EuRoC MAV","stereo frames used as Kimera input (monocular mode also supported)","Sec. II; Sec. III-A",{"category":92,"model":93,"canonical":93,"role":87,"dataset":88,"specs":94,"locator":95},"imu","EuRoC MAV IMU (model not reported in this paper)","high-rate inertial measurements; state estimates output at IMU rate (above 200 Hz)","Sec. II; Sec. III-D",{"category":97,"model":98,"canonical":98,"role":99,"dataset":88,"specs":100,"locator":101},"other","EuRoC V1 and V2 ground-truth point cloud (acquisition device not stated in this paper)","reference or ground truth","used for mesh accuracy and completeness after ICP registration","Sec. III-B",{"category":97,"model":103,"canonical":103,"role":87,"dataset":104,"specs":105,"locator":106},"Unity-based photo-realistic simulator provided by MIT Lincoln Lab","photo-realistic simulator (MIT Lincoln Lab)","ROS sensor streams with ground-truth geometry, semantics, depth and poses; 720x480 dense depth images","Sec. III-C; Sec. III-D; Acknowledgments",{"category":108,"model":109,"canonical":109,"role":110,"dataset":111,"specs":112,"locator":113},"compute","CPU (model not reported)","compute for runtime",null,"real-time CPU execution with four threads","Abstract; Sec. II; Sec. III-D",[],{"totalRows":116,"groupCount":117,"groups":118,"others":384},92,10,[119,230,283,339],{"slug":120,"group":121,"sourceId":5,"sourceLabel":6,"table":122,"selfRows":123,"metrics":124,"seqs":130,"entrants":154,"cells":162,"outcomes":224,"locators":225,"hardware":226,"wordings":227,"notes":228},"kimera2020-table-ii","kimera2020:Table II","Table II",33,[125],{"label":126,"unit":127,"statistic":128,"alignment":129},"RMSE ATE [m]","m","RMSE","SE3",[131,134,136,138,140,142,144,146,148,150,152],{"dataset":88,"sequence":132,"environment":133},"MH_01","EuRoC MAV sequences (micro aerial vehicle dataset, ref. [19]); environments not described in this paper",{"dataset":88,"sequence":135,"environment":133},"MH_02",{"dataset":88,"sequence":137,"environment":133},"MH_03",{"dataset":88,"sequence":139,"environment":133},"MH_04",{"dataset":88,"sequence":141,"environment":133},"MH_05",{"dataset":88,"sequence":143,"environment":133},"V1_01",{"dataset":88,"sequence":145,"environment":133},"V1_02",{"dataset":88,"sequence":147,"environment":133},"V1_03",{"dataset":88,"sequence":149,"environment":133},"V2_01",{"dataset":88,"sequence":151,"environment":133},"V2_02",{"dataset":88,"sequence":153,"environment":133},"V2_03",[155,158,160],{"name":156,"methodId":5,"linkable":157,"proposed":157,"self":157},"Kimera-VIO (fixed-lag smoothing)",true,{"name":159,"methodId":5,"linkable":157,"proposed":157,"self":157},"Kimera-VIO (full smoothing)",{"name":161,"methodId":5,"linkable":157,"proposed":157,"self":157},"Kimera-RPGO (loop closure)",[163,167,170,173,176,179,182,185,188,190,192,194,196,197,199,201,203,205,206,208,209,210,212,213,214,215,217,218,219,220,221,222,223],[164,164,164,165,166,164,166,166,164],0,0.11,-1,[164,164,168,169,166,164,166,166,164],1,0.1,[164,164,171,172,166,164,166,166,164],2,0.16,[164,164,174,175,166,164,166,166,164],3,0.24,[164,164,177,178,166,164,166,166,164],4,0.35,[164,164,180,181,166,164,166,166,164],5,0.05,[164,164,183,184,166,164,166,166,164],6,0.08,[164,164,186,187,166,164,166,166,164],7,0.07,[164,164,189,184,166,164,166,166,164],8,[164,164,191,169,166,164,166,166,164],9,[164,164,117,193,166,164,166,166,164],0.21,[168,164,164,195,166,164,166,166,164],0.04,[168,164,168,187,166,164,166,166,164],[168,164,171,198,166,164,166,166,164],0.12,[168,164,174,200,166,164,166,166,164],0.27,[168,164,177,202,166,164,166,166,164],0.2,[168,164,180,204,166,164,166,166,164],0.06,[168,164,183,187,166,164,166,166,164],[168,164,186,207,166,164,166,166,164],0.09,[168,164,189,187,166,164,166,166,164],[168,164,191,207,166,164,166,166,164],[168,164,117,211,166,164,166,166,164],0.19,[171,164,164,184,166,164,166,166,164],[171,164,168,207,166,164,166,166,164],[171,164,171,165,166,164,166,166,164],[171,164,174,216,166,164,166,166,164],0.15,[171,164,177,175,166,164,166,166,164],[171,164,180,181,166,164,166,166,164],[171,164,183,165,166,164,166,166,164],[171,164,186,198,166,164,166,166,164],[171,164,189,187,166,164,166,166,164],[171,164,191,169,166,164,166,166,164],[171,164,117,211,166,164,166,166,164],[],[122],[],[],[229],"EuRoC ATE RMSE grouped as fixed-lag smoothing, full smoothing and PGO with loop closure; comparator values taken from Delmerico and Scaramuzza [77] (Sim(3) alignment per text) and VINS-Mono [24]; comparators use a monocular camera while Kimera uses stereo; Kimera aligned with SE(3); loop threshold alpha = 0.001",{"slug":231,"group":232,"sourceId":5,"sourceLabel":6,"table":233,"selfRows":234,"metrics":235,"seqs":238,"entrants":247,"cells":252,"outcomes":277,"locators":278,"hardware":279,"wordings":280,"notes":281},"kimera2020-table-iv","kimera2020:Table IV","Table IV",12,[236],{"label":237,"unit":127,"statistic":128,"alignment":129},"RMSE [m] (Table IV caption: completeness [78, Sec. 4.3.3])",[239,242,243,244,245,246],{"dataset":240,"sequence":143,"environment":241},"EuRoC MAV (V1, V2 ground-truth point cloud)","EuRoC MAV V1 and V2 sequences with ground-truth point cloud (environment not described in this paper)",{"dataset":240,"sequence":145,"environment":241},{"dataset":240,"sequence":147,"environment":241},{"dataset":240,"sequence":149,"environment":241},{"dataset":240,"sequence":151,"environment":241},{"dataset":240,"sequence":153,"environment":241},[248,250],{"name":249,"methodId":5,"linkable":157,"proposed":157,"self":157},"Kimera-Mesher multi-frame mesh",{"name":251,"methodId":5,"linkable":157,"proposed":157,"self":157},"Kimera-Semantics global TSDF mesh",[253,255,257,259,261,263,265,267,269,271,273,275],[164,164,164,254,166,164,166,166,164],0.482,[164,164,168,256,166,164,166,166,164],0.374,[164,164,171,258,166,164,166,166,164],0.451,[164,164,174,260,166,164,166,166,164],0.465,[164,164,177,262,166,164,166,166,164],0.491,[164,164,180,264,166,164,166,166,164],0.53,[168,164,164,266,166,164,166,166,164],0.364,[168,164,168,268,166,164,166,166,164],0.384,[168,164,171,270,166,164,166,166,164],0.353,[168,164,174,272,166,164,166,166,164],0.48,[168,164,177,274,166,164,166,166,164],0.432,[168,164,180,276,166,164,166,166,164],0.411,[],[233],[],[],[282],"Mesh evaluated against the EuRoC ground-truth point cloud: mesh sampled at 10^3 points per m2, registered by rigid ICP in CloudCompare (ICP threshold 1.0 m); caption calls the metric completeness while the text describes the same values as average error of the global mesh (0.35 to 0.48 m); Multi-Frame mesh computed with a large VIO horizon (full smoothing)",{"slug":284,"group":285,"sourceId":5,"sourceLabel":6,"table":286,"selfRows":234,"metrics":287,"seqs":300,"entrants":304,"cells":311,"outcomes":333,"locators":334,"hardware":335,"wordings":336,"notes":337},"kimera2020-table-v","kimera2020:Table V","Table V",[288,293,296,298],{"label":289,"unit":290,"statistic":291,"alignment":292},"mIoU [%]","%","mean","none",{"label":294,"unit":290,"statistic":295,"alignment":292},"Acc [%] (portion of correctly labelled points)","not_reported",{"label":297,"unit":127,"statistic":295,"alignment":295},"ATE [m]",{"label":299,"unit":127,"statistic":128,"alignment":129},"Geometric RMSE [m] of mesh points after registration",[301],{"dataset":104,"sequence":302,"environment":303},"simulated scene (about 32 m trajectory, Sec. III-C)","photo-realistic Unity-based simulated scene; semantic classes include wall, shelf and floor (Sec. III-C, Fig. 4)",[305,307,309],{"name":306,"methodId":5,"linkable":157,"proposed":79,"self":157},"Kimera-Semantics with GT depth and GT poses",{"name":308,"methodId":5,"linkable":157,"proposed":79,"self":157},"Kimera-Semantics with GT depth and Kimera-VIO poses",{"name":310,"methodId":5,"linkable":157,"proposed":157,"self":157},"Kimera-Semantics with dense stereo and Kimera-VIO poses",[312,314,316,317,319,321,323,324,326,328,330,331],[164,164,164,313,166,164,166,166,164],80.1,[164,168,164,315,166,164,166,166,164],94.68,[164,171,164,164,166,164,166,166,164],[164,174,164,318,166,164,166,166,164],0.079,[168,164,164,320,166,164,166,166,164],80.03,[168,168,164,322,166,164,166,166,164],94.5,[168,171,164,195,166,164,166,166,164],[168,174,164,325,166,164,166,166,164],0.131,[171,164,164,327,166,164,166,166,164],57.23,[171,168,164,329,166,164,166,166,164],80.74,[171,171,164,195,166,164,166,166,164],[171,174,164,332,166,164,166,166,164],0.215,[],[286],[],[],[338],"Unity-based photo-realistic simulator (MIT Lincoln Lab) with ground-truth 2D semantics; mesh registered to ground truth and point RMSE computed as in Sec. III-B (distance direction not specified); simulated trajectory about 32 m",{"slug":340,"group":341,"sourceId":5,"sourceLabel":6,"table":342,"selfRows":117,"metrics":343,"seqs":345,"entrants":357,"cells":362,"outcomes":378,"locators":379,"hardware":380,"wordings":381,"notes":382},"kimera2020-table-iii","kimera2020:Table III","Table III",[344],{"label":126,"unit":127,"statistic":128,"alignment":129},[346,349,351,353,355],{"dataset":88,"sequence":347,"environment":348},"V1_01, alpha=10","EuRoC MAV V1_01 (environment not described in this paper)",{"dataset":88,"sequence":350,"environment":348},"V1_01, alpha=1",{"dataset":88,"sequence":352,"environment":348},"V1_01, alpha=0.1",{"dataset":88,"sequence":354,"environment":348},"V1_01, alpha=0.01",{"dataset":88,"sequence":356,"environment":348},"V1_01, alpha=0.001",[358,360],{"name":359,"methodId":5,"linkable":157,"proposed":79,"self":157},"Kimera ablation: PGO w\u002Fo PCM",{"name":361,"methodId":5,"linkable":157,"proposed":157,"self":157},"Kimera-RPGO",[363,364,366,368,370,371,372,373,374,376],[164,164,164,181,166,164,166,166,164],[164,164,168,365,166,164,166,166,164],0.45,[164,164,171,367,166,164,166,166,164],1.74,[164,164,174,369,166,164,166,166,164],1.59,[164,164,177,369,166,164,166,166,164],[168,164,164,181,166,164,166,166,164],[168,164,168,181,166,164,166,166,164],[168,164,171,181,166,164,166,166,164],[168,164,174,375,166,164,166,166,164],0.045,[168,164,177,377,166,164,166,166,164],0.049,[],[342],[],[],[383],"EuRoC V1_01 ATE RMSE versus DBoW2 loop-closure threshold alpha; smaller alpha gives more but less conservative loop closures",[385,391,397,403,409,416],{"group":386,"slug":387,"sourceLabel":6,"table":388,"selfRows":189,"datasets":389},"kimera2020:Text Sec. III-D","kimera2020-text-sec-iii-d","Text Sec. III-D",[88,390,104],"not_reported (Fig. 5 runtime breakdown)",{"group":392,"slug":393,"sourceLabel":394,"table":342,"selfRows":189,"datasets":395},"kimeramulti2022:Table III","kimeramulti2022-table-iii","Tian et al., 2022",[396],"DCIST simulation",{"group":398,"slug":399,"sourceLabel":394,"table":286,"selfRows":183,"datasets":400},"kimeramulti2022:Table V","kimeramulti2022-table-v",[401,402],"Medfield outdoor dataset (authors' own)","Stata outdoor dataset (authors' own)",{"group":404,"slug":405,"sourceLabel":406,"table":407,"selfRows":168,"datasets":408},"dmvio2022:Table I","dmvio2022-table-i","von Stumberg & Cremers, 2022","Table I",[88],{"group":410,"slug":411,"sourceLabel":412,"table":413,"selfRows":168,"datasets":414},"ghadimzadeh2025slamnde:Table 2","ghadimzadeh2025slamnde-table-2","Ghadimzadeh Alamdari et al., 2025","Table 2",[415],"Luleå SubT tunnel dataset (Koval et al. 2022)",{"group":417,"slug":418,"sourceLabel":419,"table":122,"selfRows":168,"datasets":420},"orbslam3_2021:Table II","orbslam3-2021-table-ii","Campos et al., 2021",[421],"EuRoC",1790510662322]