[{"data":1,"prerenderedAt":197},["ShallowReactive",2],{"method-okvis2x2025":3},{"method":4,"reference":62,"equipment":86,"figures":157,"results":91},{"id":5,"label":6,"shortName":7,"title":8,"year":9,"era":10,"cluster":11,"scope":12,"keyIdeaZh":13,"keyIdeaEn":14,"fulltextStatus":15,"publicationStatus":16,"recommendation":17,"constructionRelevance":18,"validationEnvironment":19,"strengths":23,"limitations":29,"sensors":34,"platform":40,"estimator":42,"association":43,"timeModel":44,"deskew":45,"loopClosure":46,"globalOptimization":47,"mapRepresentation":48,"prior":49,"outputGeometry":50,"compute":51,"codeUrl":52,"codeLicense":53,"relatedVersions":54},"okvis2x2025","Boche et al., 2025","OKVIS2-X","OKVIS2-X: Open Keyframe-Based Visual-Inertial SLAM Configurable With Dense Depth or LiDAR, and GNSS",2025,"recent","C07","full_slam_with_global_correction","OKVIS2-X 以關鍵影格式視覺慣性 SLAM（OKVIS2）為核心，可選擇加入深度網路估計的稠密深度、LiDAR 或 GNSS。系統把 supereight2 體素占據子地圖綁定在關鍵影格上，並以影格對地圖、地圖對地圖的占據對齊殘差把子地圖與狀態估計緊耦合，使迴圈閉合或 GNSS 修正後地圖可隨位姿一起更新，形成全域一致、可直接用於導航的稠密地圖。GNSS 以可觀測性準則處理座標對齊與訊號中斷，並支援相機外參線上校正。","Extends OKVIS2 visual-inertial SLAM with keyframe-anchored volumetric occupancy submaps built from learned depth or LiDAR and tightly coupled by occupancy alignment factors, plus optional GNSS with observability-aware alignment and online extrinsic calibration, giving globally consistent dense maps at large scale.","full_text_reviewed","peer_reviewed_published","main_body","OKVIS2-X 直接輸出全域一致的稠密占據地圖與網格，並在 Hilti-Oxford 資料上以地面真值點雲評估網格精度（LiDAR 版本平均 2.9 cm，完整度 78 %）；依 zhang2023hiltioxford 的資料集說明，Hilti-Oxford 包含施工現場序列，但本文未說明 exp04 至 exp06 的場景類型。這是少數同時提供軌跡與建圖精度的多感測器系統，對施工現場掃描具參考價值（推論）。",[20,21,22],"public_benchmark","independent_reference","cross_site",[24,25,26,27,28],"On VBR (1.0 to 9.0 km sequences) Ours-vil-nc average ATE RMSE 0.829 m versus 2.475 m for FAST-LIVO; visual-inertial Ours-vid-nc 2.210 m versus 5.572 m for ORB-SLAM3 and 28.770 m for OpenVINS (Table VI)","Hilti-Oxford mesh accuracy 0.029 m and completeness 78.33 % with LiDAR, 0.042 m and 64.27 % with learned depth, at a 0.2 m threshold (Table V)","Hilti 2022 challenge score 438.10 for Ours-vil-ba, above FAST-LIVO2 (402.93) and VILENS (325.77) but below LiDAR-inertial Wildcat (563.79) (Table IV)","Faster than ORB-SLAM3 on EuRoC MH05 (38.1 versus 64.7 ms per frame) and on par on Hilti exp06 (Sec. VI)","Open-source with configurable sensor setups",[30,31,32,33],"Dense mapping raises memory: up to 25.5 GB for Ours-vil on VBR Campus1 (Table IX)","Higher runtime than ORB-SLAM3 on VBR Campus1 (107.3 versus 69.1 ms per frame) (Sec. VI)","Long corridors (Hilti exp07) are degenerate for LiDAR alignment (Sec. VI)","Learned depth networks require a GPU",[35,36,37,38,39],"one or more cameras","IMU","optional learned depth (stereo and multi-view stereo networks)","optional LiDAR","optional GNSS",[41],"GVINS dataset complex_environment sequence (carrying platform not described in this paper)","OKVIS2 keyframe-based visual-inertial SLAM (BRISK features, realtime sliding-window estimator plus asynchronous full-graph optimization with posegraph edges from marginalized landmarks, DBoW2 loop closure) extended with volumetric occupancy submaps anchored to keyframes and tightly coupled through frame-to-map and map-to-map occupancy alignment factors (Tukey robustifier), optional GNSS position factors and online camera-IMU extrinsic calibration (Secs. IV-V)","BRISK keypoint matching for landmarks; dense submap alignment by occupancy residuals of depth or LiDAR points against submaps; sky segmentation (Fast-SCNN) to remove sky pixels from depth (Sec. V)","discrete multi-camera frames with IMU pre-integration between states; LiDAR points and GNSS measurements are related to states through IMU propagation (Sec. V)","not described as a separate step","yes, DBoW2 place recognition with full-graph optimization (Sec. V)","asynchronous full-graph optimization and optional final full bundle adjustment ('-ba' variants); GNSS alignment with a yaw-observability criterion and global re-alignment after GNSS dropouts (Sec. V)","supereight2 volumetric occupancy submaps anchored to keyframes, meshable into a globally consistent map (Figs. 1, 9, 10)","camera intrinsics and initial extrinsics; networks fine-tuned on synthetic data","trajectory (causal, non-causal and full-BA variants) and dense occupancy submaps or meshes","desktop with Intel i7-13700 and RTX-3080 10 GB; Ours-vi up to 47 Hz on EuRoC MH05; MVS network at 13 Hz on 8-view images; GPU memory 3.51 GB for the networks; also runs on a drone with NVIDIA Orin NX (Sec. VI)","https:\u002F\u002Fgithub.com\u002Fethz-mrl\u002FOKVIS2-X","BSD-3-Clause (LICENSE file; GitHub API reports NOASSERTION)",[55,59],{"relation":56,"title":57,"doi_or_url":58},"preprint","OKVIS2-X arXiv v1 (accepted version, CC BY 4.0)","https:\u002F\u002Farxiv.org\u002Fabs\u002F2510.04612",{"relation":60,"title":61,"doi_or_url":52},"code_release","ethz-mrl\u002FOKVIS2-X",{"id":5,"kind":63,"shortName":7,"title":8,"authors":64,"year":9,"venue":69,"venueType":70,"publisher":71,"volumeIssuePages":72,"doi":73,"arxivId":74,"url":75,"firstPublicDate":76,"publicationStatus":16,"metadataStatus":77,"fulltextStatus":15,"era":10,"classicReason":78,"codeUrl":52,"cluster":11,"topics":79,"mdpi":81,"verification":82,"label":6,"fulltextRoute":83,"versionRead":84,"addedByCensus":85},"method",[65,66,67,68],"Simon Boche","Jaehyung Jung","Sebastián Barbas Laina","Stefan Leutenegger","IEEE Transactions on Robotics","journal","IEEE","41, pp. 6064-6083","10.1109\u002Ftro.2025.3619051","2510.04612","https:\u002F\u002Fdoi.org\u002F10.1109\u002FTRO.2025.3619051","2025-10-06","metadata_verified","not_applicable",[11,80],"C08",false,"corrected","arXiv","arXiv v1 (2025-10-06), T-RO accepted version (August 2025, Special Issue on Visual SLAM); IEEE version of record not read",true,[87,94,97,104,109,115,119,122,127,130,136,140,144,148,151,154],{"category":88,"model":89,"canonical":89,"role":90,"dataset":91,"specs":92,"locator":93},"compute","i7-13700","compute for runtime",null,"desktop CPU; all timings in Sec. VI-H","Sec. VI-H",{"category":88,"model":95,"canonical":95,"role":90,"dataset":91,"specs":96,"locator":93},"NVIDIA Orin NX (on a drone)","system including networks and depth or LiDAR integration can run with parameter trade-offs (supplementary video; not quantified)",{"category":98,"model":99,"canonical":99,"role":100,"dataset":101,"specs":102,"locator":103},"stereo_camera","EuRoC MAV stereo camera (model not reported)","dataset sensor","EuRoC MAV","stereo images with IMU on a drone","Sec. VI-B",{"category":105,"model":106,"canonical":106,"role":107,"dataset":101,"specs":108,"locator":103},"other","EuRoC Vicon-room reference point clouds","reference or ground truth","mm-level accurate point clouds used for mesh accuracy and completeness",{"category":110,"model":111,"canonical":111,"role":100,"dataset":112,"specs":113,"locator":114},"camera","Hilti-Oxford handheld rig cameras (5 cameras; model not reported)","Hilti-Oxford","all 5 used by the estimator, 2 front cameras for the stereo network, front-left for MVS","Sec. VI-C",{"category":116,"model":117,"canonical":117,"role":100,"dataset":112,"specs":118,"locator":114},"lidar","Hilti-Oxford 32-channel LiDAR (model not reported)","32-channel point clouds on the handheld device; far plane 30 m in mapping",{"category":105,"model":120,"canonical":120,"role":107,"dataset":112,"specs":121,"locator":114},"Hilti-Oxford ground truth (sparse control positions and mm-accurate dense point clouds)","sparse ground-truth positions for scoring; dense point clouds for exp04 to exp06",{"category":98,"model":123,"canonical":123,"role":100,"dataset":124,"specs":125,"locator":126},"VBR stereo camera (model not reported; high resolution)","VBR (Vision Benchmark in Rome)","with IMU, LiDAR and ground-truth poses","Sec. VI-D",{"category":116,"model":128,"canonical":128,"role":100,"dataset":124,"specs":129,"locator":126},"VBR LiDAR (model not reported)","far plane limited to 30 m",{"category":131,"model":132,"canonical":132,"role":100,"dataset":133,"specs":134,"locator":135},"gnss","ZED-F9P GNSS sensor","GVINS-Dataset","raw measurements and RTK solutions provided by the GVINS dataset; SPP computed with RTKLIB","Sec. VI-E",{"category":88,"model":137,"canonical":137,"role":90,"dataset":91,"specs":138,"locator":139},"RTX-3080 10 GB","desktop GPU running the stereo and MVS networks and depth fusion; Ours-vid uses 3.51 GB GPU memory for the networks","Sec. VI-H; Fig. 12 note",{"category":141,"model":142,"canonical":142,"role":100,"dataset":101,"specs":143,"locator":103},"imu","EuRoC MAV IMU (model not reported in this paper)","IMU measurements recorded with the stereo images by a drone",{"category":141,"model":145,"canonical":145,"role":100,"dataset":112,"specs":146,"locator":147},"Hilti-Oxford handheld device IMU (model not reported in this paper)","IMU measurements; known camera-IMU time offset accounted for","Secs. VI, VI-C",{"category":141,"model":149,"canonical":149,"role":100,"dataset":124,"specs":150,"locator":126},"VBR IMU (model not reported in this paper)","IMU measurements; the driving sequences Campus* and Ciampino* contain episodes of missing IMU measurements",{"category":98,"model":152,"canonical":152,"role":100,"dataset":133,"specs":153,"locator":135},"GVINS-Dataset stereo camera (model not reported in this paper)","stereo camera recorded together with an IMU and a ZED-F9P GNSS sensor",{"category":141,"model":155,"canonical":155,"role":100,"dataset":133,"specs":156,"locator":135},"GVINS-Dataset IMU (model not reported in this paper)","IMU recorded with the stereo camera and the ZED-F9P GNSS sensor",[158,171,179,189],{"refId":5,"refLabel":6,"fig":159,"whatZh":160,"license":161,"licenseUrl":162,"sourceUrl":163,"src":164,"width":165,"height":166,"thumb":167,"thumbWidth":168,"thumbHeight":169,"modified":170},"Fig. 1","VBR Spagna 序列的三維重建：上為 LiDAR 版、下為深度網路版，黑線為估計軌跡，不同顏色代表不同子地圖","CC BY 4.0 (arXiv v1)","http:\u002F\u002Fcreativecommons.org\u002Flicenses\u002Fby\u002F4.0\u002F","https:\u002F\u002Farxiv.org\u002Fhtml\u002F2510.04612v1\u002Ffigures\u002Ftop-figure-vertical.png","\u002Ffigure-files\u002Fokvis2x2025\u002Ffig-1.webp",1400,1580,"\u002Ffigure-files\u002Fokvis2x2025\u002Ffig-1.thumb.webp",480,542,"resized to at most 1400 px wide and converted to WebP",{"refId":5,"refLabel":6,"fig":172,"whatZh":173,"license":161,"licenseUrl":162,"sourceUrl":174,"src":175,"width":165,"height":176,"thumb":177,"thumbWidth":168,"thumbHeight":178,"modified":170},"Fig. 9","由體積子地圖產生的三維重建，左為學習式深度、右為 LiDAR，上排 Hilti exp06、下排 VBR Ciampino0","https:\u002F\u002Farxiv.org\u002Fhtml\u002F2510.04612v1\u002Ffigures\u002Fevaluation\u002Fmeshes-mod.png","\u002Ffigure-files\u002Fokvis2x2025\u002Ffig-9.webp",729,"\u002Ffigure-files\u002Fokvis2x2025\u002Ffig-9.thumb.webp",250,{"refId":5,"refLabel":6,"fig":180,"whatZh":181,"license":161,"licenseUrl":162,"sourceUrl":182,"src":183,"width":184,"height":185,"thumb":186,"thumbWidth":168,"thumbHeight":187,"modified":188},"Fig. 7","EuRoC V1_02 重建網格以網格到點誤差著色，比較 OKVIS2-X 深度版與 SimpleMapping","https:\u002F\u002Farxiv.org\u002Fhtml\u002F2510.04612v1\u002Fv102_meshes.png","\u002Ffigure-files\u002Fokvis2x2025\u002Ffig-7.webp",1103,536,"\u002Ffigure-files\u002Fokvis2x2025\u002Ffig-7.thumb.webp",233,"converted to WebP",{"refId":5,"refLabel":6,"fig":190,"whatZh":191,"license":161,"licenseUrl":162,"sourceUrl":192,"src":193,"width":165,"height":194,"thumb":195,"thumbWidth":168,"thumbHeight":196,"modified":170},"Fig. 10","VBR Campus1 所有子地圖疊合的重建與加入 GNSS 的即時軌跡，標示 GNSS 可用與中斷區段","https:\u002F\u002Farxiv.org\u002Fhtml\u002F2510.04612v1\u002Ffigures\u002Fevaluation\u002Ffigure-gnss-combined.png","\u002Ffigure-files\u002Fokvis2x2025\u002Ffig-10.webp",462,"\u002Ffigure-files\u002Fokvis2x2025\u002Ffig-10.thumb.webp",158,1790510665485]