Extends OKVIS2 visual-inertial SLAM with keyframe-anchored volumetric occupancy submaps built from learned depth or LiDAR and tightly coupled by occupancy alignment factors, plus optional GNSS with observability-aware alignment and online extrinsic calibration, giving globally consistent dense maps at large scale.

技術屬性

欄位內容為文獻擷取紀錄的原文用語(英文),以原文為據;「未查證」表示本研究尚未讀到該資訊,不代表該方法不具備此能力。

OKVIS2-X 的技術屬性
感測輸入one or more cameras、IMU、optional learned depth (stereo and multi-view stereo networks)、optional LiDAR、optional GNSS
原文測試平台GVINS dataset complex_environment sequence (carrying platform not described in this paper)
狀態估計OKVIS2 keyframe-based visual-inertial SLAM (BRISK features, realtime sliding-window estimator plus asynchronous full-graph optimization with posegraph edges from marginalized landmarks, DBoW2 loop closure) extended with volumetric occupancy submaps anchored to keyframes and tightly coupled through frame-to-map and map-to-map occupancy alignment factors (Tukey robustifier), optional GNSS position factors and online camera-IMU extrinsic calibration (Secs. IV-V)
資料關聯BRISK keypoint matching for landmarks; dense submap alignment by occupancy residuals of depth or LiDAR points against submaps; sky segmentation (Fast-SCNN) to remove sky pixels from depth (Sec. V)
時間表示discrete multi-camera frames with IMU pre-integration between states; LiDAR points and GNSS measurements are related to states through IMU propagation (Sec. V)
去畸變not described as a separate step
迴圈閉合yes, DBoW2 place recognition with full-graph optimization (Sec. V)
全域最佳化asynchronous full-graph optimization and optional final full bundle adjustment ('-ba' variants); GNSS alignment with a yaw-observability criterion and global re-alignment after GNSS dropouts (Sec. V)
地圖表示supereight2 volumetric occupancy submaps anchored to keyframes, meshable into a globally consistent map (Figs. 1, 9, 10)
先驗資訊camera intrinsics and initial extrinsics; networks fine-tuned on synthetic data
可輸出幾何trajectory (causal, non-causal and full-BA variants) and dense occupancy submaps or meshes
計算需求desktop with Intel i7-13700 and RTX-3080 10 GB; Ours-vi up to 47 Hz on EuRoC MH05; MVS network at 13 Hz on 8-view images; GPU memory 3.51 GB for the networks; also runs on a drone with NVIDIA Orin NX (Sec. VI)

使用設備

原文使用的感測器、運算硬體與載具(equipment)。型號保留原文寫法,連結到設備頁中同一型號的歸併名稱;角色依原文用途分為方法輸入、資料集感測器、執行運算平台、參考或真值量測(reference or ground truth)與比較對象設備。

原文使用的設備
類別型號(原文寫法)角色資料集原文規格出處
LiDARHilti-Oxford 32-channel LiDAR (model not reported)資料集感測器Hilti-Oxford32-channel point clouds on the handheld device; far plane 30 m in mapping(Boche et al., 2025, Sec. VI-C)
LiDARVBR LiDAR (model not reported)資料集感測器VBR (Vision Benchmark in Rome)far plane limited to 30 m(Boche et al., 2025, Sec. VI-D)
慣性量測單元(IMU)EuRoC MAV IMU (model not reported in this paper)資料集感測器EuRoC MAVIMU measurements recorded with the stereo images by a drone(Boche et al., 2025, Sec. VI-B)
慣性量測單元(IMU)Hilti-Oxford handheld device IMU (model not reported in this paper)資料集感測器Hilti-OxfordIMU measurements; known camera-IMU time offset accounted for(Boche et al., 2025, Secs. VI, VI-C)
慣性量測單元(IMU)VBR IMU (model not reported in this paper)資料集感測器VBR (Vision Benchmark in Rome)IMU measurements; the driving sequences Campus* and Ciampino* contain episodes of missing IMU measurements(Boche et al., 2025, Sec. VI-D)
慣性量測單元(IMU)GVINS-Dataset IMU (model not reported in this paper)資料集感測器GVINS-DatasetIMU recorded with the stereo camera and the ZED-F9P GNSS sensor(Boche et al., 2025, Sec. VI-E)
GNSS 接收器ZED-F9P GNSS sensor資料集感測器GVINS-Datasetraw measurements and RTK solutions provided by the GVINS dataset; SPP computed with RTKLIB(Boche et al., 2025, Sec. VI-E)
相機Hilti-Oxford handheld rig cameras (5 cameras; model not reported)資料集感測器Hilti-Oxfordall 5 used by the estimator, 2 front cameras for the stereo network, front-left for MVS(Boche et al., 2025, Sec. VI-C)
雙目相機EuRoC MAV stereo camera (model not reported)資料集感測器EuRoC MAVstereo images with IMU on a drone(Boche et al., 2025, Sec. VI-B)
雙目相機VBR stereo camera (model not reported; high resolution)資料集感測器VBR (Vision Benchmark in Rome)with IMU, LiDAR and ground-truth poses(Boche et al., 2025, Sec. VI-D)
雙目相機GVINS-Dataset stereo camera (model not reported in this paper)資料集感測器GVINS-Datasetstereo camera recorded together with an IMU and a ZED-F9P GNSS sensor(Boche et al., 2025, Sec. VI-E)
運算硬體i7-13700執行運算平台未標示desktop CPU; all timings in Sec. VI-H(Boche et al., 2025, Sec. VI-H)
運算硬體NVIDIA Orin NX (on a drone)執行運算平台未標示system including networks and depth or LiDAR integration can run with parameter trade-offs (supplementary video; not quantified)(Boche et al., 2025, Sec. VI-H)
運算硬體RTX-3080 10 GB執行運算平台未標示desktop GPU running the stereo and MVS networks and depth fusion; Ours-vid uses 3.51 GB GPU memory for the networks(Boche et al., 2025, Sec. VI-H; Fig. 12 note)
其他EuRoC Vicon-room reference point clouds參考或真值量測EuRoC MAVmm-level accurate point clouds used for mesh accuracy and completeness(Boche et al., 2025, Sec. VI-B)
其他Hilti-Oxford ground truth (sparse control positions and mm-accurate dense point clouds)參考或真值量測Hilti-Oxfordsparse ground-truth positions for scoring; dense point clouds for exp04 to exp06(Boche et al., 2025, Sec. VI-C)

論文圖片

只收錄原文以開放授權(open license)釋出的圖片,並依授權條件標示出處、圖號、授權與修改方式。

  • VBR Spagna 序列的三維重建:上為 LiDAR 版、下為深度網路版,黑線為估計軌跡,不同顏色代表不同子地圖

    Fig. 1VBR Spagna 序列的三維重建:上為 LiDAR 版、下為深度網路版,黑線為估計軌跡,不同顏色代表不同子地圖

    出處:Boche et al., 2025,Fig. 1。授權:CC BY 4.0 (arXiv v1)。原始圖檔。修改:縮小至寬度不超過 1400 px,並轉存為 WebP 格式。

  • 由體積子地圖產生的三維重建,左為學習式深度、右為 LiDAR,上排 Hilti exp06、下排 VBR Ciampino0

    Fig. 9由體積子地圖產生的三維重建,左為學習式深度、右為 LiDAR,上排 Hilti exp06、下排 VBR Ciampino0

    出處:Boche et al., 2025,Fig. 9。授權:CC BY 4.0 (arXiv v1)。原始圖檔。修改:縮小至寬度不超過 1400 px,並轉存為 WebP 格式。

  • EuRoC V1_02 重建網格以網格到點誤差著色,比較 OKVIS2-X 深度版與 SimpleMapping

    Fig. 7EuRoC V1_02 重建網格以網格到點誤差著色,比較 OKVIS2-X 深度版與 SimpleMapping

    出處:Boche et al., 2025,Fig. 7。授權:CC BY 4.0 (arXiv v1)。原始圖檔。修改:轉存為 WebP 格式。

  • VBR Campus1 所有子地圖疊合的重建與加入 GNSS 的即時軌跡,標示 GNSS 可用與中斷區段

    Fig. 10VBR Campus1 所有子地圖疊合的重建與加入 GNSS 的即時軌跡,標示 GNSS 可用與中斷區段

    出處:Boche et al., 2025,Fig. 10。授權:CC BY 4.0 (arXiv v1)。原始圖檔。修改:縮小至寬度不超過 1400 px,並轉存為 WebP 格式。

作者報告的優勢與限制

優勢

限制

營建工程相關證據

OKVIS2-X 直接輸出全域一致的稠密占據地圖與網格,並在 Hilti-Oxford 資料上以地面真值點雲評估網格精度(LiDAR 版本平均 2.9 cm,完整度 78 %);依 Zhang et al., 2023c 的資料集說明,Hilti-Oxford 包含施工現場序列,但本文未說明 exp04 至 exp06 的場景類型。這是少數同時提供軌跡與建圖精度的多感測器系統,對施工現場掃描具參考價值(推論)。

原文驗證環境:公開基準、獨立參考量測、跨場域

報告的性能數據

性能數據仍在分批查證,目前尚未收錄此方法的報告值。

來源

  • Boche et al., 2025

    Simon Boche, Jaehyung Jung, Sebastián Barbas Laina, Stefan Leutenegger(2025)OKVIS2-X: Open Keyframe-Based Visual-Inertial SLAM Configurable With Dense Depth or LiDAR, and GNSSIEEE Transactions on Robotics, 41, pp. 6064-6083

    同儕審查已出版已讀全文近十年查證後修正

回到方法圖鑑

選擇開啟Esc關閉