OKVIS2-X
OKVIS2-X 以關鍵影格式視覺慣性 SLAM(OKVIS2)為核心,可選擇加入深度網路估計的稠密深度、LiDAR 或 GNSS。系統把 supereight2 體素占據子地圖綁定在關鍵影格上,並以影格對地圖、地圖對地圖的占據對齊殘差把子地圖與狀態估計緊耦合,使迴圈閉合或 GNSS 修正後地圖可隨位姿一起更新,形成全域一致、可直接用於導航的稠密地圖。GNSS 以可觀測性準則處理座標對齊與訊號中斷,並支援相機外參線上校正。
本頁內容
Extends OKVIS2 visual-inertial SLAM with keyframe-anchored volumetric occupancy submaps built from learned depth or LiDAR and tightly coupled by occupancy alignment factors, plus optional GNSS with observability-aware alignment and online extrinsic calibration, giving globally consistent dense maps at large scale.
技術屬性
欄位內容為文獻擷取紀錄的原文用語(英文),以原文為據;「未查證」表示本研究尚未讀到該資訊,不代表該方法不具備此能力。
| 感測輸入 | one or more cameras、IMU、optional learned depth (stereo and multi-view stereo networks)、optional LiDAR、optional GNSS |
|---|---|
| 原文測試平台 | GVINS dataset complex_environment sequence (carrying platform not described in this paper) |
| 狀態估計 | OKVIS2 keyframe-based visual-inertial SLAM (BRISK features, realtime sliding-window estimator plus asynchronous full-graph optimization with posegraph edges from marginalized landmarks, DBoW2 loop closure) extended with volumetric occupancy submaps anchored to keyframes and tightly coupled through frame-to-map and map-to-map occupancy alignment factors (Tukey robustifier), optional GNSS position factors and online camera-IMU extrinsic calibration (Secs. IV-V) |
| 資料關聯 | BRISK keypoint matching for landmarks; dense submap alignment by occupancy residuals of depth or LiDAR points against submaps; sky segmentation (Fast-SCNN) to remove sky pixels from depth (Sec. V) |
| 時間表示 | discrete multi-camera frames with IMU pre-integration between states; LiDAR points and GNSS measurements are related to states through IMU propagation (Sec. V) |
| 去畸變 | not described as a separate step |
| 迴圈閉合 | yes, DBoW2 place recognition with full-graph optimization (Sec. V) |
| 全域最佳化 | asynchronous full-graph optimization and optional final full bundle adjustment ('-ba' variants); GNSS alignment with a yaw-observability criterion and global re-alignment after GNSS dropouts (Sec. V) |
| 地圖表示 | supereight2 volumetric occupancy submaps anchored to keyframes, meshable into a globally consistent map (Figs. 1, 9, 10) |
| 先驗資訊 | camera intrinsics and initial extrinsics; networks fine-tuned on synthetic data |
| 可輸出幾何 | trajectory (causal, non-causal and full-BA variants) and dense occupancy submaps or meshes |
| 計算需求 | desktop with Intel i7-13700 and RTX-3080 10 GB; Ours-vi up to 47 Hz on EuRoC MH05; MVS network at 13 Hz on 8-view images; GPU memory 3.51 GB for the networks; also runs on a drone with NVIDIA Orin NX (Sec. VI) |
使用設備
原文使用的感測器、運算硬體與載具(equipment)。型號保留原文寫法,連結到設備頁中同一型號的歸併名稱;角色依原文用途分為方法輸入、資料集感測器、執行運算平台、參考或真值量測(reference or ground truth)與比較對象設備。
| 類別 | 型號(原文寫法) | 角色 | 資料集 | 原文規格 | 出處 |
|---|---|---|---|---|---|
| LiDAR | Hilti-Oxford 32-channel LiDAR (model not reported) | 資料集感測器 | Hilti-Oxford | 32-channel point clouds on the handheld device; far plane 30 m in mapping | (Boche et al., 2025, Sec. VI-C) |
| LiDAR | VBR LiDAR (model not reported) | 資料集感測器 | VBR (Vision Benchmark in Rome) | far plane limited to 30 m | (Boche et al., 2025, Sec. VI-D) |
| 慣性量測單元(IMU) | EuRoC MAV IMU (model not reported in this paper) | 資料集感測器 | EuRoC MAV | IMU measurements recorded with the stereo images by a drone | (Boche et al., 2025, Sec. VI-B) |
| 慣性量測單元(IMU) | Hilti-Oxford handheld device IMU (model not reported in this paper) | 資料集感測器 | Hilti-Oxford | IMU measurements; known camera-IMU time offset accounted for | (Boche et al., 2025, Secs. VI, VI-C) |
| 慣性量測單元(IMU) | VBR IMU (model not reported in this paper) | 資料集感測器 | VBR (Vision Benchmark in Rome) | IMU measurements; the driving sequences Campus* and Ciampino* contain episodes of missing IMU measurements | (Boche et al., 2025, Sec. VI-D) |
| 慣性量測單元(IMU) | GVINS-Dataset IMU (model not reported in this paper) | 資料集感測器 | GVINS-Dataset | IMU recorded with the stereo camera and the ZED-F9P GNSS sensor | (Boche et al., 2025, Sec. VI-E) |
| GNSS 接收器 | ZED-F9P GNSS sensor | 資料集感測器 | GVINS-Dataset | raw measurements and RTK solutions provided by the GVINS dataset; SPP computed with RTKLIB | (Boche et al., 2025, Sec. VI-E) |
| 相機 | Hilti-Oxford handheld rig cameras (5 cameras; model not reported) | 資料集感測器 | Hilti-Oxford | all 5 used by the estimator, 2 front cameras for the stereo network, front-left for MVS | (Boche et al., 2025, Sec. VI-C) |
| 雙目相機 | EuRoC MAV stereo camera (model not reported) | 資料集感測器 | EuRoC MAV | stereo images with IMU on a drone | (Boche et al., 2025, Sec. VI-B) |
| 雙目相機 | VBR stereo camera (model not reported; high resolution) | 資料集感測器 | VBR (Vision Benchmark in Rome) | with IMU, LiDAR and ground-truth poses | (Boche et al., 2025, Sec. VI-D) |
| 雙目相機 | GVINS-Dataset stereo camera (model not reported in this paper) | 資料集感測器 | GVINS-Dataset | stereo camera recorded together with an IMU and a ZED-F9P GNSS sensor | (Boche et al., 2025, Sec. VI-E) |
| 運算硬體 | i7-13700 | 執行運算平台 | 未標示 | desktop CPU; all timings in Sec. VI-H | (Boche et al., 2025, Sec. VI-H) |
| 運算硬體 | NVIDIA Orin NX (on a drone) | 執行運算平台 | 未標示 | system including networks and depth or LiDAR integration can run with parameter trade-offs (supplementary video; not quantified) | (Boche et al., 2025, Sec. VI-H) |
| 運算硬體 | RTX-3080 10 GB | 執行運算平台 | 未標示 | desktop GPU running the stereo and MVS networks and depth fusion; Ours-vid uses 3.51 GB GPU memory for the networks | (Boche et al., 2025, Sec. VI-H; Fig. 12 note) |
| 其他 | EuRoC Vicon-room reference point clouds | 參考或真值量測 | EuRoC MAV | mm-level accurate point clouds used for mesh accuracy and completeness | (Boche et al., 2025, Sec. VI-B) |
| 其他 | Hilti-Oxford ground truth (sparse control positions and mm-accurate dense point clouds) | 參考或真值量測 | Hilti-Oxford | sparse ground-truth positions for scoring; dense point clouds for exp04 to exp06 | (Boche et al., 2025, Sec. VI-C) |
論文圖片
只收錄原文以開放授權(open license)釋出的圖片,並依授權條件標示出處、圖號、授權與修改方式。

Fig. 1VBR Spagna 序列的三維重建:上為 LiDAR 版、下為深度網路版,黑線為估計軌跡,不同顏色代表不同子地圖
出處:Boche et al., 2025,Fig. 1。授權:CC BY 4.0 (arXiv v1)。原始圖檔。修改:縮小至寬度不超過 1400 px,並轉存為 WebP 格式。

Fig. 9由體積子地圖產生的三維重建,左為學習式深度、右為 LiDAR,上排 Hilti exp06、下排 VBR Ciampino0
出處:Boche et al., 2025,Fig. 9。授權:CC BY 4.0 (arXiv v1)。原始圖檔。修改:縮小至寬度不超過 1400 px,並轉存為 WebP 格式。

Fig. 7EuRoC V1_02 重建網格以網格到點誤差著色,比較 OKVIS2-X 深度版與 SimpleMapping
出處:Boche et al., 2025,Fig. 7。授權:CC BY 4.0 (arXiv v1)。原始圖檔。修改:轉存為 WebP 格式。

Fig. 10VBR Campus1 所有子地圖疊合的重建與加入 GNSS 的即時軌跡,標示 GNSS 可用與中斷區段
出處:Boche et al., 2025,Fig. 10。授權:CC BY 4.0 (arXiv v1)。原始圖檔。修改:縮小至寬度不超過 1400 px,並轉存為 WebP 格式。
作者報告的優勢與限制
優勢
- On VBR (1.0 to 9.0 km sequences) Ours-vil-nc average ATE RMSE 0.829 m versus 2.475 m for FAST-LIVO; visual-inertial Ours-vid-nc 2.210 m versus 5.572 m for ORB-SLAM3 and 28.770 m for OpenVINS (Table VI)
- Hilti-Oxford mesh accuracy 0.029 m and completeness 78.33 % with LiDAR, 0.042 m and 64.27 % with learned depth, at a 0.2 m threshold (Table V)
- Hilti 2022 challenge score 438.10 for Ours-vil-ba, above FAST-LIVO2 (402.93) and VILENS (325.77) but below LiDAR-inertial Wildcat (563.79) (Table IV)
- Faster than ORB-SLAM3 on EuRoC MH05 (38.1 versus 64.7 ms per frame) and on par on Hilti exp06 (Sec. VI)
- Open-source with configurable sensor setups
限制
- Dense mapping raises memory: up to 25.5 GB for Ours-vil on VBR Campus1 (Table IX)
- Higher runtime than ORB-SLAM3 on VBR Campus1 (107.3 versus 69.1 ms per frame) (Sec. VI)
- Long corridors (Hilti exp07) are degenerate for LiDAR alignment (Sec. VI)
- Learned depth networks require a GPU
營建工程相關證據
OKVIS2-X 直接輸出全域一致的稠密占據地圖與網格,並在 Hilti-Oxford 資料上以地面真值點雲評估網格精度(LiDAR 版本平均 2.9 cm,完整度 78 %);依 Zhang et al., 2023c 的資料集說明,Hilti-Oxford 包含施工現場序列,但本文未說明 exp04 至 exp06 的場景類型。這是少數同時提供軌跡與建圖精度的多感測器系統,對施工現場掃描具參考價值(推論)。
原文驗證環境:公開基準、獨立參考量測、跨場域
報告的性能數據
性能數據仍在分批查證,目前尚未收錄此方法的報告值。
來源
Boche et al., 2025
(2025)OKVIS2-X: Open Keyframe-Based Visual-Inertial SLAM Configurable With Dense Depth or LiDAR, and GNSSIEEE Transactions on Robotics, 41, pp. 6064-6083
DOI 10.1109/tro.2025.3619051arXiv 2510.04612程式碼
同儕審查已出版已讀全文近十年查證後修正
相關版本
- 預印本:OKVIS2-X arXiv v1 (accepted version, CC BY 4.0) https://arxiv.org/abs/2510.04612
- 程式碼釋出:ethz-mrl/OKVIS2-X https://github.com/ethz-mrl/OKVIS2-X
程式碼:https://github.com/ethz-mrl/OKVIS2-X(授權:BSD-3-Clause (LICENSE file; GitHub API reports NOASSERTION))。有公開程式碼不等於已被重現,也不代表目前版本與論文版本相同。