Deep feature-enhanced LVIO (tunnel)
此研究針對隧道幾何特徵稀疏、結構重複而導致光達里程計退化的問題,提出光達、視覺與慣性融合的里程計。光達端以曲率區分邊緣與平面特徵,採點對線與點對面配準;將資訊矩陣求逆得到共變異數後,分別對旋轉與平移子區塊做特徵分解,以經驗門檻判定六自由度中哪些方向退化。視覺端以 SuperPoint 搭配自適應門檻擷取特徵、以 LightGlue 匹配,並結合 IMU 預積分,但不做後端最佳化。條件式擴展卡爾曼濾波器只在偵測到退化時,以選擇矩陣保留退化方向上的視覺慣性資訊來更新狀態。系統只做里程計,不含迴圈閉合;驗證完全使用 MIT 校園隧道(KMCT)與 WHU-Helmet 隧道、地鐵兩組公開資料。
本頁內容
Detects 6-DoF degeneracy of feature-based LiDAR odometry in tunnels from the covariance of the rotation and translation blocks, and fuses a SuperPoint and LightGlue visual-inertial odometry only along the degenerate directions through a conditional EKF.
技術屬性
欄位內容為文獻擷取紀錄的原文用語(英文),以原文為據;「未查證」表示本研究尚未讀到該資訊,不代表該方法不具備此能力。
| 感測輸入 | 3D LiDAR (mechanical spinning model assumed in the LO derivation; Velodyne on KMCT, Livox on WHU-Helmet)、camera (Intel RealSense D455 RGB-D on KMCT; helmet cameras on WHU-Helmet)、IMU (preintegrated in the VIO) |
|---|---|
| 原文測試平台 | no own platform; offline evaluation on public datasets recorded by Clearpath Jackal ground robots (KMCT) and a helmet-mounted rig (WHU-Helmet) |
| 狀態估計 | Conditional EKF: LiDAR odometry runs continuously; when covariance-based detection flags degenerate rotation or translation directions, a selection matrix keeps only the visual-inertial information along those directions and the state is updated with a FAST-LIO style Kalman gain; the VIO itself has no back-end optimization (Secs. 3.2, 3.3) |
| 資料關聯 | LiDAR: curvature from five horizontal neighbours on each side separates edge and planar points, matched point-to-line and point-to-plane to global feature maps and solved by Gauss-Newton; degeneracy from eigen-decomposition of the rotation and translation blocks of the inverted information matrix against empirical thresholds; visual: SuperPoint features with an adaptive score threshold and LightGlue matching inside an ORB-SLAM3 style tracking thread (Secs. 3.1 to 3.2) |
| 時間表示 | discrete scan poses; constant-velocity camera prediction; IMU preintegration between frames with high-rate IMU propagation of the output pose (Secs. 3.1.1, 3.2) |
| 去畸變 | linear interpolation of the inter-scan transform across the sweep by point index (Eq. 3, Sec. 3.1.1) |
| 迴圈閉合 | none; the authors argue loop closure is often infeasible in tunnels and choose an odometry design (Secs. 2.3, 3.3.1) |
| 全域最佳化 | none (no factor graph or bundle adjustment) |
| 地圖表示 | point cloud map built by accumulating registered scans, with global edge and planar feature maps used for matching (Secs. 3, 3.1.1) |
| 先驗資訊 | none (official pre-trained SuperPoint and LightGlue models; no prior map) |
| 可輸出幾何 | dense LiDAR point cloud map evaluated by Mean Map Entropy (local consistency only) and, on one KMCT sequence, by voxel error against the ground-truth map (Secs. 4.2.2, 4.3, Fig. 9) |
| 計算需求 | Intel i9-12900H CPU, NVIDIA 3070 Ti GPU, 32 GB RAM, Ubuntu 20.04 with CUDA, cuDNN and ONNX inference: 25.4 to 29.0 ms per frame in total (VIO about 20 to 23 ms, LO about 5 to 6 ms, EKF about 0.1 ms) versus 28.8 to 35.8 ms for R3LIVE++ and 49.7 to 67.6 ms for LVI-SAM (Table 5) |
使用設備
原文使用的感測器、運算硬體與載具(equipment)。型號保留原文寫法,連結到設備頁中同一型號的歸併名稱;角色依原文用途分為方法輸入、資料集感測器、執行運算平台、參考或真值量測(reference or ground truth)與比較對象設備。
| 類別 | 型號(原文寫法) | 角色 | 資料集 | 原文規格 | 出處 |
|---|---|---|---|---|---|
| LiDAR | Velodyne LiDAR (model 原文未報告) | 資料集感測器 | Kimera-Multi Campus-Tunnel (KMCT) | 原文未報告 | (Yan et al., 2026a, Sec. 4.1) |
| LiDAR | LIVOX LiDAR (model 原文未報告) | 資料集感測器 | WHU-Helmet (WHUH) | 原文未報告 | (Yan et al., 2026a, Sec. 4.1) |
| 慣性量測單元(IMU) | IMU (model 原文未報告)歸入:IMU (model not reported) | 資料集感測器 | WHU-Helmet (WHUH) | 原文未報告 | (Yan et al., 2026a, Sec. 4.1) |
| GNSS 接收器 | GNSS receiver (model 原文未報告)歸入:GNSS (receiver model not reported) | 資料集感測器 | WHU-Helmet (WHUH) | 原文未報告 | (Yan et al., 2026a, Sec. 4.1) |
| 相機 | cameras (models 原文未報告) | 資料集感測器 | WHU-Helmet (WHUH) | Tunnel sequence 12,304 images in 1403 s; Subway 15,685 images in 1580 s | (Yan et al., 2026a, Sec. 4.1) |
| RGB-D 相機 | Intel RealSense D455 | 資料集感測器 | Kimera-Multi Campus-Tunnel (KMCT) | RGB-D camera on each robot | (Yan et al., 2026a, Sec. 4.1) |
| 載具平台 | Clearpath Jackal mobile robots (eight) | 資料集感測器 | Kimera-Multi Campus-Tunnel (KMCT) | collectively travelled 6753 m in about 30 min in the MIT campus tunnel | (Yan et al., 2026a, Sec. 4.1) |
| 載具平台 | helmet (per the dataset name and the title of ref. [21]) | 資料集感測器 | WHU-Helmet (WHUH) | Wuhan University helmet-based multisensor dataset | (Yan et al., 2026a, Sec. 4.1, ref. [21]) |
| 運算硬體 | Intel i9-12900H CPU with NVIDIA 3070 Ti | 執行運算平台 | 未標示 | 32 GB RAM, Ubuntu 20.04, CUDA 11.3, cuDNN 8.9.6, ONNX 1.16.3 | (Yan et al., 2026a, Sec. 4.1) |
作者報告的優勢與限制
優勢
- Lowest average ATE on both datasets: 3.21 m on KMCT (Fast-LIO2 3.26 m) and 3.74 m on WHU-Helmet (BALM 4.46 m, a 16.14% ATE reduction that the abstract labels as RPE) (Tables 1 to 2, Sec. 4.3)
- Average RPE 1.77% on KMCT and 1.93% on WHU-Helmet per 100 m, 7.33% and 17.87% lower than Fast-LIO2 (Sec. 4.3, Tables 1 to 2)
- Lowest average Mean Map Entropy, -8.58 on KMCT and -7.03 on WHU-Helmet (Tables 3 to 4); 3σ voxel error 0.12 versus 0.18 for LVI-SAM on a KMCT partial map (Sec. 4.3, Fig. 9)
- About 25 to 29 ms per frame, 13.98% faster on average than R3LIVE++ (Table 5)
- Removing the degeneracy module raises ATE by 18.00% and RPE by 24.37% on three KMCT sequences (Sec. 5)
限制
- Dynamic objects in tunnels are not considered; multi-robot collaborative mapping is left to future work (Sec. 6)
- Not best on every sequence: KMCT 07_ac2-005 ATE 2.89 m versus 2.70 m for Fast-LIO2, 07_sob-002 1.18 m versus 1.13 m for COIN-LIO, 07_tho-001 8.48 m versus 8.37 m for Fast-LIO2 (Table 1)
- Each sequence was run five times and the best trial reported (Sec. 4.1)
- Degeneracy thresholds are set empirically (Sec. 3.1.2)
- Map quality is assessed mainly by MME, which captures local consistency only (Sec. 4.2.2)
- (inference) Only public datasets are used, absolute ATE is several metres, and code is not released (data on request)
營建工程相關證據
論文以隧道營運維護與機器人巡檢為應用情境(摘要與引言明確提及營運維護),刊於本文目標期刊。驗證資料為 MIT 校園地下隧道(KMCT,八台 Clearpath Jackal 機器人,總長 6753 m)與 WHU-Helmet 的 Tunnel(790.32 m)與 Subway(854.24 m)序列,並非施工中隧道;作者未自行蒐集資料,軌跡真值由資料集提供,參考量測方式未在文中說明。
原文驗證環境:地下或隧道、公開基準
報告的性能數據
以下是原文作者報告的性能數值(author-reported results),不是本研究重新量測的結果。每張圖只並列同一個比較組(comparison group,同一張表、同一組實驗設定)內的方法;不同比較組之間的數值不可直接比較,也不構成排名。
本方法共出現在 8 個比較組,合計 35 筆紀錄。以下列出本方法紀錄最多的 4 組,其餘 4 組列在最後,並連到性能比較頁。
Yan et al., 2026a · Table 1 本方法 9 筆
表格設定(擷取紀錄原文):ATE (RMSE, m) on KMCT; each ROS bag run five times and the best trial reported; 'Failed' = failure to run (Yan et al., 2026a, Table 1)
ATE (m), RMSE of global absolute trajectory error,Kimera-Multi Campus-Tunnel (KMCT) · 07_ac2-005
只並列這張表在相同設定下報告的方法;以「本方法:」開頭者為本頁方法。失敗、未執行與未報告以標記呈現,不是 0。
- 失敗
按 Tab 進入圖表後,用上下方向鍵逐一瀏覽各類別,Esc 關閉提示框;也可開啟表格檢視閱讀全部數值。
這些是 Yan et al., 2026a 在此表設定下報告的數值(author-reported results),只能在同一個比較組內對照,不代表方法在其他資料或設定下的表現。
資料來源作者報告值(Yan et al., 2026a, Table 1)
| 方法(原文寫法) | 報告值 | 出處 |
|---|---|---|
| ORB-SLAM3 | 無數值失敗註記(擷取紀錄):failed | (Yan et al., 2026a, Table 1) |
| VINS-Mono | 3.96 m | (Yan et al., 2026a, Table 1) |
| LeGO-LOAM | 3.63 m | (Yan et al., 2026a, Table 1) |
| LIO-SAM | 3.41 m | (Yan et al., 2026a, Table 1) |
| Fast-LIO2 | 2.7 m | (Yan et al., 2026a, Table 1) |
| LVI-SAM | 2.77 m | (Yan et al., 2026a, Table 1) |
| COIN-LIO | 2.71 m | (Yan et al., 2026a, Table 1) |
| VINS-FEN | 3.05 m | (Yan et al., 2026a, Table 1) |
| Proposed本方法原文提出 | 2.89 m | (Yan et al., 2026a, Table 1) |
Yan et al., 2026a · Table 5 本方法 9 筆
指標Time consumption (ms)
表格設定(擷取紀錄原文):Time consumption per frame; proposed method component times (VIO, LO, EKF) not extracted, only their sum (Yan et al., 2026a, Table 5)
Time consumption (ms),Kimera-Multi Campus-Tunnel (KMCT) · ac2-005
只並列這張表在相同設定下報告的方法;以「本方法:」開頭者為本頁方法。失敗、未執行與未報告以標記呈現,不是 0。
按 Tab 進入圖表後,用上下方向鍵逐一瀏覽各類別,Esc 關閉提示框;也可開啟表格檢視閱讀全部數值。
這些是 Yan et al., 2026a 在此表設定下報告的數值(author-reported results),只能在同一個比較組內對照,不代表方法在其他資料或設定下的表現。
資料來源作者報告值(Yan et al., 2026a, Table 5)
| 方法(原文寫法) | 報告值 | 出處 |
|---|---|---|
| This work (Sum of VIO, LO and EKF)本方法原文提出硬體:Intel i9-12900H CPU, NVIDIA 3070 Ti, 32 GB RAM, Ubuntu 20.04 | 27.51 ms | (Yan et al., 2026a, Table 5) |
| LVI-SAM硬體:Intel i9-12900H CPU, NVIDIA 3070 Ti, 32 GB RAM, Ubuntu 20.04 | 51.36 ms | (Yan et al., 2026a, Table 5) |
| R3LIVE++硬體:Intel i9-12900H CPU, NVIDIA 3070 Ti, 32 GB RAM, Ubuntu 20.04 | 30.53 ms | (Yan et al., 2026a, Table 5) |
Yan et al., 2026a · Table 6 本方法 6 筆
資料集與序列Kimera-Multi Campus-Tunnel (KMCT) · Average (3 sequences)
表格設定(擷取紀錄原文):Ablation of deep-feature VIO and CEKF; averages over KMCT 07_api-003, 07_sob-002 and 07_spl-007; per-sequence rows not extracted (Yan et al., 2026a, Table 6)
Average ATE (m),Kimera-Multi Campus-Tunnel (KMCT) · Average (3 sequences)
只並列這張表在相同設定下報告的方法;以「本方法:」開頭者為本頁方法。失敗、未執行與未報告以標記呈現,不是 0。
按 Tab 進入圖表後,用上下方向鍵逐一瀏覽各類別,Esc 關閉提示框;也可開啟表格檢視閱讀全部數值。
這些是 Yan et al., 2026a 在此表設定下報告的數值(author-reported results),只能在同一個比較組內對照,不代表方法在其他資料或設定下的表現。
資料來源作者報告值(Yan et al., 2026a, Table 6)
| 方法(原文寫法) | 報告值 | 出處 |
|---|---|---|
| Standard VIO + Base (VINS-Mono front end)本方法 | 1.84 m | (Yan et al., 2026a, Table 6) |
| Base + EKF (standard EKF without selection)本方法 | 1.59 m | (Yan et al., 2026a, Table 6) |
| Proposed method本方法原文提出 | 1.24 m | (Yan et al., 2026a, Table 6) |
Yan et al., 2026a · Table 2 本方法 4 筆
表格設定(擷取紀錄原文):ATE (RMSE, m) on WHU-Helmet with dataset ground-truth trajectory; LiDAR baselines use variants adapted to the Livox configuration; 'Failed' = failure to run (Yan et al., 2026a, Table 2)
ATE (m),WHU-Helmet (WHUH) · Tunnel
只並列這張表在相同設定下報告的方法;以「本方法:」開頭者為本頁方法。失敗、未執行與未報告以標記呈現,不是 0。
- 失敗
按 Tab 進入圖表後,用上下方向鍵逐一瀏覽各類別,Esc 關閉提示框;也可開啟表格檢視閱讀全部數值。
這些是 Yan et al., 2026a 在此表設定下報告的數值(author-reported results),只能在同一個比較組內對照,不代表方法在其他資料或設定下的表現。
資料來源作者報告值(Yan et al., 2026a, Table 2)
| 方法(原文寫法) | 報告值 | 出處 |
|---|---|---|
| ORB-SLAM3 | 5.92 m | (Yan et al., 2026a, Table 2) |
| VINS-Mono | 無數值失敗註記(擷取紀錄):failed | (Yan et al., 2026a, Table 2) |
| LOAM | 無數值失敗註記(擷取紀錄):failed | (Yan et al., 2026a, Table 2) |
| BALM | 4.53 m | (Yan et al., 2026a, Table 2) |
| LIO-Mapping | 5.95 m | (Yan et al., 2026a, Table 2) |
| Fast-LIO2 | 4.21 m | (Yan et al., 2026a, Table 2) |
| R3live++ | 5.3 m | (Yan et al., 2026a, Table 2) |
| COIN-LIO | 4.05 m | (Yan et al., 2026a, Table 2) |
| VINS-FEN | 4.39 m | (Yan et al., 2026a, Table 2) |
| This work本方法原文提出 | 4.04 m | (Yan et al., 2026a, Table 2) |
其他比較組
來源
Yan et al., 2026a
(2026)Deep feature-enhanced LiDAR-visual-inertial odometry for robust mapping in tunnel environmentAutomation in Construction, 184: 106825
DOI 10.1016/j.autcon.2026.106825
同儕審查已出版已讀全文近十年查證後修正
相關版本
- 預印本:Deep feature-enhanced LiDAR-visual-inertial odometry for robotic inspection and mapping in degraded tunnel environments (SSRN) 10.2139/ssrn.5412661