Semantics-aided incremental meshing (LIO + RGB)
此方法以 OneFormer 視覺基礎模型對每張 RGB 影像做全景分割,再利用 FAST-LIO2 的 IMU 狀態把標籤投影到已完成運動畸變校正(deskew)的 LiDAR 掃描點,並以遮罩侵蝕、邊界距離與深度不連續檢查剔除不可靠的投影。帶標籤的點進入改良的 TSDF:每個體素保存標籤直方圖,截斷距離依類別與量測距離調整(例如欄杆較窄、牆面較寬),最後以 Marching Cubes 產生帶語意的網格,可匯出為 USD 資產。作者刻意不用迴圈閉合,以免位姿跳動使增量式 TSDF 失效。
本頁內容
Projects OneFormer panoptic labels from RGB frames onto FAST-LIO2-deskewed LiDAR scans and fuses them into a Voxblox-style TSDF with per-voxel label histograms and class- and range-dependent truncation, then extracts a labelled marching-cubes mesh exportable to USD; loop closure is deliberately omitted.
技術屬性
欄位內容為文獻擷取紀錄的原文用語(英文),以原文為據;「未查證」表示本研究尚未讀到該資訊,不代表該方法不具備此能力。
| 感測輸入 | 3D LiDAR、IMU、monocular camera |
|---|---|
| 原文測試平台 | ["public multi-sensor datasets only (Oxford Spires Christ Church College sequence 2、NTU VIRAL NYA01)、the carrying platforms are not described in the paper"] |
| 狀態估計 | FAST-LIO2 tightly coupled LiDAR-inertial odometry (Sec. 3.1) |
| 資料關聯 | LiDAR registration inherited from FAST-LIO2 (nearest-neighbour search on ikd-Tree, Sec. 3.1); image-to-LiDAR label association by ego-motion-compensated pinhole-plus-distortion projection filtered by mask erosion, boundary-distance rejection and depth-discontinuity checks (Sec. 3.2); voxel labels by majority vote plus greedy segment-to-map-label overlap assignment (Sec. 3.3) |
| 時間表示 | IMU state forward-propagated from the LiDAR timestamp to the camera timestamp to form a dynamic LiDAR-to-camera transform, with per-point motion correction from the instantaneous camera velocity (Sec. 3.2, Eq. 1) |
| 去畸變 | labels are transferred onto point clouds deskewed by FAST-LIO2 (Sec. 3) |
| 迴圈閉合 | none (authors deliberately use LIO without loop closure to avoid discontinuous pose corrections that would require TSDF re-integration, Sec. 3.1) |
| 全域最佳化 | none |
| 地圖表示 | Voxblox-based TSDF with 0.1 m voxels, per-voxel label histogram with range-dependent log-normal confidence, class- and range-dependent truncation, base weighting either constant or inverse-square (1/z^2 of the point coordinate in the sensor frame; the mode used in the experiments is not stated), behind-surface weight dropoff, sparsity compensation factor 1.5 and weight cap 255, plus per-voxel semantic and geometric uncertainty scores; mesh via marching cubes |
| 先驗資訊 | vision foundation model labels (OneFormer) on RGB frames (learned prior) |
| 可輸出幾何 | semantically labelled triangle mesh (USD assets mentioned) |
| 計算需求 | 原文未報告 (no processor, GPU or timing figures are given; the pipeline is only stated to operate at LiDAR rate, Sec. 3) |
使用設備
原文使用的感測器、運算硬體與載具(equipment)。型號保留原文寫法,連結到設備頁中同一型號的歸併名稱;角色依原文用途分為方法輸入、資料集感測器、執行運算平台、參考或真值量測(reference or ground truth)與比較對象設備。
| 類別 | 型號(原文寫法) | 角色 | 資料集 | 原文規格 | 出處 |
|---|---|---|---|---|---|
| LiDAR | 3D LiDAR (model not reported) | 方法輸入 | Oxford Spires; NTU VIRAL | maximum range about 150 m stated for the indoor LiDAR setup | (Affan et al., 2026, Sec. 3; Sec. 5.2.2) |
| 地面雷射掃描儀(TLS) | terrestrial laser scanner (model not reported) | 參考或真值量測 | Oxford Spires | TLS ground truth of the Christ Church College scene | (Affan et al., 2026, Sec. 5.2) |
| 慣性量測單元(IMU) | IMU (model not reported) | 方法輸入 | Oxford Spires; NTU VIRAL | 原文未報告 | (Affan et al., 2026, Sec. 3, Sec. 3.2) |
| 相機 | RGB camera (model not reported) | 方法輸入 | Oxford Spires; NTU VIRAL | pinhole-plus-distortion projection model; calibrated LiDAR-to-camera extrinsics | (Affan et al., 2026, Sec. 3.2) |
論文圖片
只收錄原文以開放授權(open license)釋出的圖片,並依授權條件標示出處、圖號、授權與修改方式。

Fig. 2系統流程圖:RGB 全景分割、FAST-LIO2 里程計與標籤轉移、語意感知 TSDF 融合與網格化
出處:Affan et al., 2026,Fig. 2。授權:CC BY 4.0。原始圖檔。修改:轉存為 WebP 格式。

Fig. 3(c)本方法在 Oxford Spires(Christ Church College)重建的網格,可與 ImMesh、Voxblox 對照
出處:Affan et al., 2026,Fig. 3(c)。授權:CC BY 4.0。原始圖檔。修改:縮小至寬度不超過 1400 px,並轉存為 WebP 格式。

Fig. 4(c)本方法在 NTU VIRAL(NYA01)重建的網格,作者標註完整度較高的區域
出處:Affan et al., 2026,Fig. 4(c)。授權:CC BY 4.0。原始圖檔。修改:轉存為 WebP 格式。

Fig. 5(a)NTU VIRAL 重建的幾何不確定度分布,集中於深度不連續與高曲率處
出處:Affan et al., 2026,Fig. 5(a)。授權:CC BY 4.0。原始圖檔。修改:縮小至寬度不超過 1400 px,並轉存為 WebP 格式。
作者報告的優勢與限制
優勢
- On Oxford Spires Christ Church College (sequence 2) with TLS ground truth: accuracy RMSE 14.54 cm, completeness 98.55% and F1 98.58%, versus Voxblox 17.24 cm/97.17%/96.98% and ImMesh 17.40 cm/95.99%/96.31% (Sec. 5.2, Table 1).
限制
- ["Authors list odometry drift on longer trajectories, dependence on camera field-of-view coverage and calibration, and thin structures, specular materials and heavy clutter (Sec. 6).", "Meshes were aligned to the TLS ground truth by manual coarse registration plus ICP before scoring, so the metrics exclude global pose error (Sec. 5.2).", "Geometric and semantic uncertainty overlap only slightly, and the authors note that mismatched assumptions between LIO back-end, deskewing and meshing libraries can propagate into the mesh (Sec. 5.3, Table 2).", "Class multipliers, the sparsity factor and the weight cap were set empirically (Sec. 3.4).", "(observation) The statement that LIO drift stays within voxel-resolution tolerance on the evaluated trajectories (Sec. 3.1) is not backed by a reported trajectory error metric.", "(inference) Learned semantic priors can alter reconstructed geometry at boundaries
- independence of geometric evidence must be checked for engineering use."]
營建工程相關證據
定量評估僅用 Oxford Spires 的 Christ Church College(sequence 2,TLS 真值);作者在 Sec. 5.2 稱其為室內文化資產場景,但 Sec. 5.1.1 的定性描述同時涵蓋室外與室內空間(尖塔與拱形結構)。場景屬完工建築而非施工中工地;NTU VIRAL(NYA01)只作定性分析。引言提及 AEC 與數位分身用途,但論文未做營建場域驗證。
原文驗證環境:公開基準、已完工建築、獨立參考量測
報告的性能數據
以下是原文作者報告的性能數值(author-reported results),不是本研究重新量測的結果。每張圖只並列同一個比較組(comparison group,同一張表、同一組實驗設定)內的方法;不同比較組之間的數值不可直接比較,也不構成排名。
本方法共出現在 1 個比較組,合計 3 筆紀錄。
Affan et al., 2026 · Table 1 本方法 3 筆
資料集與序列Oxford Spires · Christ Church College (sequence 2)
表格設定(擷取紀錄原文):Meshes sampled to 500,000 points, manually coarse-registered then ICP-aligned to TLS ground truth; outlier radius filter; inlier threshold tau = 0.30 m; accuracy as inlier RMSE, completeness as recall percentage, F1 from precision and recall at tau (Affan et al., 2026, Table 1)
Acc. (cm), inlier RMSE reconstruction to ground truth,Oxford Spires · Christ Church College (sequence 2)
只並列這張表在相同設定下報告的方法;以「本方法:」開頭者為本頁方法。失敗、未執行與未報告以標記呈現,不是 0。
按 Tab 進入圖表後,用上下方向鍵逐一瀏覽各類別,Esc 關閉提示框;也可開啟表格檢視閱讀全部數值。
這些是 Affan et al., 2026 在此表設定下報告的數值(author-reported results),只能在同一個比較組內對照,不代表方法在其他資料或設定下的表現。
資料來源作者報告值(Affan et al., 2026, Table 1)
| 方法(原文寫法) | 報告值 | 出處 |
|---|---|---|
| ImMesh | 17.4 cm | (Affan et al., 2026, Table 1) |
| Voxblox | 17.24 cm | (Affan et al., 2026, Table 1) |
| Ours本方法原文提出 | 14.54 cm | (Affan et al., 2026, Table 1) |
來源
Affan et al., 2026
(2026)Incremental Semantics-Aided Meshing from LiDAR-Inertial Odometry and RGB Direct Label TransferThe International Archives of the Photogrammetry, Remote Sensing and Spatial Information Sciences, XLIX-B2-2026:335-342
DOI 10.5194/isprs-archives-xlix-b2-2026-335-2026arXiv 2604.09478
狀態未定已讀全文近十年查證後修正
相關版本
- 預印本:arXiv:2604.09478 https://arxiv.org/abs/2604.09478