Feed-forward monocular dense reconstruction that regresses keyframe pointmaps from sliding-window clips (I2P) and registers them into a global frame with retrieved scene frames (L2W), all initialized from DUSt3R and running at over 20 FPS without explicit camera parameters or bundle adjustment.

技術屬性

欄位內容為文獻擷取紀錄的原文用語(英文),以原文為據;「未查證」表示本研究尚未讀到該資訊,不代表該方法不具備此能力。

SLAM3R 的技術屬性
感測輸入monocular RGB camera (video)
原文測試平台原文未報告 (7-Scenes, ScanNet, ETH3D and in-the-wild videos; capture platforms not described)、simulation (Replica)
狀態估計feed-forward networks without explicit camera parameters: an Image-to-Points (I2P) ViT with multi-view cross-attention regresses keyframe pointmaps from sliding-window clips (L = 11), and a Local-to-World (L2W) network registers each keyframe's pointmap into the global frame using retrieved scene frames; no pose optimization or bundle adjustment (Sec. 3)
資料關聯implicit through cross-attention between keyframe and supporting or scene-frame tokens; a learned retrieval module selects the top-K scene frames from a reservoir of registered frames (Sec. 3.2, Supp. A)
時間表示不適用 (no explicit camera poses; poses can be derived afterwards by PnP-RANSAC)
去畸變不適用
迴圈閉合none as optimization; retrieval of long-term scene frames acts as implicit re-localization during registration (Sec. 4.2)
全域最佳化none; authors state that removing camera parameters prevents global bundle adjustment (Sec. 5)
地圖表示global dense point cloud built from registered pointmaps with per-pixel confidence (confidence threshold 3 in evaluation) (Sec. 3, Supp. B)
先驗資訊networks initialized from DUSt3R (224 x 224) weights and trained on about 850K clips from ScanNet++, Aria Synthetic Environments and CO3D-v2 (Sec. 4, Supp. A)
可輸出幾何dense 3D point cloud (colored by input frames); camera poses only as a derived by-product via PnP-RANSAC (Sec. 4.1)
計算需求about 24-25 FPS on a single NVIDIA 4090D at 224 x 224 input; training on 8 NVIDIA 4090D (24 GB) for about one day (Sec. 4, Tables 1-2)

使用設備

原文使用的感測器、運算硬體與載具(equipment)。型號保留原文寫法,連結到設備頁中同一型號的歸併名稱;角色依原文用途分為方法輸入、資料集感測器、執行運算平台、參考或真值量測(reference or ground truth)與比較對象設備。

原文使用的設備
類別型號(原文寫法)角色資料集原文規格出處
運算硬體NVIDIA 4090D執行運算平台未標示single GPU for FPS measurements; training on 8 GPUs with 24 GB each(Liu et al., 2025, Sec. 4, Sec. 4.1)

作者報告的優勢與限制

優勢

限制

營建工程相關證據

論文未涉及營建場域;定量評估限於 7-Scenes、Replica 與補充材料中 ScanNet、Tanks and Temples、ETH3D 的少數場景,且重建先以 Umeyama 與 ICP 對齊真值再計算公分級誤差,並未檢驗公制尺度或以獨立量測驗證。其不需相機參數即可即時由影片產生稠密點雲,對以一般手機或相機快速記錄工地狀況有吸引力,但輸入解析度低、無全域平差,量測用途仍需另行校核尺度與精度(推論)。

原文驗證環境:公開基準、模擬

報告的性能數據

性能數據仍在分批查證,目前尚未收錄此方法的報告值。

來源

  • Liu et al., 2025

    Yuzheng Liu, Siyan Dong, Shuzhe Wang, Yingda Yin, Yanchao Yang, Qingnan Fan, Baoquan Chen(2025)SLAM3R: Real-Time Dense Scene Reconstruction from Monocular RGB Videos2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 16651-16662

    同儕審查已出版已讀全文近十年查證後修正

回到方法圖鑑

選擇開啟Esc關閉