[{"data":1,"prerenderedAt":91},["ShallowReactive",2],{"method-slam3r2025":3},{"method":4,"reference":57,"equipment":82,"figures":90,"results":87},{"id":5,"label":6,"shortName":7,"title":8,"year":9,"era":10,"cluster":11,"scope":12,"keyIdeaZh":13,"keyIdeaEn":14,"fulltextStatus":15,"publicationStatus":16,"recommendation":17,"constructionRelevance":18,"validationEnvironment":19,"strengths":22,"limitations":27,"sensors":32,"platform":34,"estimator":37,"association":38,"timeModel":39,"deskew":40,"loopClosure":41,"globalOptimization":42,"mapRepresentation":43,"prior":44,"outputGeometry":45,"compute":46,"codeUrl":47,"codeLicense":48,"relatedVersions":49},"slam3r2025","Liu et al., 2025","SLAM3R","SLAM3R: Real-Time Dense Scene Reconstruction from Monocular RGB Videos",2025,"recent","C09","map_representation_or_reconstruction","SLAM3R 以前饋式神經網路直接從單眼 RGB 影片產生稠密點雲，而不求解任何相機參數。影片先以滑動視窗切成重疊片段，影像對點雲（I2P）網路以多視角交叉注意力，從 11 張影像回歸視窗中間關鍵影格的點雲；局部對世界（L2W）網路再參考以檢索模組從儲存池挑出的歷史影格，把局部點雲逐步配準到全域座標。兩個網路都以 DUSt3R 權重初始化，在單張 4090D 上約每秒 24 至 25 影格，但因沒有相機參數，無法做全域光束法平差，推導出的位姿也不如專門的 SLAM。","Feed-forward monocular dense reconstruction that regresses keyframe pointmaps from sliding-window clips (I2P) and registers them into a global frame with retrieved scene frames (L2W), all initialized from DUSt3R and running at over 20 FPS without explicit camera parameters or bundle adjustment.","full_text_reviewed","peer_reviewed_published","supplementary","論文未涉及營建場域；定量評估限於 7-Scenes、Replica 與補充材料中 ScanNet、Tanks and Temples、ETH3D 的少數場景，且重建先以 Umeyama 與 ICP 對齊真值再計算公分級誤差，並未檢驗公制尺度或以獨立量測驗證。其不需相機參數即可即時由影片產生稠密點雲，對以一般手機或相機快速記錄工地狀況有吸引力，但輸入解析度低、無全域平差，量測用途仍需另行校核尺度與精度（推論）。",[20,21],"public_benchmark","simulation",[23,24,25,26],"Best accuracy and completeness among real-time methods on 7 Scenes (average 2.13 \u002F 2.34 cm at about 25 FPS) (Table 1)","On Replica, accuracy and completeness (3.57 \u002F 2.62 cm) comparable to optimization-based NICER-SLAM and DUSt3R while running at about 24 FPS (Table 2)","Learned L2W registration with retrieval is more accurate and faster than DUSt3R global alignment or Umeyama plus ICP (Table 5)","Much less drift than the concurrent Spann3R (Replica ATE 6.61 versus 32.79 cm) (Table 3, Supp. Table 7)",[28,29,30,31],"No camera parameters, so global bundle adjustment cannot be performed (Sec. 5)","Derived poses fall short of dedicated SLAM systems (Replica ATE 6.61 cm versus 0.33 cm for DROID-SLAM) (Sec. 5, Table 3)","Input images are center-cropped to 224 x 224 (Sec. 4)","Evaluation aligns reconstructions to ground truth with Umeyama and ICP, so metric scale and absolute drift are not assessed directly (Sec. 4.1) (inference)",[33],"monocular RGB camera (video)",[35,36],"not_reported (7-Scenes, ScanNet, ETH3D and in-the-wild videos; capture platforms not described)","simulation (Replica)","feed-forward networks without explicit camera parameters: an Image-to-Points (I2P) ViT with multi-view cross-attention regresses keyframe pointmaps from sliding-window clips (L = 11), and a Local-to-World (L2W) network registers each keyframe's pointmap into the global frame using retrieved scene frames; no pose optimization or bundle adjustment (Sec. 3)","implicit through cross-attention between keyframe and supporting or scene-frame tokens; a learned retrieval module selects the top-K scene frames from a reservoir of registered frames (Sec. 3.2, Supp. A)","not_applicable (no explicit camera poses; poses can be derived afterwards by PnP-RANSAC)","not_applicable","none as optimization; retrieval of long-term scene frames acts as implicit re-localization during registration (Sec. 4.2)","none; authors state that removing camera parameters prevents global bundle adjustment (Sec. 5)","global dense point cloud built from registered pointmaps with per-pixel confidence (confidence threshold 3 in evaluation) (Sec. 3, Supp. B)","networks initialized from DUSt3R (224 x 224) weights and trained on about 850K clips from ScanNet++, Aria Synthetic Environments and CO3D-v2 (Sec. 4, Supp. A)","dense 3D point cloud (colored by input frames); camera poses only as a derived by-product via PnP-RANSAC (Sec. 4.1)","about 24-25 FPS on a single NVIDIA 4090D at 224 x 224 input; training on 8 NVIDIA 4090D (24 GB) for about one day (Sec. 4, Tables 1-2)","https:\u002F\u002Fgithub.com\u002FPKU-VCL-3DV\u002FSLAM3R","CC BY-NC-SA 4.0 per the repository LICENSE text (GitHub API reports NOASSERTION)",[50,54],{"relation":51,"title":52,"doi_or_url":53},"preprint","arXiv:2412.09401v3","https:\u002F\u002Farxiv.org\u002Fabs\u002F2412.09401",{"relation":55,"title":56,"doi_or_url":47},"code_release","PKU-VCL-3DV\u002FSLAM3R",{"id":5,"kind":58,"shortName":7,"title":8,"authors":59,"year":9,"venue":67,"venueType":68,"publisher":69,"volumeIssuePages":70,"doi":71,"arxivId":72,"url":73,"firstPublicDate":74,"publicationStatus":16,"metadataStatus":75,"fulltextStatus":15,"era":10,"classicReason":40,"codeUrl":47,"cluster":11,"topics":76,"mdpi":77,"verification":78,"label":6,"fulltextRoute":79,"versionRead":80,"addedByCensus":81},"method",[60,61,62,63,64,65,66],"Yuzheng Liu","Siyan Dong","Shuzhe Wang","Yingda Yin","Yanchao Yang","Qingnan Fan","Baoquan Chen","2025 IEEE\u002FCVF Conference on Computer Vision and Pattern Recognition (CVPR)","conference","IEEE","pp. 16651-16662","10.1109\u002Fcvpr52734.2025.01552","2412.09401","https:\u002F\u002Fapi.crossref.org\u002Fworks\u002F10.1109\u002FCVPR52734.2025.01552","2024-12-12","metadata_verified",[11],false,"corrected","arXiv","arXiv v3 (2412.09401v3, 23 Mar 2025, labelled CVPR 2025) including the supplementary material; cross-checked against the CVPR 2025 CVF open-access paper and supplementary (all extracted values match)",true,[83],{"category":84,"model":85,"canonical":85,"role":86,"dataset":87,"specs":88,"locator":89},"compute","NVIDIA 4090D","compute for runtime",null,"single GPU for FPS measurements; training on 8 GPUs with 24 GB each","Sec. 4, Sec. 4.1",[],1790510664349]