[{"data":1,"prerenderedAt":103},["ShallowReactive",2],{"method-nicerslam2024":3},{"method":4,"reference":62,"equipment":87,"figures":102,"results":92},{"id":5,"label":6,"shortName":7,"title":8,"year":9,"era":10,"cluster":11,"scope":12,"keyIdeaZh":13,"keyIdeaEn":14,"fulltextStatus":15,"publicationStatus":16,"recommendation":17,"constructionRelevance":18,"validationEnvironment":19,"strengths":22,"limitations":27,"sensors":33,"platform":35,"estimator":38,"association":39,"timeModel":40,"deskew":41,"loopClosure":42,"globalOptimization":43,"mapRepresentation":44,"prior":45,"outputGeometry":46,"compute":47,"codeUrl":48,"codeLicense":49,"relatedVersions":50},"nicerslam2024","Zhu et al., 2024","NICER-SLAM","NICER-SLAM: Neural Implicit Scene Encoding for RGB SLAM",2024,"recent","C09","odometry_with_local_mapping","NICER-SLAM 是只用單眼 RGB 影像的神經隱式 SLAM，追蹤與建圖共用同一個階層式 SDF 表示：粗層為 32 立方的稠密特徵格網，細層以多解析度網格學習殘差 SDF，另以多解析度網格表示顏色。因為沒有深度量測，建圖時額外加入 Omnidata 的單眼深度與法向量、GMFlow 光流、影像扭曲與 Eikonal 等損失來消除歧義，並依各體素取樣次數局部調整 SDF 轉密度的參數。系統在 Replica 上的幾何品質接近 RGB-D 方法，但追蹤不如 DROID-SLAM，未做迴圈閉合，也遠非即時。","RGB-only neural implicit SLAM that tracks and maps with one hierarchical SDF and colour grid, disambiguated by monocular depth and normal priors, optical flow and a warping loss, with a locally adaptive SDF-to-density transform; no loop closure and not real-time.","full_text_reviewed","peer_reviewed_published","supplementary","論文未涉及營建場域；定量評估只在合成 Replica 資料，3DV 版另以 Azure Kinect 自行拍攝 6 個戶外場景，但只做定性比較，且作者指出戶外深度量測不可靠。其以單眼深度與法向量先驗補足 RGB 幾何的做法，說明僅靠影像取得公分級網格仍依賴學習先驗，且未做迴圈閉合、速度遠非即時；對以一般相機記錄工地的情境可作為上限參考，但無法支持施工驗收等級的幾何主張（推論）。",[20,21],"public_benchmark","simulation",[23,24,25,26],"Best Replica geometry among RGB-only methods (accuracy 3.65 cm, completion 4.16 cm, completion ratio 79.37%, normal consistency 90.27%), ahead of TANDEM, NeRF-SLAM, DIM-SLAM* and DROID-SLAM and close to RGB-D NICE-SLAM (3.87, 3.87, 82.41, 89.93) (Table 1)","Replica tracking on par with RGB-D NICE-SLAM (1.88 versus 1.95 cm average ATE) without depth input (Table 3, Sec. 4.1)","Best average novel-view PSNR on Replica for both extrapolated (23.93 dB) and interpolated (25.41 dB) views, above RGB-D NICE-SLAM (23.26 and 24.42 dB) (Table 2)","On the self-captured outdoor Azure Kinect scenes it reconstructs textureless walls and small details such as a handrail where most monocular baselines fail (qualitative, Sec. 4.1, Fig. 6)",[28,29,30,31,32],"Not optimized for real-time operation (Sec. 5); about 496 ms per mapping iteration on an A100 (arXiv v1 Sec. 3.4)","No loop closure, so tracking could be improved (Sec. 5)","Tracking is clearly worse than systems built for tracking: Replica 1.88 cm versus 0.33 (DROID-SLAM), 0.55 (DSO), 0.46 (DIM-SLAM) and 1.15 cm (TANDEM) (Table 3); 7-Scenes 8.55 versus 5.66 cm for DROID-SLAM (arXiv v1 Table 4)","Removing the monocular depth or normal loss sharply degrades mapping and tracking, showing dependence on learned priors (arXiv v1 Table 5a; the 3DV version moves ablations to the supplementary)","The outdoor evaluation is qualitative only, and the Azure Kinect depth is unreliable outdoors (Sec. 4.1, Fig. 6)",[34],"monocular RGB camera",[36,37],"simulation (Replica)","not_reported (7-Scenes and the self-captured outdoor Azure Kinect dataset; carrier not described)","end-to-end optimization through differentiable volume rendering: tracking optimizes the current pose with an RGB rendering loss (100 iterations, 1024 pixels) with the map fixed; mapping runs a 3-stage optimization with RGB, warping, optical-flow, monocular depth, monocular normal and Eikonal losses, ending with local bundle adjustment over 16 selected frames of which half are frozen (Sec. 3.3; iteration counts from arXiv v1 Sec. 3.4)","direct photometric rendering loss plus dense correspondence cues: RGB warping between keyframes and optical flow from GMFlow (Sec. 3.3)","discrete poses","not_applicable","none (stated as a limitation, Sec. 5)","none; local BA over selected mapping frames only (Sec. 3.3-3.4)","hierarchical neural implicit SDF: coarse 32^3 dense feature grid plus 8-level fine residual grids (32-128) and a 16-level colour grid (16-2048) with small MLP decoders; VolSDF-style SDF-to-density with a locally adaptive beta from per-voxel sample counts (Sec. 3.1-3.2)","monocular depth and normal predictions from an off-the-shelf predictor (Omnidata in arXiv v1), optical flow from GMFlow; COLMAP used to obtain intrinsics for 7-Scenes and the self-captured outdoor dataset (Sec. 3.3, Sec. 4)","camera trajectory and a triangle mesh extracted by marching cubes at 512^3; rendered novel views (Sec. 3.4)","not reported in the 3DV main text; arXiv v1 Sec. 3.4 reports a single NVIDIA A100 with on average 496 ms per mapping iteration and 147 ms per tracking iteration (100 iterations each); not real-time (Sec. 5)","https:\u002F\u002Fgithub.com\u002Fcvg\u002Fnicer-slam","Apache-2.0 (LICENSE file checked)",[51,55,59],{"relation":52,"title":53,"doi_or_url":54},"version_of_record","3DV 2024 proceedings (IEEE Xplore 10550721)","https:\u002F\u002Fdoi.org\u002F10.1109\u002F3DV62453.2024.00096",{"relation":56,"title":57,"doi_or_url":58},"preprint","arXiv:2302.03594v1","https:\u002F\u002Farxiv.org\u002Fabs\u002F2302.03594",{"relation":60,"title":61,"doi_or_url":48},"code_release","cvg\u002Fnicer-slam",{"id":5,"kind":63,"shortName":7,"title":8,"authors":64,"year":9,"venue":72,"venueType":73,"publisher":74,"volumeIssuePages":75,"doi":76,"arxivId":77,"url":78,"firstPublicDate":79,"publicationStatus":16,"metadataStatus":80,"fulltextStatus":15,"era":10,"classicReason":41,"codeUrl":48,"cluster":11,"topics":81,"mdpi":82,"verification":83,"label":6,"fulltextRoute":84,"versionRead":85,"addedByCensus":86},"method",[65,66,67,68,69,70,71],"Zihan Zhu","Songyou Peng","Viktor Larsson","Zhaopeng Cui","Martin R. Oswald","Andreas Geiger","Marc Pollefeys","2024 International Conference on 3D Vision (3DV)","conference","IEEE","pp. 42-52","10.1109\u002F3dv62453.2024.00096","2302.03594","https:\u002F\u002Fapi.crossref.org\u002Fworks\u002F10.1109\u002F3DV62453.2024.00096","2023-02-07","metadata_verified",[11],false,"corrected","NTU institutional (Chrome)","3DV 2024 version of record (IEEE Xplore HTML full text; Tables 1-3 read from the table images) via NTU institutional access; arXiv v1 (2302.03594v1, 7 Feb 2023) had been read in full and was compared. Results follow the version of record except six 7-Scenes rows kept from arXiv v1",true,[88,95],{"category":89,"model":90,"canonical":90,"role":91,"dataset":92,"specs":93,"locator":94},"compute","A100","compute for runtime",null,"single GPU; 496 ms per mapping iteration, 147 ms per tracking iteration","arXiv v1 Sec. 3.4 (implementation details are not in the 3DV main text)",{"category":96,"model":97,"canonical":97,"role":98,"dataset":99,"specs":100,"locator":101},"rgbd","Azure Kinect","method input","SCO (self-captured outdoor)","used to capture the self-captured outdoor (SCO) dataset of 6 scenes with 800 to 2700 frames; only RGB images are input, the depth is shown for visualization and is unreliable outdoors","Sec. 4 Datasets; Sec. 4.1; Fig. 6",[],1790510664221]