[{"data":1,"prerenderedAt":104},["ShallowReactive",2],{"method-nerfslam2023":3},{"method":4,"reference":61,"equipment":82,"figures":90,"results":87},{"id":5,"label":6,"shortName":7,"title":8,"year":9,"era":10,"cluster":11,"scope":12,"keyIdeaZh":13,"keyIdeaEn":14,"fulltextStatus":15,"publicationStatus":16,"recommendation":17,"constructionRelevance":18,"validationEnvironment":19,"strengths":22,"limitations":28,"sensors":33,"platform":35,"estimator":37,"association":38,"timeModel":39,"deskew":40,"loopClosure":41,"globalOptimization":42,"mapRepresentation":43,"prior":44,"outputGeometry":45,"compute":46,"codeUrl":47,"codeLicense":48,"relatedVersions":49},"nerfslam2023","Rosinol et al., 2023","NeRF-SLAM","NeRF-SLAM: Real-Time Dense Monocular SLAM with Neural Radiance Fields",2023,"recent","C09","odometry_with_local_mapping","NeRF-SLAM 把稠密單眼 SLAM 與即時雜湊式神經輻射場串接：追蹤端直接採用 DROID-SLAM 的學習式光流與稠密光束法平差，並依 σ-Fusion 的做法由 Hessian 結構計算每個深度與位姿的邊際共變異數。建圖端以 Instant-NGP 表示場景，損失同時包含顏色誤差與以共變異數加權的深度誤差，使雜訊大的單眼深度不致把幾何拉偏，並在同一執行緒中微調關鍵影格位姿。兩個執行緒平行運作，在單張 RTX 2080 Ti 上約每秒 10 影格，但系統沒有迴圈閉合。","Real-time monocular pipeline that feeds DROID-SLAM poses, dense depths and their marginal covariances into an Instant-NGP radiance field trained with a covariance-weighted depth loss, jointly refining poses and map; no loop closure.","full_text_reviewed","peer_reviewed_published","supplementary","論文未涉及營建場域；評估只使用 Replica 渲染序列與 Blender 合成的 Cube-Diorama，幾何品質以渲染深度的 L1 誤差代替，未報告軌跡誤差，也沒有網格或點雲對獨立參考量測的精度。其以單眼相機估計深度並依共變異數加權的做法，對低成本影像記錄工地有參考價值（推論），但約 11 GB GPU 記憶體需求與缺乏迴圈閉合限制了在大範圍工地的使用。",[20,21],"public_benchmark","simulation",[23,24,25,26,27],"Replica average Depth L1 9.29 cm and PSNR 42.03 dB from monocular input, versus 14.18 cm and 17.76 dB for NICE-SLAM without depth (IROS version Table I; arXiv v1 printed 4.49 cm and 41.40 dB)","Better geometric accuracy than TSDF-Fusion (21.88 cm) and sigma-Fusion (20.10 cm) built from the same poses and depths (Table I)","Up to 178% better PSNR (office-1) and 75% better Depth L1 (room-2) than NICE-SLAM (abstract, Sec. IV-C)","Uncertainty weighting avoids the slower and biased convergence seen when raw dense depths supervise the field (Sec. IV-D, Figs. 4-5)","Runs at about 10 FPS with tracking and mapping in parallel on one GPU (Sec. IV-E)",[29,30,31,32],"Needs about 11 GB of GPU memory because of dense correlation volumes and hierarchical grids, which can be prohibitive for low-compute robots such as drones (Sec. V)","NICE-SLAM supervised with ground-truth depth has a much lower average Depth L1 (4.08 cm) than NeRF-SLAM (9.29 cm), and in office-1 NeRF-SLAM (16.32 cm) is worse than NICE-SLAM without depth (10.24 cm) (Table I)","No trajectory accuracy (ATE) is reported and evaluation uses only synthetic Replica and Cube-Diorama data (Sec. IV) (inference)","No loop closure or global bundle adjustment; GO-SLAM later points out this limitation (Sec. III-C; GO-SLAM Sec. 2) (inference)",[34],"monocular camera",[36],"simulation (Replica rendered sequences and Blender Cube-Diorama)","DROID-SLAM dense bundle adjustment over a sliding window of at most 8 keyframes (Schur complement and Cholesky solve), with marginal covariances of dense depths and poses computed as in sigma-Fusion; the mapping thread minimizes photometric plus covariance-weighted depth loss jointly over poses and radiance-field parameters (Sec. III-A to III-C)","dense learned optical flow with per-measurement weights from a RAFT-style ConvGRU (DROID-SLAM) (Sec. III-A)","discrete poses (keyframes)","not_applicable","none described; tracking uses only a sliding window of keyframes (Sec. III-C)","none in tracking; the mapping thread optimizes all received keyframes (poses and map) through the rendering losses (Sec. III-B to III-C)","Instant-NGP hash-based hierarchical volumetric neural radiance field (density and colour), supervised by RGB and depth weighted by its marginal covariance (Sec. III-B, Sec. III-D)","pretrained DROID-SLAM weights for tracking (Sec. III-D)","keyframe poses, dense keyframe depth maps with uncertainty and a radiance field rendered to colour and depth; no mesh or point-cloud accuracy evaluation, geometry assessed by rendered Depth L1 (Sec. IV-C)","single NVIDIA RTX 2080 Ti (11 GB) shared by tracking and mapping; about 10 FPS overall at 640x480 (tracking 15 FPS, mapping 10 FPS; the section also states 12 FPS); needs about 11 GB GPU memory (Sec. III-D, Sec. IV-E, Sec. V)","https:\u002F\u002Fgithub.com\u002FToniRV\u002FNeRF-SLAM","BSD-2-Clause (GitHub license detection of LICENSE.BSD)",[50,54,58],{"relation":51,"title":52,"doi_or_url":53},"version_of_record","IROS 2023 proceedings (IEEE Xplore 10341922)","https:\u002F\u002Fdoi.org\u002F10.1109\u002FIROS55552.2023.10341922",{"relation":55,"title":56,"doi_or_url":57},"preprint","arXiv:2210.13641v1","https:\u002F\u002Farxiv.org\u002Fabs\u002F2210.13641",{"relation":59,"title":60,"doi_or_url":47},"code_release","ToniRV\u002FNeRF-SLAM",{"id":5,"kind":62,"shortName":7,"title":8,"authors":63,"year":9,"venue":67,"venueType":68,"publisher":69,"volumeIssuePages":70,"doi":71,"arxivId":72,"url":73,"firstPublicDate":74,"publicationStatus":16,"metadataStatus":75,"fulltextStatus":15,"era":10,"classicReason":40,"codeUrl":47,"cluster":11,"topics":76,"mdpi":77,"verification":78,"label":6,"fulltextRoute":79,"versionRead":80,"addedByCensus":81},"method",[64,65,66],"Antoni Rosinol","John J. Leonard","Luca Carlone","2023 IEEE\u002FRSJ International Conference on Intelligent Robots and Systems (IROS)","conference","IEEE","pp. 3437-3444","10.1109\u002Firos55552.2023.10341922","2210.13641","https:\u002F\u002Fapi.crossref.org\u002Fworks\u002F10.1109\u002FIROS55552.2023.10341922","2022-10-24","metadata_verified",[11],false,"corrected","NTU institutional (Chrome)","IROS 2023 version of record (IEEE Xplore HTML full text, Table I read from the table image) via NTU institutional access; arXiv v1 (2210.13641v1, CC BY 4.0) had been read in full and was compared. Values follow the version of record",true,[83],{"category":84,"model":85,"canonical":85,"role":86,"dataset":87,"specs":88,"locator":89},"compute","RTX 2080 Ti","compute for runtime",null,"GPU with 11 Gb memory, used for tracking and mapping","Sec. III-D",[91],{"refId":5,"refLabel":6,"fig":92,"whatZh":93,"license":94,"licenseUrl":95,"sourceUrl":96,"src":97,"width":98,"height":99,"thumb":100,"thumbWidth":101,"thumbHeight":102,"modified":103},"Fig. 2","系統架構：DROID-SLAM 追蹤迴路、邊際共變異數計算與以 Instant-NGP 擬合神經輻射場的資訊流","CC BY 4.0 (arXiv v1 version of the figure; the IROS version is copyright IEEE)","https:\u002F\u002Fcreativecommons.org\u002Flicenses\u002Fby\u002F4.0\u002F","https:\u002F\u002Farxiv.org\u002Fhtml\u002F2210.13641v1\u002Fimg\u002Farchitecture_7.png","\u002Ffigure-files\u002Fnerfslam2023\u002Ffig-2.webp",1400,1218,"\u002Ffigure-files\u002Fnerfslam2023\u002Ffig-2.thumb.webp",480,418,"resized to at most 1400 px wide and converted to WebP",1790510664209]