[{"data":1,"prerenderedAt":95},["ShallowReactive",2],{"method-gsslam2024":3},{"method":4,"reference":58,"equipment":83,"figures":94,"results":88},{"id":5,"label":6,"shortName":7,"title":8,"year":9,"era":10,"cluster":11,"scope":12,"keyIdeaZh":13,"keyIdeaEn":14,"fulltextStatus":15,"publicationStatus":16,"recommendation":17,"constructionRelevance":18,"validationEnvironment":19,"strengths":22,"limitations":27,"sensors":33,"platform":35,"estimator":38,"association":39,"timeModel":40,"deskew":41,"loopClosure":42,"globalOptimization":43,"mapRepresentation":44,"prior":45,"outputGeometry":46,"compute":47,"codeUrl":48,"codeLicense":49,"relatedVersions":50},"gsslam2024","Yan et al., 2024","GS-SLAM (Yan et al.)","GS-SLAM: Dense Visual SLAM with 3D Gaussian Splatting",2024,"recent","C09","odometry_with_local_mapping","GS-SLAM 將三維高斯潑濺（3D Gaussian Splatting）用於 RGB-D 稠密 SLAM：場景由帶不透明度與一階球諧係數的各向異性高斯表示，位姿則透過作者推導的潑濺解析梯度直接最佳化。建圖時依累積不透明度偏低或渲染深度與量測深度不符的像素新增高斯，並降低不在表面附近之漂浮高斯的不透明度；追蹤先以半解析度渲染取得粗略位姿，再只用與深度一致的可靠高斯做全解析度精修。系統沒有迴圈閉合，評估用的網格是由估計位姿與深度經 TSDF 融合而得。","RGB-D SLAM that represents the scene with 3D Gaussians, adds or suppresses Gaussians according to rendered opacity and depth error, and tracks the camera by coarse-to-fine optimization through analytical splatting gradients; no loop closure.","full_text_reviewed","peer_reviewed_published","main_body","論文未涉及營建場域；評估只用 Replica 合成場景與 TUM RGB-D 三段桌面及辦公室序列。作者自承方法依賴高品質深度且大場景記憶體需求高，在真實 TUM 資料上的 ATE 為公分級並落後多個基準；幾何評估使用 1 cm 門檻的精確率與召回率，網格則來自 TSDF 融合而非高斯本身。對營建點雲而言，可作為三維高斯 SLAM 早期設計的參照，但無法直接支持施工尺度的幾何精度主張（推論）。",[20,21],"public_benchmark","simulation",[23,24,25,26],"Replica average ATE RMSE 0.50 cm, slightly below Point-SLAM (0.54 cm), at 8.34 FPS versus 0.42 FPS (Tables 1 and 4)","Best average Depth L1 (1.16 cm) and precision (74.0%) among compared neural RGB-D methods on Replica (Table 3)","Rendering at about 387 FPS with the best PSNR and SSIM and the lowest LPIPS among compared methods on Replica (Table 6)","A light variant with zero-order spherical harmonics cuts TUM memory to 18.8-22.3 MB (Table 5)",[28,29,30,31,32],"Relies on high-quality depth data (Sec. 5)","High memory for large scenes: 198.04 MB scene representation on Replica room0, about 4 times NICE-SLAM (Sec. 4.4, Sec. 5)","On real TUM RGB-D data the average ATE (3.7 cm) is worse than ESLAM, Co-SLAM and Point-SLAM, and far from classical RGB-D SLAM (Table 2, Sec. 4.2)","Poses are still sometimes degraded by improper Gaussians during joint optimization (Supp. Sec. 5)","No loop closure or global optimization (Sec. 3.3) (inference from the method description)",[34],"RGB-D camera",[36,37],"simulation (Replica synthetic sequences)","not_reported (TUM RGB-D capture platform not described)","gradient-based pose optimization (Adam on quaternion and translation) through analytical derivatives of differentiable Gaussian splatting; constant-velocity initialization, coarse stage on half-resolution renders, fine stage on full-resolution renders from depth-consistent (reliable) Gaussians (Sec. 3.3, Supp. Sec. 1-2, 5)","direct: L1 photometric loss on rendered colour for tracking; mapping and BA use L1 depth and colour rendering losses (Eqs. 6, 10, 13)","discrete poses (keyframes)","not_applicable","none","none; joint map and pose adjustment over K = 10 keyframes drawn at random from the keyframe database, poses optimized only in the second half of the iterations (Sec. 3.3, Supp. Sec. 5)","anisotropic 3D Gaussians with opacity and first-degree spherical harmonics (12 coefficients); adaptive expansion adds Gaussians at pixels with low cumulative opacity or depth mismatch and suppresses floaters by opacity decay (Sec. 3.1-3.2)","none; Gaussians initialized from sensor depth (half of the first-frame pixels back-projected) (Sec. 3.2)","3D Gaussian map with rendered RGB and depth; meshes for evaluation are produced by TSDF fusion of estimated poses and depth, since meshing the Gaussians directly was unsatisfactory (Supp. Sec. 5, Fig. 10)","Intel Core i9-13900K 5.5 GHz and NVIDIA RTX 4090; 8.34 FPS system rate and 198.04 MB scene memory on Replica room0 (Table 4); about 387 FPS rendering (Table 6)","https:\u002F\u002Fgithub.com\u002Fyanchi-3dv\u002Fdiff-gaussian-rasterization-for-gsslam","only the modified rasterizer is public, under the Inria and MPII Gaussian-Splatting License (research and non-commercial use); full SLAM code not released",[51,55],{"relation":52,"title":53,"doi_or_url":54},"preprint","arXiv:2311.11700v4","https:\u002F\u002Farxiv.org\u002Fabs\u002F2311.11700",{"relation":56,"title":57,"doi_or_url":48},"code_release","Modified differential Gaussian rasterization only (yanchi-3dv\u002Fdiff-gaussian-rasterization-for-gsslam)",{"id":5,"kind":59,"shortName":7,"title":8,"authors":60,"year":9,"venue":68,"venueType":69,"publisher":70,"volumeIssuePages":71,"doi":72,"arxivId":73,"url":74,"firstPublicDate":75,"publicationStatus":16,"metadataStatus":76,"fulltextStatus":15,"era":10,"classicReason":41,"codeUrl":48,"cluster":11,"topics":77,"mdpi":78,"verification":79,"label":6,"fulltextRoute":80,"versionRead":81,"addedByCensus":82},"method",[61,62,63,64,65,66,67],"Chi Yan","Delin Qu","Dan Xu","Bin Zhao","Zhigang Wang","Dong Wang","Xuelong Li","2024 IEEE\u002FCVF Conference on Computer Vision and Pattern Recognition (CVPR)","conference","IEEE","pp. 19595-19604","10.1109\u002Fcvpr52733.2024.01853","2311.11700","https:\u002F\u002Fapi.crossref.org\u002Fworks\u002F10.1109\u002FCVPR52733.2024.01853","2023-11-20","metadata_verified",[11],false,"corrected","arXiv","arXiv v4 (2311.11700v4, 7 Apr 2024), marked 'Accepted to CVPR 2024 (highlight)', including the supplementary material; cross-checked against the CVPR 2024 CVF open-access paper and supplementary (all extracted values match)",true,[84,91],{"category":85,"model":86,"canonical":86,"role":87,"dataset":88,"specs":89,"locator":90},"compute","Intel Core i9-13900K","compute for runtime",null,"5.50 GHz CPU","Sec. 4.1",{"category":85,"model":92,"canonical":92,"role":87,"dataset":88,"specs":93,"locator":90},"NVIDIA RTX 4090","GPU",[],1790510663985]