[{"data":1,"prerenderedAt":109},["ShallowReactive",2],{"method-rtgslam2024":3},{"method":4,"reference":60,"equipment":85,"figures":108,"results":97},{"id":5,"label":6,"shortName":7,"title":8,"year":9,"era":10,"cluster":11,"scope":12,"keyIdeaZh":13,"keyIdeaEn":14,"fulltextStatus":15,"publicationStatus":16,"recommendation":17,"constructionRelevance":18,"validationEnvironment":19,"strengths":23,"limitations":28,"sensors":34,"platform":36,"estimator":40,"association":41,"timeModel":42,"deskew":43,"loopClosure":44,"globalOptimization":45,"mapRepresentation":46,"prior":47,"outputGeometry":48,"compute":49,"codeUrl":50,"codeLicense":51,"relatedVersions":52},"rtgslam2024","Peng et al., 2024","RTG-SLAM","RTG-SLAM: Real-time 3D Reconstruction at Scale using Gaussian Splatting",2024,"recent","C09","full_slam_with_global_correction","RTG-SLAM 是以 RGB-D 相機即時重建大範圍室內場景的三維高斯 SLAM。每個高斯只能是不透明或近乎透明：不透明高斯被視為橢圓圓盤，深度以射線與圓盤交點計算，使單一高斯即可貼合一塊局部表面，透明高斯只補足殘餘顏色，因此所需高斯數量與記憶體大幅減少。系統只對新觀測、顏色誤差大或深度誤差大的像素新增高斯，並只最佳化尚未穩定的高斯與其覆蓋的像素，追蹤則採用傳統的影格對模型 ICP，並以沿用 ORB-SLAM2 的後端做特徵地標圖最佳化。","Real-time RGB-D Gaussian-splatting SLAM for large indoor scenes using compact opaque (disc-depth) and transparent (residual-colour) Gaussians, adding Gaussians only where new or erroneous and optimizing only unstable ones, with frame-to-model ICP tracking and an ORB-SLAM2-style landmark back end.","full_text_reviewed","peer_reviewed_published","main_body","論文以 Azure Kinect 手持掃描 43 至 100 m2 的走廊、倉庫、旅館房間、住家與辦公室，屬大範圍室內建物尺度，並以 ScanNet++ 雷射掃描模型評估幾何（使用真值位姿時精度約 0.95 cm、3 cm 內比例 96%）。語料中的地下工程研究 [yan2026_underground3dgsslam] 也以它為 3DGS 基準（Table III、Table VI）。但論文未在施工現場驗證，自掃資料沒有真值，且 ScanNet++ 的幾何評估排除了追蹤誤差，工地使用仍需獨立的精度檢核（推論）。",[20,21,22],"public_benchmark","simulation","completed_building",[24,25,26,27],"About twice the speed and half the memory of Co-SLAM on the Azure home scene (17.90 FPS and 8782 MB versus 8.65 FPS and 17342 MB), while SplaTAM and ESLAM run out of memory there (Table 1)","Lowest TUM RGB-D average tracking error among compared neural methods (1.06 cm), close to ORB-SLAM2 (1.00 cm) (Table 2)","ScanNet++ geometry accuracy 0.95 cm with 96.41% within 3 cm using ground-truth poses, second only to Point-SLAM which uses correct depth (Table 3)","Real-time reconstruction of self-scanned scenes of 43-100 m2 at around 16 FPS without post-processing (Sec. 1, Supp. Table 4)",[29,30,31,32,33],"Rendering quality is degraded compared with original Gaussians because only opaque and transparent Gaussians are used (Sec. 5)","Reflective or transparent materials make Gaussians switch states and optimize poorly (Sec. 5)","Outdoor scenes, dynamic objects, fast camera motion and changing lighting are not yet handled (Sec. 5)","Tracking fails on ScanNet++ because cameras are far apart, so only geometry with ground-truth poses is evaluated there (Supp. E.1)","Self-scanned Azure dataset has no ground truth and is used qualitatively except for time and memory (Supp. D)",[35],"RGB-D camera (Microsoft Azure Kinect for the self-scanned dataset)",[37,38,39],"handheld (Azure Kinect tethered to a laptop, frames streamed to a desktop)","simulation (Replica)","not_reported (TUM RGB-D and ScanNet++ capture platforms not described)","multi-level frame-to-model point-to-plane ICP against depth and normals rendered from the Gaussians (front end), plus an ORB-SLAM2-derived back-end graph optimization over 3D ORB landmarks in a separate C++ thread; mapping optimizes only unstable Gaussians with L1 colour and depth losses (Sec. 3.2, Supp. C)","projective point-to-plane ICP correspondences between the current depth frame and the rendered model; ORB feature landmarks in the back end (Sec. 3.2)","discrete poses","not_applicable","not_reported (the back end is inherited from ORB-SLAM2, but loop detection is not described in the paper)","back-end graph optimization over ORB landmarks (ORB-SLAM2 style); global Gaussian optimization on keyframes during scanning and over all keyframes at the end (Sec. 3.2, Supp. C)","compact 3D Gaussians forced to be opaque (alpha 0.99, fitting surface and dominant colour, depth rendered by ray intersection with the Gaussian's ellipsoid disc) or nearly transparent (alpha 0.1, residual colour); stable and unstable states with confidence counts; spherical harmonics colour (Sec. 3.1, 3.2)","none; Gaussians initialized from sensor depth, vertices and normals (Sec. 3.2, Supp. A)","Gaussian map rendered to colour, depth and normals; geometry evaluated from points sampled uniformly from the Gaussians (Sec. 4.2)","Intel i9 13900KF with NVIDIA RTX 4090; 17.24 FPS and 2751 MB on Replica office0, 17.90 FPS and 8782 MB on the Azure home scene, 21.74 FPS and 3563 MB on TUM; capture laptop Intel i7 10750-H with NVIDIA 2070 (Table 1, Supp. C, Supp. Table 7)","https:\u002F\u002Fgithub.com\u002FMisEty\u002FRTG-SLAM","GPL-3.0 (LICENSE file checked)",[53,57],{"relation":54,"title":55,"doi_or_url":56},"preprint","arXiv:2404.19706v3","https:\u002F\u002Farxiv.org\u002Fabs\u002F2404.19706",{"relation":58,"title":59,"doi_or_url":50},"code_release","MisEty\u002FRTG-SLAM",{"id":5,"kind":61,"shortName":7,"title":8,"authors":62,"year":9,"venue":70,"venueType":71,"publisher":72,"volumeIssuePages":73,"doi":74,"arxivId":75,"url":76,"firstPublicDate":77,"publicationStatus":16,"metadataStatus":78,"fulltextStatus":15,"era":10,"classicReason":43,"codeUrl":50,"cluster":11,"topics":79,"mdpi":80,"verification":81,"label":6,"fulltextRoute":82,"versionRead":83,"addedByCensus":84},"method",[63,64,65,66,67,68,69],"Zhexi Peng","Tianjia Shao","Yong Liu","Jingke Zhou","Yin Yang","Jingdong Wang","Kun Zhou","SIGGRAPH '24 Conference Papers (Special Interest Group on Computer Graphics and Interactive Techniques Conference Conference Papers)","conference","ACM","Article pp. 1-11","10.1145\u002F3641519.3657455","2404.19706","https:\u002F\u002Fapi.crossref.org\u002Fworks\u002F10.1145\u002F3641519.3657455","2024-04-30","metadata_verified",[11],false,"corrected","arXiv","arXiv v3 (2404.19706v3, 9 May 2024) with the ACM SIGGRAPH Conference Papers '24 header and appended supplementary materials; ACM version of record not compared",true,[86,93,100,104],{"category":87,"model":88,"canonical":88,"role":89,"dataset":90,"specs":91,"locator":92},"rgbd","Microsoft Azure Kinect","method input","Azure dataset (self-scanned)","RGB-D camera for real-time scanning of the self-collected Azure dataset","Sec. 1, Sec. 4.1, Supp. C",{"category":94,"model":95,"canonical":95,"role":96,"dataset":97,"specs":98,"locator":99},"compute","intel i9 13900KF","compute for runtime",null,"desktop CPU running SLAM","Sec. 4.1",{"category":94,"model":101,"canonical":102,"role":96,"dataset":97,"specs":103,"locator":99},"Nvidia RTX 4090","NVIDIA RTX 4090","desktop GPU running SLAM",{"category":94,"model":105,"canonical":105,"role":96,"dataset":90,"specs":106,"locator":107},"intel i7 10750-H with nvidia 2070","laptop for data acquisition and viewing; frames sent to the desktop over wireless network","Supp. C",[],1790510660363]