MASt3R
MASt3R 在 DUSt3R 上增加輸出稠密局部特徵的分支並以匹配損失訓練,同時提出快速互為最近鄰匹配以降低二次複雜度。與 DUSt3R 不同,當訓練真值為公制時不做尺度正規化,使模型可輸出公制尺度點圖。其 DTU 結果是以真值相機對匹配點三角化取得,並非完全無相機的重建。
本頁內容
Adds a dense feature head and metric-aware regression to DUSt3R for 3D-grounded matching; the DTU MVS numbers use GT cameras for triangulation.
技術屬性
欄位內容為文獻擷取紀錄的原文用語(英文),以原文為據;「未查證」表示本研究尚未讀到該資訊,不代表該方法不具備此能力。
| 感測輸入 | monocular camera (image pairs) |
|---|---|
| 原文測試平台 | 未記錄 |
| 狀態估計 | feed-forward pointmap regression (DUSt3R backbone) with an added dense local-feature head |
| 資料關聯 | dense learned local features with fast reciprocal nearest-neighbour matching |
| 時間表示 | 不適用 |
| 去畸變 | 不適用 |
| 迴圈閉合 | 不適用 |
| 全域最佳化 | none |
| 地圖表示 | pairwise pointmaps with confidence and dense descriptors |
| 先驗資訊 | Learned prior from a mixture of 14 training datasets (10 with metric ground truth), initialized from the public DUSt3R checkpoint (ViT-Large encoder, ViT-Base decoder); regression normalization dropped when ground truth is metric (Sec. 3.1, 4.1) |
| 可輸出幾何 | pointmaps (metric-scale when trained on metric data), dense matches; DTU point clouds by triangulating matches with GT cameras |
| 計算需求 | No GPU or runtime hardware reported; matching time is reported only on a single CPU core (Fig. 2 right in the VoR); fast reciprocal matching with k = 3000 speeds matching about 64 times on Map-free; the network handles at most 512 px on the largest side, so high-resolution images need coarse-to-fine window matching (Sec. 3.3, 3.4, 4.2) |
使用設備
原文使用的感測器、運算硬體與載具(equipment)。型號保留原文寫法,連結到設備頁中同一型號的歸併名稱;角色依原文用途分為方法輸入、資料集感測器、執行運算平台、參考或真值量測(reference or ground truth)與比較對象設備。
| 類別 | 型號(原文寫法) | 角色 | 資料集 | 原文規格 | 出處 |
|---|---|---|---|---|---|
| 相機 | iPhone 7 | 資料集感測器 | InLoc | 329 InLoc query images | (Leroy et al., 2024, Sec. 4.4) |
| 相機 | mobile phones (models not named) | 資料集感測器 | Aachen Day-Night | 824 daytime and 98 nighttime Aachen query images | (Leroy et al., 2024, Sec. 4.4) |
| 相機 | hand-held cameras (models not named) | 資料集感測器 | Aachen Day-Night | 4,328 Aachen reference images | (Leroy et al., 2024, Sec. 4.4) |
| 運算硬體 | single CPU core (model not reported) | 執行運算平台 | Map-free relocalization | platform for the matching-time versus accuracy plot | (Leroy et al., 2024, Fig. 2 right (VoR); Fig. 3 right (arXiv v1)) |
作者報告的優勢與限制
優勢
- Map-free test: VCRE AUC 0.933 and median translation error 0.36 m using MASt3R's own metric depth, vs 0.697 and 0.97 m for DUSt3R (Table 2, VoR)
- Zero-shot DTU MVS overall Chamfer 0.374 mm vs 1.741 mm for DUSt3R and within the range of DTU-trained methods (0.295 to 0.462 mm) (Table 3 right, VoR)
- InLoc: top-40 retrieval reaches 56.1/79.3/90.9% (DUC1) and 71.0/87.0/91.6% (DUC2), above the listed baselines (Table 4)
- Fast reciprocal matching speeds matching and improves pose accuracy through more uniform match coverage (Sec. 3.3; arXiv v1 App. B)
限制
- DTU reconstruction relies on GT cameras for triangulation (DTU paragraph)
- Follow-up (Murai et al., 2025) found scale often inconsistent across MASt3R predictions (MASt3R-SLAM Sec. 3.1)
- Network input limited to 512 px on the largest side; coarse-only matching nearly doubles DTU errors (Sec. 3.4; arXiv v1 App. C Table 5)
- Direct pointmap regression (PnP on the pointmap) gives poor localization in larger scenes such as Aachen and InLoc (Sec. 4.4, Table 4)
- Pairwise only in this paper; DUSt3R's multi-view global alignment is not used (Sec. 3.1)
- (reviewer inference) The VoR text still reports a 30-point VCRE AUC gain over LoFTR+KBR (0.634), but the VoR Table 2 also lists Mickey at 0.748, so the margin over the best listed baseline is about 19 points
營建工程相關證據
論文未涉及營建場域。
原文驗證環境:公開基準
報告的性能數據
以下是原文作者報告的性能數值(author-reported results),不是本研究重新量測的結果。每張圖只並列同一個比較組(comparison group,同一張表、同一組實驗設定)內的方法;不同比較組之間的數值不可直接比較,也不構成排名。
本方法共出現在 9 個比較組,合計 59 筆紀錄。以下列出本方法紀錄最多的 4 組,其餘 5 組列在最後,並連到性能比較頁。
Leroy et al., 2024 · Table 2 本方法 21 筆
資料集與序列Map-free relocalization · test set (130 scenes)
表格設定(擷取紀錄原文):Map-free relocalization test set (VoR table, which adds FAR, RoMa and Mickey compared with arXiv v1). VCRE = virtual correspondence reprojection error, precision and AUC at VCRE < 90 px; pose precision and AUC at < 25 cm and 5 deg; median translation and rotation error; the depth column gives the metric-scale source (DPT fine-tuned on KITTI, KBR, or MASt3R's own depth, 'auto'). (Leroy et al., 2024, Table 2)
VCRE Reproj.,Map-free relocalization · test set (130 scenes)
只並列這張表在相同設定下報告的方法;以「本方法:」開頭者為本頁方法。失敗、未執行與未報告以標記呈現,不是 0。
按 Tab 進入圖表後,用上下方向鍵逐一瀏覽各類別,Esc 關閉提示框;也可開啟表格檢視閱讀全部數值。
這些是 Leroy et al., 2024 在此表設定下報告的數值(author-reported results),只能在同一個比較組內對照,不代表方法在其他資料或設定下的表現。
資料來源作者報告值(Leroy et al., 2024, Table 2)
| 方法(原文寫法) | 報告值 | 出處 |
|---|---|---|
| RPR [5] (DPT depth) | 147.1 px | (Leroy et al., 2024, Table 2 (VoR)) |
| SIFT [54] (DPT depth) | 222.8 px | (Leroy et al., 2024, Table 2 (VoR)) |
| SP+SG [78] (DPT depth) | 160.3 px | (Leroy et al., 2024, Table 2 (VoR)) |
| LoFTR [87] (KBR depth) | 165 px | (Leroy et al., 2024, Table 2 (VoR)) |
| FAR [75] (auto) | 137 px | (Leroy et al., 2024, Table 2 (VoR)) |
| RoMa [29] (DPT depth) | 128.8 px | (Leroy et al., 2024, Table 2 (VoR)) |
| Mickey [8] (auto) | 129.5 px | (Leroy et al., 2024, Table 2 (VoR)) |
| DUSt3R [106] (DPT depth) | 116 px | (Leroy et al., 2024, Table 2 (VoR)) |
| MASt3R (DPT depth)本方法原文提出 | 104 px | (Leroy et al., 2024, Table 2 (VoR)) |
| MASt3R (auto, own metric depth)本方法原文提出 | 48.7 px | (Leroy et al., 2024, Table 2 (VoR)) |
| MASt3R (direct reg., PnP on pointmap)本方法原文提出 | 53.2 px | (Leroy et al., 2024, Table 2 (VoR)) |
Liu et al., 2025 · Table 1 本方法 17 筆
表格設定(擷取紀錄原文):7 Scenes, one-twentieth of frames of each test sequence as input video; accuracy and completeness in cm against back-projected ground-truth depth; SLAM3R filters points with confidence threshold 3, SLAM3R-NoConf keeps all (Liu et al., 2025, Table 1)
Acc.,7-Scenes · Chess
只並列這張表在相同設定下報告的方法;以「本方法:」開頭者為本頁方法。失敗、未執行與未報告以標記呈現,不是 0。
按 Tab 進入圖表後,用上下方向鍵逐一瀏覽各類別,Esc 關閉提示框;也可開啟表格檢視閱讀全部數值。
這些是 Liu et al., 2025 在此表設定下報告的數值(author-reported results),只能在同一個比較組內對照,不代表方法在其他資料或設定下的表現。
資料來源作者報告值(Liu et al., 2025, Table 1)
| 方法(原文寫法) | 報告值 | 出處 |
|---|---|---|
| DUSt3R [ 64 ] | 2.26 cm | (Liu et al., 2025, Table 1) |
| MASt3R [ 28 ]本方法 | 2.08 cm | (Liu et al., 2025, Table 1) |
| Spann3R [ 61 ] | 2.23 cm | (Liu et al., 2025, Table 1) |
| SLAM3R-NoConf (Ours)原文提出 | 2.12 cm | (Liu et al., 2025, Table 1) |
| SLAM3R (Ours)原文提出 | 1.63 cm | (Liu et al., 2025, Table 1) |
Wang et al., 2025b · Table 3 本方法 4 筆
資料集與序列ETH3D · 10 random frames per scene
表格設定(擷取紀錄原文,這些數值分屬表中不同部分):(Wang et al., 2025b, Table 3)
- Point map estimation on ETH3D, 10 random frames per scene, predicted cloud aligned to GT with the Umeyama algorithm (similarity or rigid not stated), invalid points filtered with official masks; global alignment; units not stated
- Point map estimation on ETH3D, 10 random frames per scene, predicted cloud aligned to GT with the Umeyama algorithm (similarity or rigid not stated), invalid points filtered with official masks; point map head, feed-forward; units not stated
- Point map estimation on ETH3D, 10 random frames per scene, predicted cloud aligned to GT with the Umeyama algorithm (similarity or rigid not stated), invalid points filtered with official masks; depth head unprojected with camera head, feed-forward; units not stated
Acc.,ETH3D · 10 random frames per scene
只並列這張表在相同設定下報告的方法;以「本方法:」開頭者為本頁方法。失敗、未執行與未報告以標記呈現,不是 0。
按 Tab 進入圖表後,用上下方向鍵逐一瀏覽各類別,Esc 關閉提示框;也可開啟表格檢視閱讀全部數值。
這些是 Wang et al., 2025b 在此表設定下報告的數值(author-reported results),只能在同一個比較組內對照,不代表方法在其他資料或設定下的表現。
資料來源作者報告值(Wang et al., 2025b, Table 3)
| 方法(原文寫法) | 報告值 | 出處 |
|---|---|---|
| DUSt3R | 1.167 | (Wang et al., 2025b, Table 3) |
| MASt3R本方法 | 0.968 | (Wang et al., 2025b, Table 3) |
| Ours (Point)原文提出 | 0.901 | (Wang et al., 2025b, Table 3) |
| Ours (Depth + Cam)原文提出 | 0.873 | (Wang et al., 2025b, Table 3) |
Leroy et al., 2024 · Table 3 right 本方法 3 筆
資料集與序列DTU · evaluation set (average)
表格設定(擷取紀錄原文):DTU dense MVS (mm): accuracy, completeness and overall Chamfer (average of the two) with the benchmark's evaluation code; MASt3R and DUSt3R zero-shot; MASt3R matches are computed without camera knowledge but triangulated with ground-truth cameras in the ground-truth frame, then filtered by geometric consistency. Groups: (c) handcrafted, (d) learning-based trained on DTU, (e) zero-shot. VoR adds CasMVSNet and TransMVSNet compared with arXiv v1. (Leroy et al., 2024, Table 3 right)
Acc.,DTU · evaluation set (average)
只並列這張表在相同設定下報告的方法;以「本方法:」開頭者為本頁方法。失敗、未執行與未報告以標記呈現,不是 0。
按 Tab 進入圖表後,用上下方向鍵逐一瀏覽各類別,Esc 關閉提示框;也可開啟表格檢視閱讀全部數值。
這些是 Leroy et al., 2024 在此表設定下報告的數值(author-reported results),只能在同一個比較組內對照,不代表方法在其他資料或設定下的表現。
資料來源作者報告值(Leroy et al., 2024, Table 3 right)
| 方法(原文寫法) | 報告值 | 出處 |
|---|---|---|
| Camp [14] (c) | 0.835 mm | (Leroy et al., 2024, Table 3 right (VoR)) |
| Furu [32] (c) | 0.613 mm | (Leroy et al., 2024, Table 3 right (VoR)) |
| Tola [95] (c) | 0.342 mm | (Leroy et al., 2024, Table 3 right (VoR)) |
| Gipuma [33] (c) | 0.283 mm | (Leroy et al., 2024, Table 3 right (VoR)) |
| MVSNet [114] (d) | 0.396 mm | (Leroy et al., 2024, Table 3 right (VoR)) |
| CVP-MVSNet [113] (d) | 0.296 mm | (Leroy et al., 2024, Table 3 right (VoR)) |
| UCS-Net [18] (d) | 0.338 mm | (Leroy et al., 2024, Table 3 right (VoR)) |
| CER-MVS [57] (d) | 0.359 mm | (Leroy et al., 2024, Table 3 right (VoR)) |
| CIDER [111] (d) | 0.417 mm | (Leroy et al., 2024, Table 3 right (VoR)) |
| PatchmatchNet [103] (d) | 0.427 mm | (Leroy et al., 2024, Table 3 right (VoR)) |
| CasMVSNet [36] (d) | 0.325 mm | (Leroy et al., 2024, Table 3 right (VoR)) |
| TransMVSNet [22] (d) | 0.321 mm | (Leroy et al., 2024, Table 3 right (VoR)) |
| GeoMVSNet [122] (d) | 0.331 mm | (Leroy et al., 2024, Table 3 right (VoR)) |
| DUSt3R [106] (e) | 2.677 mm | (Leroy et al., 2024, Table 3 right (VoR)) |
| MASt3R (e)本方法原文提出 | 0.403 mm | (Leroy et al., 2024, Table 3 right (VoR)) |
其他比較組
來源
Leroy et al., 2024
(2024)Grounding Image Matching in 3D with MASt3RComputer Vision - ECCV 2024 (Lecture Notes in Computer Science), LNCS, pp. 71-91 (Crossref published-print 2025)
DOI 10.1007/978-3-031-73220-1_5arXiv 2406.09756程式碼
同儕審查已出版已讀全文近十年查證後修正
相關版本
- 預印本:arXiv:2406.09756 https://arxiv.org/abs/2406.09756
程式碼:https://github.com/naver/mast3r(授權:CC BY-NC-SA 4.0)。有公開程式碼不等於已被重現,也不代表目前版本與論文版本相同。