Adds a dense feature head and metric-aware regression to DUSt3R for 3D-grounded matching; the DTU MVS numbers use GT cameras for triangulation.

技術屬性

欄位內容為文獻擷取紀錄的原文用語(英文),以原文為據;「未查證」表示本研究尚未讀到該資訊,不代表該方法不具備此能力。

MASt3R 的技術屬性
感測輸入monocular camera (image pairs)
原文測試平台未記錄
狀態估計feed-forward pointmap regression (DUSt3R backbone) with an added dense local-feature head
資料關聯dense learned local features with fast reciprocal nearest-neighbour matching
時間表示不適用
去畸變不適用
迴圈閉合不適用
全域最佳化none
地圖表示pairwise pointmaps with confidence and dense descriptors
先驗資訊Learned prior from a mixture of 14 training datasets (10 with metric ground truth), initialized from the public DUSt3R checkpoint (ViT-Large encoder, ViT-Base decoder); regression normalization dropped when ground truth is metric (Sec. 3.1, 4.1)
可輸出幾何pointmaps (metric-scale when trained on metric data), dense matches; DTU point clouds by triangulating matches with GT cameras
計算需求No GPU or runtime hardware reported; matching time is reported only on a single CPU core (Fig. 2 right in the VoR); fast reciprocal matching with k = 3000 speeds matching about 64 times on Map-free; the network handles at most 512 px on the largest side, so high-resolution images need coarse-to-fine window matching (Sec. 3.3, 3.4, 4.2)

使用設備

原文使用的感測器、運算硬體與載具(equipment)。型號保留原文寫法,連結到設備頁中同一型號的歸併名稱;角色依原文用途分為方法輸入、資料集感測器、執行運算平台、參考或真值量測(reference or ground truth)與比較對象設備。

原文使用的設備
類別型號(原文寫法)角色資料集原文規格出處
相機iPhone 7資料集感測器InLoc329 InLoc query images(Leroy et al., 2024, Sec. 4.4)
相機mobile phones (models not named)資料集感測器Aachen Day-Night824 daytime and 98 nighttime Aachen query images(Leroy et al., 2024, Sec. 4.4)
相機hand-held cameras (models not named)資料集感測器Aachen Day-Night4,328 Aachen reference images(Leroy et al., 2024, Sec. 4.4)
運算硬體single CPU core (model not reported)執行運算平台Map-free relocalizationplatform for the matching-time versus accuracy plot(Leroy et al., 2024, Fig. 2 right (VoR); Fig. 3 right (arXiv v1))

作者報告的優勢與限制

優勢

限制

營建工程相關證據

論文未涉及營建場域。

原文驗證環境:公開基準

報告的性能數據

以下是原文作者報告的性能數值(author-reported results),不是本研究重新量測的結果。每張圖只並列同一個比較組(comparison group,同一張表、同一組實驗設定)內的方法;不同比較組之間的數值不可直接比較,也不構成排名。

本方法共出現在 9 個比較組,合計 59 筆紀錄。以下列出本方法紀錄最多的 4 組,其餘 5 組列在最後,並連到性能比較頁。

Leroy et al., 2024 · Table 2 本方法 21 筆

資料集與序列Map-free relocalization · test set (130 scenes)

表格設定(擷取紀錄原文):Map-free relocalization test set (VoR table, which adds FAR, RoMa and Mickey compared with arXiv v1). VCRE = virtual correspondence reprojection error, precision and AUC at VCRE < 90 px; pose precision and AUC at < 25 cm and 5 deg; median translation and rotation error; the depth column gives the metric-scale source (DPT fine-tuned on KITTI, KBR, or MASt3R's own depth, 'auto'). (Leroy et al., 2024, Table 2)

VCRE Reproj.,Map-free relocalization · test set (130 scenes)

只並列這張表在相同設定下報告的方法;以「本方法:」開頭者為本頁方法。失敗、未執行與未報告以標記呈現,不是 0。

按 Tab 進入圖表後,用上下方向鍵逐一瀏覽各類別,Esc 關閉提示框;也可開啟表格檢視閱讀全部數值。

這些是 Leroy et al., 2024 在此表設定下報告的數值(author-reported results),只能在同一個比較組內對照,不代表方法在其他資料或設定下的表現。

統計量:原文未報告;對齊方式:未對齊;單位:px;場景:Map-free relocalization test set (130 scenes, each with two video sequences); metric relative pose from a single reference image; qualitative pairs with viewpoint changes up to 180 deg

資料來源作者報告值(Leroy et al., 2024, Table 2)

數值與出處
方法(原文寫法)報告值出處
RPR [5] (DPT depth)147.1 px(Leroy et al., 2024, Table 2 (VoR))
SIFT [54] (DPT depth)222.8 px(Leroy et al., 2024, Table 2 (VoR))
SP+SG [78] (DPT depth)160.3 px(Leroy et al., 2024, Table 2 (VoR))
LoFTR [87] (KBR depth)165 px(Leroy et al., 2024, Table 2 (VoR))
FAR [75] (auto)137 px(Leroy et al., 2024, Table 2 (VoR))
RoMa [29] (DPT depth)128.8 px(Leroy et al., 2024, Table 2 (VoR))
Mickey [8] (auto)129.5 px(Leroy et al., 2024, Table 2 (VoR))
DUSt3R [106] (DPT depth)116 px(Leroy et al., 2024, Table 2 (VoR))
MASt3R (DPT depth)本方法原文提出104 px(Leroy et al., 2024, Table 2 (VoR))
MASt3R (auto, own metric depth)本方法原文提出48.7 px(Leroy et al., 2024, Table 2 (VoR))
MASt3R (direct reg., PnP on pointmap)本方法原文提出53.2 px(Leroy et al., 2024, Table 2 (VoR))

Liu et al., 2025 · Table 1 本方法 17 筆

表格設定(擷取紀錄原文):7 Scenes, one-twentieth of frames of each test sequence as input video; accuracy and completeness in cm against back-projected ground-truth depth; SLAM3R filters points with confidence threshold 3, SLAM3R-NoConf keeps all (Liu et al., 2025, Table 1)

Acc.,7-Scenes · Chess

只並列這張表在相同設定下報告的方法;以「本方法:」開頭者為本頁方法。失敗、未執行與未報告以標記呈現,不是 0。

按 Tab 進入圖表後,用上下方向鍵逐一瀏覽各類別,Esc 關閉提示框;也可開啟表格檢視閱讀全部數值。

這些是 Liu et al., 2025 在此表設定下報告的數值(author-reported results),只能在同一個比較組內對照,不代表方法在其他資料或設定下的表現。

統計量:平均值(mean);對齊方式:other: Umeyama similarity alignment followed by ICP to the ground-truth point cloud;單位:cm;場景:real indoor rooms

資料來源作者報告值(Liu et al., 2025, Table 1)

數值與出處
方法(原文寫法)報告值出處
DUSt3R [ 64 ]2.26 cm(Liu et al., 2025, Table 1)
MASt3R [ 28 ]本方法2.08 cm(Liu et al., 2025, Table 1)
Spann3R [ 61 ]2.23 cm(Liu et al., 2025, Table 1)
SLAM3R-NoConf (Ours)原文提出2.12 cm(Liu et al., 2025, Table 1)
SLAM3R (Ours)原文提出1.63 cm(Liu et al., 2025, Table 1)

Wang et al., 2025b · Table 3 本方法 4 筆

資料集與序列ETH3D · 10 random frames per scene

表格設定(擷取紀錄原文,這些數值分屬表中不同部分):(Wang et al., 2025b, Table 3)

  • Point map estimation on ETH3D, 10 random frames per scene, predicted cloud aligned to GT with the Umeyama algorithm (similarity or rigid not stated), invalid points filtered with official masks; global alignment; units not stated
  • Point map estimation on ETH3D, 10 random frames per scene, predicted cloud aligned to GT with the Umeyama algorithm (similarity or rigid not stated), invalid points filtered with official masks; point map head, feed-forward; units not stated
  • Point map estimation on ETH3D, 10 random frames per scene, predicted cloud aligned to GT with the Umeyama algorithm (similarity or rigid not stated), invalid points filtered with official masks; depth head unprojected with camera head, feed-forward; units not stated

Acc.,ETH3D · 10 random frames per scene

只並列這張表在相同設定下報告的方法;以「本方法:」開頭者為本頁方法。失敗、未執行與未報告以標記呈現,不是 0。

按 Tab 進入圖表後,用上下方向鍵逐一瀏覽各類別,Esc 關閉提示框;也可開啟表格檢視閱讀全部數值。

這些是 Wang et al., 2025b 在此表設定下報告的數值(author-reported results),只能在同一個比較組內對照,不代表方法在其他資料或設定下的表現。

統計量:原文未報告;對齊方式:原文未報告;單位:原文未報告;場景:not described in the paper

資料來源作者報告值(Wang et al., 2025b, Table 3)

數值與出處
方法(原文寫法)報告值出處
DUSt3R1.167(Wang et al., 2025b, Table 3)
MASt3R本方法0.968(Wang et al., 2025b, Table 3)
Ours (Point)原文提出0.901(Wang et al., 2025b, Table 3)
Ours (Depth + Cam)原文提出0.873(Wang et al., 2025b, Table 3)

Leroy et al., 2024 · Table 3 right 本方法 3 筆

資料集與序列DTU · evaluation set (average)

表格設定(擷取紀錄原文):DTU dense MVS (mm): accuracy, completeness and overall Chamfer (average of the two) with the benchmark's evaluation code; MASt3R and DUSt3R zero-shot; MASt3R matches are computed without camera knowledge but triangulated with ground-truth cameras in the ground-truth frame, then filtered by geometric consistency. Groups: (c) handcrafted, (d) learning-based trained on DTU, (e) zero-shot. VoR adds CasMVSNet and TransMVSNet compared with arXiv v1. (Leroy et al., 2024, Table 3 right)

Acc.,DTU · evaluation set (average)

只並列這張表在相同設定下報告的方法;以「本方法:」開頭者為本頁方法。失敗、未執行與未報告以標記呈現,不是 0。

按 Tab 進入圖表後,用上下方向鍵逐一瀏覽各類別,Esc 關閉提示框;也可開啟表格檢視閱讀全部數值。

這些是 Leroy et al., 2024 在此表設定下報告的數值(author-reported results),只能在同一個比較組內對照,不代表方法在其他資料或設定下的表現。

統計量:平均值(mean);對齊方式:未對齊;單位:mm;場景:DTU MVS evaluation set, zero-shot (no DTU training); input images 1200 x 1600

資料來源作者報告值(Leroy et al., 2024, Table 3 right)

數值與出處
方法(原文寫法)報告值出處
Camp [14] (c)0.835 mm(Leroy et al., 2024, Table 3 right (VoR))
Furu [32] (c)0.613 mm(Leroy et al., 2024, Table 3 right (VoR))
Tola [95] (c)0.342 mm(Leroy et al., 2024, Table 3 right (VoR))
Gipuma [33] (c)0.283 mm(Leroy et al., 2024, Table 3 right (VoR))
MVSNet [114] (d)0.396 mm(Leroy et al., 2024, Table 3 right (VoR))
CVP-MVSNet [113] (d)0.296 mm(Leroy et al., 2024, Table 3 right (VoR))
UCS-Net [18] (d)0.338 mm(Leroy et al., 2024, Table 3 right (VoR))
CER-MVS [57] (d)0.359 mm(Leroy et al., 2024, Table 3 right (VoR))
CIDER [111] (d)0.417 mm(Leroy et al., 2024, Table 3 right (VoR))
PatchmatchNet [103] (d)0.427 mm(Leroy et al., 2024, Table 3 right (VoR))
CasMVSNet [36] (d)0.325 mm(Leroy et al., 2024, Table 3 right (VoR))
TransMVSNet [22] (d)0.321 mm(Leroy et al., 2024, Table 3 right (VoR))
GeoMVSNet [122] (d)0.331 mm(Leroy et al., 2024, Table 3 right (VoR))
DUSt3R [106] (e)2.677 mm(Leroy et al., 2024, Table 3 right (VoR))
MASt3R (e)本方法原文提出0.403 mm(Leroy et al., 2024, Table 3 right (VoR))

其他比較組

列出其餘 5 個比較組

來源

  • Leroy et al., 2024

    Vincent Leroy, Yohann Cabon, Jerome Revaud(2024)Grounding Image Matching in 3D with MASt3RComputer Vision - ECCV 2024 (Lecture Notes in Computer Science), LNCS, pp. 71-91 (Crossref published-print 2025)

    同儕審查已出版已讀全文近十年查證後修正

回到方法圖鑑

選擇開啟Esc關閉