Feed-forward multi-view geometry Transformer predicting cameras, depth and point maps in normalized scale; evaluated after similarity alignment.

技術屬性

欄位內容為文獻擷取紀錄的原文用語(英文),以原文為據;「未查證」表示本研究尚未讀到該資訊,不代表該方法不具備此能力。

VGGT 的技術屬性
感測輸入monocular camera (one to hundreds of views)
原文測試平台未記錄
狀態估計feed-forward Transformer with alternating frame/global attention predicting cameras, depth, point maps and tracks
資料關聯implicit (learned); 3D point tracks
時間表示不適用
去畸變不適用
迴圈閉合不適用
全域最佳化none in the feed-forward model; the paper also reports an optional post-hoc bundle-adjustment variant ('Ours (with BA)', Table 1)
地圖表示per-view depth and point maps in the first-camera frame
先驗資訊learned prior trained on a large mixture of 3D datasets
可輸出幾何point clouds (point head or depth+camera unprojection), camera parameters, depth maps; normalized, non-metric scale
計算需求Feature backbone on one H100 (flash attention v3, 336x518 images): 0.04 s and 1.88 GB for 1 frame, 1.04 s and 11.41 GB for 50 frames, 8.75 s and 40.63 GB for 200 frames (arXiv v1 Table 9); about 0.2 s feed-forward and about 1.8 s with BA for 10 frames (Table 1); trained on 64 A100 GPUs for nine days (Sec. 3.4); (Maggio et al., 2025) reports about 60 images on a 24 GB RTX 4090

使用設備

原文使用的感測器、運算硬體與載具(equipment)。型號保留原文寫法,連結到設備頁中同一型號的歸併名稱;角色依原文用途分為方法輸入、資料集感測器、執行運算平台、參考或真值量測(reference or ground truth)與比較對象設備。

原文使用的設備
類別型號(原文寫法)角色資料集原文規格出處
運算硬體NVIDIA H100執行運算平台未標示Single GPU with flash attention v3 for runtime and memory measurements; also used for Table 1 timings(Wang et al., 2025b, Table 1 caption; Table 9)

作者報告的優勢與限制

優勢

限制

營建工程相關證據

論文未涉及營建場域。

原文驗證環境:公開基準

報告的性能數據

以下是原文作者報告的性能數值(author-reported results),不是本研究重新量測的結果。每張圖只並列同一個比較組(comparison group,同一張表、同一組實驗設定)內的方法;不同比較組之間的數值不可直接比較,也不構成排名。

本方法共出現在 4 個比較組,合計 35 筆紀錄。

Wang et al., 2025b · Table 9 本方法 18 筆

表格設定(擷取紀錄原文):Feature-backbone inference time and peak GPU memory versus number of input frames (arXiv v1 only; not in the CVF main paper) (Wang et al., 2025b, Table 9)

Backbone inference time for N frames,不適用 · 1 input frames

這張表在此指標與資料序列只列出本方法一筆,沒有可並列的其他方法,因此不畫圖,數值與出處見下表。這是 Wang et al., 2025b 在此表設定下報告的數值(author-reported results),不代表方法在其他資料或設定下的表現。

統計量:原文未報告;對齊方式:未對齊;單位:s

數值與出處
方法(原文寫法)報告值出處
VGGT本方法原文提出硬體:single NVIDIA H100 with flash attention v3; 336x518 images0.04 s(Wang et al., 2025b, Table 9 (arXiv v1, Sec. 5))

Wang et al., 2025b · Table 3 本方法 8 筆

資料集與序列ETH3D · 10 random frames per scene

表格設定(擷取紀錄原文,這些數值分屬表中不同部分):(Wang et al., 2025b, Table 3)

  • Point map estimation on ETH3D, 10 random frames per scene, predicted cloud aligned to GT with the Umeyama algorithm (similarity or rigid not stated), invalid points filtered with official masks; global alignment; units not stated
  • Point map estimation on ETH3D, 10 random frames per scene, predicted cloud aligned to GT with the Umeyama algorithm (similarity or rigid not stated), invalid points filtered with official masks; point map head, feed-forward; units not stated
  • Point map estimation on ETH3D, 10 random frames per scene, predicted cloud aligned to GT with the Umeyama algorithm (similarity or rigid not stated), invalid points filtered with official masks; depth head unprojected with camera head, feed-forward; units not stated

Acc.,ETH3D · 10 random frames per scene

只並列這張表在相同設定下報告的方法;以「本方法:」開頭者為本頁方法。失敗、未執行與未報告以標記呈現,不是 0。

按 Tab 進入圖表後,用上下方向鍵逐一瀏覽各類別,Esc 關閉提示框;也可開啟表格檢視閱讀全部數值。

這些是 Wang et al., 2025b 在此表設定下報告的數值(author-reported results),只能在同一個比較組內對照,不代表方法在其他資料或設定下的表現。

統計量:原文未報告;對齊方式:原文未報告;單位:原文未報告;場景:not described in the paper

資料來源作者報告值(Wang et al., 2025b, Table 3)

數值與出處
方法(原文寫法)報告值出處
DUSt3R1.167(Wang et al., 2025b, Table 3)
MASt3R0.968(Wang et al., 2025b, Table 3)
Ours (Point)本方法原文提出0.901(Wang et al., 2025b, Table 3)
Ours (Depth + Cam)本方法原文提出0.873(Wang et al., 2025b, Table 3)

Wang et al., 2025b · Table 1 本方法 6 筆

表格設定(擷取紀錄原文):Camera pose estimation with 10 random frames per scene, AUC@30 combining relative rotation and translation accuracy; no method trained on RealEstate10K; per-method time column not extracted; values from arXiv v1; the CVF version prints 83.4 for FLARE on CO3Dv2 and marks MV-DUSt3R's CO3Dv2 value as not trained on CO3D (Wang et al., 2025b, Table 1)

AUC@30 (RRA and RTA),RealEstate10K (unseen) · 10 random frames per scene

只並列這張表在相同設定下報告的方法;以「本方法:」開頭者為本頁方法。失敗、未執行與未報告以標記呈現,不是 0。

按 Tab 進入圖表後,用上下方向鍵逐一瀏覽各類別,Esc 關閉提示框;也可開啟表格檢視閱讀全部數值。

這些是 Wang et al., 2025b 在此表設定下報告的數值(author-reported results),只能在同一個比較組內對照,不代表方法在其他資料或設定下的表現。

統計量:原文未報告;對齊方式:未對齊;單位:AUC score (0 to 100);場景:not described in the paper

資料來源作者報告值(Wang et al., 2025b, Table 1)

數值與出處
方法(原文寫法)報告值出處
Colmap+SPSG45.2(Wang et al., 2025b, Table 1)
PixSfM49.4(Wang et al., 2025b, Table 1)
PoseDiff48(Wang et al., 2025b, Table 1)
DUSt3R67.7(Wang et al., 2025b, Table 1)
MASt3R76.4(Wang et al., 2025b, Table 1)
VGGSfM v278.9(Wang et al., 2025b, Table 1)
MV-DUSt3R71.3(Wang et al., 2025b, Table 1)
CUT3R75.3(Wang et al., 2025b, Table 1)
FLARE78.8(Wang et al., 2025b, Table 1)
Fast3R72.7(Wang et al., 2025b, Table 1)
Ours (Feed-Forward)本方法原文提出85.3(Wang et al., 2025b, Table 1)
Ours (with BA)本方法原文提出93.5(Wang et al., 2025b, Table 1)

Wang et al., 2025b · Table 2 本方法 3 筆

資料集與序列DTU · evaluation scenes

表格設定(擷取紀錄原文,這些數值分屬表中不同部分):(Wang et al., 2025b, Table 2)

  • Dense MVS estimation on DTU; method uses known GT cameras; units not stated in the paper
  • Dense MVS estimation on DTU; method does not know GT cameras; units not stated in the paper

Acc.,DTU · evaluation scenes

只並列這張表在相同設定下報告的方法;以「本方法:」開頭者為本頁方法。失敗、未執行與未報告以標記呈現,不是 0。

按 Tab 進入圖表後,用上下方向鍵逐一瀏覽各類別,Esc 關閉提示框;也可開啟表格檢視閱讀全部數值。

這些是 Wang et al., 2025b 在此表設定下報告的數值(author-reported results),只能在同一個比較組內對照,不代表方法在其他資料或設定下的表現。

統計量:原文未報告;對齊方式:原文未報告;單位:原文未報告;場景:not described in the paper

資料來源作者報告值(Wang et al., 2025b, Table 2)

數值與出處
方法(原文寫法)報告值出處
Gipuma0.283(Wang et al., 2025b, Table 2)
MVSNet0.396(Wang et al., 2025b, Table 2)
CIDER0.417(Wang et al., 2025b, Table 2)
PatchmatchNet0.427(Wang et al., 2025b, Table 2)
MASt3R0.403(Wang et al., 2025b, Table 2)
GeoMVSNet0.331(Wang et al., 2025b, Table 2)
DUSt3R2.677(Wang et al., 2025b, Table 2)
Ours本方法原文提出0.389(Wang et al., 2025b, Table 2)

來源

  • Wang et al., 2025b

    Jianyuan Wang, Minghao Chen, Nikita Karaev, Andrea Vedaldi, Christian Rupprecht, David Novotny(2025)VGGT: Visual Geometry Grounded Transformer2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 5294-5306

    同儕審查已出版已讀全文近十年查證後修正

回到方法圖鑑

選擇開啟Esc關閉