Novel view synthesis (NVS) models can render realistic images of the same scene from different viewpoints, yet the generated views are not always geometrically consistent. Multi-view (MV) consistency has been used to assess NVS quality, but its potential for multimedia forensics—specifically for localizing geometric inconsistencies across wide-baseline image pairs—remains largely unexplored.
To address this gap, we introduce DeformView, a wide-baseline MV dataset that provides pixel‑level annotations of geometric inconsistencies. The dataset serves as a unified benchmark for the localization task.
Using DeformView, we evaluate state‑of‑the‑art MV consistency‑scoring methods and find that they transfer poorly to the forensic localization scenario, exhibiting high false‑positive rates.
We therefore propose DEFECt3R, a lightweight learning‑based classifier that leverages cross‑view feature relationships to predict geometric inconsistencies at the pixel level. Training incorporates hard negatives—geometrically consistent but visually deformed views—providing explicit supervision. Experiments show that DEFECt3R markedly improves localization accuracy while substantially reducing false positives.
Ablation studies reveal that both the choice of feature representation and the quality of cross‑view correspondences critically affect performance.
In summary, MV geometric consistency is a promising yet underexploited signal for forensics. This work establishes the first benchmark and baseline for wide‑baseline geometric inconsistency localization, paving the way for future research.
Review: MV geometric consistency offers strong forensic cues, and DEFECt3R demonstrates an effective approach to harnessing them.