This work investigates the pivotal design choices for multimodal misinformation detection. By running 3,375 experiments across three benchmark datasets with a variety of pretrained vision and language backbones, we systematically compare feature‑fusion methods, cross‑modal alignment strategies, and model scale, and conduct robustness analyses to identify which choices boost performance and which silently fail under certain perturbations. The study addresses four research questions: (1) which vision‑language backbone combinations are most effective? (2) how do the depth and type of fusion layers affect results? (3) how robust are models to image tampering or textual noise? (4) which component of the pipeline most strongly determines behavior? The findings provide practical guidelines for building stronger, more dependable multimodal misinformation detectors.
Review