Modern generative models such as GANs, diffusion networks and autoregressive systems can now synthesize facial images that are virtually indistinguishable from real photographs, making forged‑image detection increasingly challenging and raising concerns about identity theft, fraud and misinformation. This study concentrates on GAN‑generated synthetic faces and investigates detection methods that rely solely on image analysis. Existing detectors mainly use convolutional neural networks (CNNs) or global vision transformers (ViT). CNNs excel at extracting local texture cues but struggle with broader contextual reasoning, whereas ViTs capture long‑range structures at the cost of heavy computation. We therefore explore three Swin‑Transformer‑based designs: a compact Swin trained from scratch, ImageNet‑1K pre‑trained Swin‑Tiny and Swin‑Small fine‑tuned for binary classification, and a novel hybrid that feeds EfficientNet‑B0 convolutional features into a Swin backend.
All models are evaluated on a 140 K real‑and‑fake face dataset comprising StyleGAN‑generated fakes, Flickr photos and DFDC real images, with balanced training, validation and test splits. The EfficientNet‑B0+Swin hybrid reaches 99% accuracy and 99.44% recall on a 5 000‑image test set, outperforming both pure Swin variants and a prior CNN‑only baseline. Results indicate that merging hierarchical CNN features with shifted‑window self‑attention yields an efficient, lightweight solution for detecting GAN‑synthesized faces.
Review: The paper highlights the promise of cross‑modal feature fusion, offering a practical approach for deep‑fake detection in resource‑constrained settings.