MM-FGDNet is a multimodal framework designed for large‑scale live‑streaming e‑commerce environments, modeling abnormal behavior from temporal‑evolution and group‑structure perspectives. The cross‑modal temporal alignment module projects video, text, audio, and user‑action streams into a unified temporal semantic space, enabling synchronized representations across modalities. The temporal fraud‑pattern module captures the progression from weak early signals to sudden outbreaks, typically using recurrent networks or attention mechanisms to model signal intensity over time. The cooperative manipulation module employs a graph neural network to encode coordinated interactions among organized user groups and automated accounts, revealing potential collusive behavior. Experiments on real‑world multi‑platform datasets show that MM-FGDNet achieves an AUC of 0.927, F1 of 0.847, precision of 0.861, recall of 0.834, and an Early Detection Score of 0.689, while substantially reducing false alarms. Ablation studies confirm the contribution of each component, and cross‑domain tests demonstrate stable generalization to new streamers, product categories, and platforms. Overall, the results indicate that MM-FGDNet offers an effective and scalable proactive detection solution for coordinated abnormal activities in live‑streaming systems.
Review