Multimodal emotion recognition is a fundamental problem in affective computing, with applications such as sentiment analysis, intelligent customer service, and human‑computer interaction. Existing approaches often rely on single‑modality features or employ naïve multimodal fusion, which fails to capture the synergy between global semantic context and fine‑grained inter‑modality interactions, limiting the depth of emotion understanding.
To address this limitation, we introduce a hybrid framework called Transformer‑GAT. The Transformer component captures global semantic information across utterances via self‑attention, modeling long‑range dependencies. Meanwhile, a Graph Attention Network (GAT) treats each modality (e.g., speech, text, video) as a node in a graph and learns attention‑weighted edges to model fine‑grained relationships between modalities, thereby enriching emotional feature representations. The outputs of both modules are combined in a weighted fusion layer, balancing global context with local modality details.
Extensive experiments on the IEMOCAP and MELD benchmarks show that our model achieves weighted F1 scores of 72.45% and 77.37%, respectively, surpassing state‑of‑the‑art baselines. These results demonstrate that Transformer‑GAT effectively integrates multimodal cues and provides deeper emotional insights, opening new directions for multimodal emotion computing.
Review