NeFut Logo NeFut
Admin Login

[CS.AI] Affect-Prototype Guided Fusion for Open-Vocabulary Incomplete Multi-modal Emotion Recognition

Published at: 2026-09-16 22:00 Last updated: 2026-09-18 00:46
#AI #Machine Learning #LLM

Open‑vocabulary multimodal emotion recognition (OV‑MER) aims to generate natural‑language emotion labels from multimodal affective cues. In real‑world settings, however, acquiring complete and synchronized modal data is often infeasible due to device limitations and privacy concerns. Existing OV‑MER approaches assume full‑modal inputs and therefore struggle to fuse features when some modalities are missing. Moreover, fusion methods designed for incomplete modalities are typically built for fixed‑label tasks and cannot accommodate the need to fuse emotional cues guided by arbitrary emotion semantics in an OV‑MER context. To address these challenges, we propose an Affect‑Prototype‑Conditioned Fusion (APCF) framework for incomplete open‑vocabulary emotion recognition. APCF is a candidate‑free generative framework that extends modal contribution learning to scenarios guided by any emotional semantics. Specifically, we construct an affect‑prototype library that explicitly models multimodal contribution patterns for diverse emotions, providing dynamic constraints for fusion under different semantic perspectives. Conditional retrieval and feature aggregation are then performed based on the available modal features, yielding refined fused affective representations. These representations are fed into a large language model (LLM) decoder to produce open‑vocabulary emotion labels. Experiments on the OV‑MERD+ and MER‑FG datasets demonstrate that APCF substantially outperforms state‑of‑the‑art baselines across all metrics.

Review

Original Source: https://arxiv.org/abs/2609.16962

[h] Back to Home