Multimodal emotion recognition (MER) aims to infer human affective states by integrating complementary cues from modalities such as audio and text. In practice, affective cues are entangled with speaker style and lexical content, and cross‑modal disagreement further complicates evidence integration. Conventional discriminative fusion compresses multimodal evidence into a final prediction, which insufficiently preserves modality‑specific cues and conflict information. By contrast, large generative affective models embed reasoning within language decoding, leaving emotion evidence implicit and hard to verify in a structured space. To overcome these limitations, we propose BiCFlow-MER (Bidirectional Conditional Flow for Multimodal Emotion Recognition), formulating audio‑text MER as generative evidence transport within a structured emotion space. BiCFlow-MER first disentangles emotion‑oriented evidence from speaker‑style and lexical‑content factors, constructing a conflict‑aware affective condition. Guided by this condition, a bidirectional rectified flow transports each utterance to an explicit emotion‑space endpoint. Candidate emotions are jointly verified by (1) adaptive prototype‑cloud scoring of the transported endpoint and (2) backward class‑to‑condition consistency with the original multimodal condition, enabling conflict‑aware recognition. Experiments on IEMOCAP, MELD, and the zero‑shot CASE benchmark show that BiCFlow-MER outperforms all compared methods, demonstrating the synergy of discriminative recognition and generative evidence modeling via conditional transport and establishing a new MER paradigm.
Review