基于动态语义路由与对称双向增强的鲁棒多模态情感分析
Robust Multimodal Sentiment Analysis Based on Dynamic Semantic Routing and Symmetric Bidirectional Enhancement
-
摘要: 多模态情感分析(Multimodal Sentiment Analysis, MSA)致力于整合文本、声学和视觉等多源信息以识别复杂的人类情感,但在处理现实世界中广泛存在的非对齐和高歧义性数据时,仍面临跨模态语义层级错位以及对文本模态的单向依赖偏差两大核心挑战。传统融合方法往往隐含假设不同模态在神经网络的相同物理深度上具备语义对应关系,忽略了文本深层意图与视听浅层线索之间的异质性关联;且现有模型普遍采用以文本为中心的单向交互范式,当文本本身存在缺陷或具有误导性时,缺乏利用非语言线索进行反向修正的能力。针对上述局限,本文提出了一种基于动态语义路由与对称双向增强的鲁棒多模态情感分析网络(Dynamic Semantic Routing and Symmetric-enhanced MCEN, DySym-MCEN)。该架构受人类认知中“预测与纠错”双重机制启发,首先,设计了动态语义路由(Dynamic Semantic Routing, DSR)机制,通过计算跨模态特征的全局亲和度矩阵,打破了物理层级固定的刚性假设,自适应地实现了异构模态在逻辑语义层面的弹性对齐;其次,提出了对称双向增强(Symmetric Bidirectional Enhancement, SBE)模块,构建了数学上完全对称的双流交互路径,并配合自适应互补门控与循环一致性损失约束,赋予模型在模态冲突场景下的双向纠错能力;最后,在MOSI、MOSEI和CH-SIMS三个主流基准数据集上的广泛实验结果表明,DySym-MCEN在各项指标上表现优秀,特别是在包含大量非对齐样本的复杂场景中展现出了卓越的鲁棒性与泛化性能。Abstract: Multimodal sentiment analysis aims to recognize complex human emotions by effectively synthesizing heterogeneous information sources, including textual, acoustic, and visual modalities. However, when processing real-world data characterized by ubiquitous non-alignment and high ambiguity, existing studies still grapple with two fundamental challenges: cross-modal semantic level misalignment and unimodal dependency bias toward the textual modality. Traditional fusion methodologies often implicitly assume that distinct modalities share semantic correspondence at the exact same physical depth within neural networks, thereby neglecting the inherent heterogeneity between deep textual intents and shallow audio-visual cues, such as micro-expressions. Furthermore, prevalent models predominantly adopt a text-centric unidirectional interaction paradigm. Consequently, when the textual modality itself is defective or deceptive—as seen in sarcasm—these models lack the capability to utilize non-verbal cues to inversely correct the misleading semantics. To address these limitations, this paper proposes a robust framework named dynamic semantic routing and symmetric-enhanced modality-consistent embedding network (DySym-MCEN). Drawing inspiration from the dual cognitive mechanism of “prediction and error-correction”, our model constructs a robust feature interaction manifold through two key innovations. First, we introduce a dynamic semantic routing (DSR) mechanism. By calculating a global cross-modal affinity matrix, DSR effectively breaks the rigid assumption of fixed physical layer alignment, adaptively achieving the elastic alignment of heterogeneous modalities at a logical semantic level. Second, we propose a symmetric bidirectional enhancement (SBE) module, which constructs a mathematically symmetric dual-stream interaction architecture. Coupled with an adaptive complementary gating unit and a cycle consistency loss constraint, this approach endows the model with bidirectional error-correction capabilities in scenarios involving modal conflict. Finally, extensive experiments conducted on three mainstream benchmark datasets—MOSI, MOSEI, and CH-SIMS—demonstrated that DySym-MCEN achieved excellent performance across all evaluation metrics. Notably, the model exhibited superior robustness and generalization performance, particularly in complex scenarios characterized by a large number of unaligned samples.
下载: