• 中国期刊全文数据库
  • 中国学术期刊综合评价数据库
  • 中国科技论文与引文数据库
  • 中国核心期刊(遴选)数据库
欧阳宁, 张恩泽, 林乐平. 基于协同注意力机制的多模态情感分析模型J. 桂林电子科技大学学报, xxxx, x(x): 1-8. DOI: 10.16725/j.1673-808X.202587
引用本文: 欧阳宁, 张恩泽, 林乐平. 基于协同注意力机制的多模态情感分析模型J. 桂林电子科技大学学报, xxxx, x(x): 1-8. DOI: 10.16725/j.1673-808X.202587
OUYANG Ning, ZHANG Enze, LIN Leping. A multimodal sentiment analysis model based on collaborative attention mechanismsJ. Journal of Guilin University of Electronic Technology, xxxx, x(x): 1-8. DOI: 10.16725/j.1673-808X.202587
Citation: OUYANG Ning, ZHANG Enze, LIN Leping. A multimodal sentiment analysis model based on collaborative attention mechanismsJ. Journal of Guilin University of Electronic Technology, xxxx, x(x): 1-8. DOI: 10.16725/j.1673-808X.202587

基于协同注意力机制的多模态情感分析模型

A multimodal sentiment analysis model based on collaborative attention mechanisms

  • 摘要: 随着社交媒体的普及和数据来源的多样化,基于文本、语音和图像等多模态数据的情感分析逐渐兴起。相比单一模态,利用多模态信息可以更全面地感知和分析情感状态。然而,在多模态情感分析中,图像特征的空间分布蕴含着多层次的语义信息,如何有效利用这些多尺度特征是当前的一个关键挑战。针对图像信息的空间分布和语义细节常被忽视的问题,提出了一种基于协同注意力机制的多模态情感分析模型。该模型设计了一种多维特征协同注意力模块,结合多尺度空间感知和逐步信道优化策略,充分激发空间和通道注意力之间的协同效应,以实现对图像情感信息的精确捕捉。同时,通过引入多重投影变换模块,采用低投影、中层转换和高投影的三级结构,进一步增强了图像模态的表达能力。实验结果在公开数据集MOSI和MOSEI上的验证表明,所提模型在准确捕获多尺度图像特征方面表现优越。

     

    Abstract: With the popularity of social media and the diversification of data sources, sentiment analysis based on multimodal data such as text, speech, and images is gradually emerging. Compared with single modality, the emotional state can be perceived and analyzed more comprehensively using multimodal information. However, in multimodal sentiment analysis, the spatial distribution of image features contains multilevel semantic information, and how to effectively utilize these multiscale features is a key challenge at present. Aiming at the problem that the spatial distribution and semantic details of image information are often neglected, a multimodal sentiment analysis model based on the mechanism of collaborative attention is proposed. The model designs a multidimensional feature co-attention module, which combines multi-scale spatial perception and stepwise channel optimization strategies to fully stimulate the synergistic effect between spatial and channel attention in order to achieve accurate capture of image emotion information. Meanwhile, the expression ability of image modality is further enhanced by introducing a multiple projection transformation module with a three-level structure of low projection, mid-level transformation and high projection. The validation of the experimental results on the public datasets MOSI and MOSEI shows that the proposed model performs superiorly in accurately capturing multi-scale image features.

     

/

返回文章
返回