Abstract:
With the popularity of social media and the diversification of data sources, sentiment analysis based on multimodal data such as text, speech, and images is gradually emerging. Compared with single modality, the emotional state can be perceived and analyzed more comprehensively using multimodal information. However, in multimodal sentiment analysis, the spatial distribution of image features contains multilevel semantic information, and how to effectively utilize these multiscale features is a key challenge at present. Aiming at the problem that the spatial distribution and semantic details of image information are often neglected, a multimodal sentiment analysis model based on the mechanism of collaborative attention is proposed. The model designs a multidimensional feature co-attention module, which combines multi-scale spatial perception and stepwise channel optimization strategies to fully stimulate the synergistic effect between spatial and channel attention in order to achieve accurate capture of image emotion information. Meanwhile, the expression ability of image modality is further enhanced by introducing a multiple projection transformation module with a three-level structure of low projection, mid-level transformation and high projection. The validation of the experimental results on the public datasets MOSI and MOSEI shows that the proposed model performs superiorly in accurately capturing multi-scale image features.