• 中国期刊全文数据库
  • 中国学术期刊综合评价数据库
  • 中国科技论文与引文数据库
  • 中国核心期刊(遴选)数据库
赵胤铎, 翟仲毅, 赵岭忠. 基于时序增强多模态注意力网络的行为识别模型J. 桂林电子科技大学学报, 2026, 46(3): 278-283. DOI: 10.16725/j.1673-808X.202359
引用本文: 赵胤铎, 翟仲毅, 赵岭忠. 基于时序增强多模态注意力网络的行为识别模型J. 桂林电子科技大学学报, 2026, 46(3): 278-283. DOI: 10.16725/j.1673-808X.202359
Zhao Yinduo, Zhai Zhongyi, Zhao Lingzhong. Enhanced temporal multimodal attention network for action recognitionJ. Journal of Guilin University of Electronic Technology, 2026, 46(3): 278-283. DOI: 10.16725/j.1673-808X.202359
Citation: Zhao Yinduo, Zhai Zhongyi, Zhao Lingzhong. Enhanced temporal multimodal attention network for action recognitionJ. Journal of Guilin University of Electronic Technology, 2026, 46(3): 278-283. DOI: 10.16725/j.1673-808X.202359

基于时序增强多模态注意力网络的行为识别模型

Enhanced temporal multimodal attention network for action recognition

  • 摘要: 行为识别是视觉智能的主流应用场景之一,其比静态图像分类更关注时间特征。目前,许多基于2D卷积神经网络的轻量级行为识别模型被提出,然而,这些模型未能很好地捕捉到运动特征的时间关系。为解决此问题,提出了一种时序增强多模态注意力网络E-TMAN,通过两阶段结构充分构建时间特征。E-TMAN包括局部运动增强(LME)和差分时间交互(DTI)2个主要模块。LME负责利用时间注意力机制将时间信息融合到图像特征中,以增强短期运动特征的表达能力。DTI采用差分范式学习长期的时间特征,采用了2种注意力机制,即通道注意力和空间注意力。这2种注意力机制通过不同的形式对语义特征进行补充。最后,在3个公共数据集(SomethingV2、UCF101和HMDB51)上进行实验来评估E-TMAN。实验结果表明,该模型能够有效利用时间特征,与现有方法相比,能够以较低的开销获得较好的性能。

     

    Abstract: Action recognition is one of the major application areas of computer vision and focuses more on temporal dynamics than static image classification. Recently, many lightweight temporal modeling methods based on 2D convolutional neural networks (CNN) have been proposed for action recognition. However, these models often fail to effectively capture long-range temporal dependencies in motion sequences. To address this limitation, an enhanced temporal multimodal attention network (E-TMAN) is proposed to model temporal features through a two-stage architecture. E-TMAN consists of two main modules: local motion enhancement (LME) and difference time interaction (DTI). The LME module employs a temporal attention mechanism to fuse temporal information into spatial features, thereby enhancing the representation of short-term motion patterns. The DTI module adopts a differential learning strategy to capture long-term temporal dependencies through two complementary attention mechanisms, namely channel attention and spatial attention. By integrating temporal and contextual information in different forms, the proposed network achieves more effective temporal feature modeling. Extensive experiments are conducted on three public benchmarks: SomethingV2, UCF101 and HMDB51. The results show that E-TMAN effectively exploits temporal features and achieves competitive performance with lower computational cost compared with state-of-the-art methods.

     

/

返回文章
返回