• 中国期刊全文数据库
  • 中国学术期刊综合评价数据库
  • 中国科技论文与引文数据库
  • 中国核心期刊(遴选)数据库
冯程, 首照宇, 袁小虎, 等. 基于多尺度差异与对称注意力的学生表情识别J. 桂林电子科技大学学报, xxxx, x(x): 1-12. DOI: 10.16725/j.1673-808X.202620
引用本文: 冯程, 首照宇, 袁小虎, 等. 基于多尺度差异与对称注意力的学生表情识别J. 桂林电子科技大学学报, xxxx, x(x): 1-12. DOI: 10.16725/j.1673-808X.202620
FENG Cheng, SHOU Zhaoyu, YUAN Xiaohu, et al. Student Expression Recognition Based on Multi-Scale Difference and Symmetric AttentionJ. Journal of Guilin University of Electronic Technology, xxxx, x(x): 1-12. DOI: 10.16725/j.1673-808X.202620
Citation: FENG Cheng, SHOU Zhaoyu, YUAN Xiaohu, et al. Student Expression Recognition Based on Multi-Scale Difference and Symmetric AttentionJ. Journal of Guilin University of Electronic Technology, xxxx, x(x): 1-12. DOI: 10.16725/j.1673-808X.202620

基于多尺度差异与对称注意力的学生表情识别

Student Expression Recognition Based on Multi-Scale Difference and Symmetric Attention

  • 摘要: 针对通用表情识别模型在教室场景下易出现多尺度特征融合不稳定与注意力误激活的问题,提出了一种面向教室场景的学生表情识别模型MDSA-FER,并设计两种注意力机制模块:一是基于差异感知的动态多尺度融合模块(DAMF),该模块创新性地采用“共识-差异”分解范式,首先提取跨尺度的全局共识特征并进行坐标注意力增强,进而利用空间自适应门控机制动态评估各尺度相对于共识的差异重要性,通过残差重构实现多尺度互补信息的选择性融合,有效抑制了冗余噪声并保留了关键的细粒度表情线索;二是对称跨尺度注意力模块(SCSA),引入轻量级水平软对齐机制以校正姿态偏移,并通过计算左右特征的通道感知镜像相似度生成注意力图,结合多尺度策略进一步增强了模型对面部关键区域的提取能力。基于公开数据集与课堂场景数据集进行对比实验与消融分析,结果表明所提方法在识别准确性与鲁棒性方面优于基线模型,为教室场景下的学生表情识别提供了一种有效方案。

     

    Abstract: To address the problems of unstable multi-scale feature fusion and erroneous attention activation exhibited by general facial expression recognition models in classroom scenarios, a student facial expression recognition model for classroom environments, termed MDSA-FER, is proposed, along with two attention mechanism modules. The first is a Difference-Aware Dynamic Multi-scale Fusion (DAMF) module. A “consensus–difference” decomposition paradigm is adopted, in which cross-scale global consensus features are first extracted and enhanced via coordinate attention. Subsequently, a spatially adaptive gating mechanism is employed to dynamically assess the importance of scale-wise differences relative to the consensus. Through residual reconstruction, complementary multi-scale information is selectively fused, thereby effectively suppressing redundant noise while preserving critical fine-grained expression cues. The second is a Symmetric Cross-Scale Attention (SCSA) module, in which a lightweight horizontal soft-alignment mechanism is introduced to correct pose deviations. Attention maps are generated by computing channel-aware mirror similarity between left and right facial features, and a multi-scale strategy is further incorporated to enhance the extraction of key facial regions. Comparative experiments and ablation analyses conducted on a public dataset and a classroom-scene dataset demonstrate that the proposed method achieves superior recognition accuracy and robustness compared with baseline models, providing an effective solution for student facial expression recognition in classroom scenarios.

     

/

返回文章
返回