Abstract:
To address the problems of unstable multi-scale feature fusion and erroneous attention activation exhibited by general facial expression recognition models in classroom scenarios, a student facial expression recognition model for classroom environments, termed MDSA-FER, is proposed, along with two attention mechanism modules. The first is a Difference-Aware Dynamic Multi-scale Fusion (DAMF) module. A “consensus–difference” decomposition paradigm is adopted, in which cross-scale global consensus features are first extracted and enhanced via coordinate attention. Subsequently, a spatially adaptive gating mechanism is employed to dynamically assess the importance of scale-wise differences relative to the consensus. Through residual reconstruction, complementary multi-scale information is selectively fused, thereby effectively suppressing redundant noise while preserving critical fine-grained expression cues. The second is a Symmetric Cross-Scale Attention (SCSA) module, in which a lightweight horizontal soft-alignment mechanism is introduced to correct pose deviations. Attention maps are generated by computing channel-aware mirror similarity between left and right facial features, and a multi-scale strategy is further incorporated to enhance the extraction of key facial regions. Comparative experiments and ablation analyses conducted on a public dataset and a classroom-scene dataset demonstrate that the proposed method achieves superior recognition accuracy and robustness compared with baseline models, providing an effective solution for student facial expression recognition in classroom scenarios.