Abstract:To address the challenge of recognizing pilot emotions in high-noise, low-resource cockpit environments, this paper proposes Wav-MHA, a model integrating multi-head attention mechanisms. First, data augmentation is performed using Gaussian white noise injection with signal-to-noise ratio adjustment and pitch shifting to expand sample size and improve generalization. Next, a pre-trained Wav2vec 2.0 model serves as the backbone, with a layer-wise freezing strategy applied to reduce parameters and mitigate overfitting. A task-oriented multi-head attention module is then introduced to focus on emotionally salient temporal segments, enhancing feature discriminability. Finally, on the self-constructed six-emotion cockpit speech dataset PSED, the proposed Wav-MHA model achieves a weighted accuracy of 74.89%, with particularly strong performance in recognizing neutral and fatigue states. Comparative and ablation experiments validate the model’s effectiveness and the necessity of each module. In cross-domain tests, the model attains accuracies of 95% and 87.83% on the RAVDESS and EMO-DB datasets, respectively, demonstrating superior cross-linguistic transferability and adaptability to complex environments.