Abstract:Several challenges are encountered in driver emotion monitoring, including severe noise interference in multimodal physiological signals, ambiguous emotional boundaries, and limited adaptability to real-world driving scenarios. To address these challenges, MCG-Net, an emotion recognition method integrating multimodal physiological signals with adaptive temporal modeling, was proposed. Based on the Valence–Arousal dimensional model, a three-class emotion recognition framework was constructed by fusing electroencephalography (EEG), electromyography (EMG), and electrodermal activity (EDA) signals. Meanwhile, multi-scale feature enhancement, cross-layer gated temporal modeling, and a global attention mechanism were introduced to construct a multimodal recognition framework and enhance emotional feature representation. To validate the effectiveness and scene adaptability of the proposed model, this study conducts standardized emotion recognition evaluations on public datasets and further examines its application capability in complex driving scenarios using a self-constructed real-world driving dataset. The results show that MCG-Net achieves superior overall performance on public datasets in Valence, Arousal, and fused Valence–Arousal dimensional emotion recognition tasks. On the self-constructed real-world driving dataset, the trimodal fusion model achieves an accuracy of 82.35% and a Cohen’s Kappa coefficient of 0.7354. Further analyses based on the confusion matrix, predicted probability distributions, and feature-space visualizations demonstrate that the proposed method can effectively exploit the complementary information provided by EEG, EMG, and EDA signals, thereby enhancing the stability of emotion-boundary discrimination in complex driving scenarios. These findings provide a methodological reference for scene-adaptive multimodal modeling in driver emotion monitoring.