Title: Deep learning-based multimodal vocal emotion recognition
Authors: Lei Zhang
Addresses: Henan Academy of Drama, Henan University, Zhengzhou, 451464, China
Abstract: In the field of intelligent interaction and mental health screening, accurately identifying the emotions in singing is crucial. However, traditional methods rely solely on voice features, which are prone to misjudgement when recognising complex emotions (such as 'sarcastic' or 'mixed emotions'). To address this, this paper proposes a multimodal deep learning model that integrates singing and lyrics text, achieving deep collaboration of the two types of information through a cross-modal attention mechanism. Experiments show that the sentiment recognition accuracy of the model proposed in this paper reaches 81.3%, which is significantly higher than that of the model using only voice (accuracy 72.1%) and the early fusion method (accuracy 78.5%). Its comprehensive discrimination ability is also superior to the comparison baseline. This confirms that multimodal fusion can more comprehensively capture emotional cues, providing a reliable solution for achieving more refined human-computer interaction.
Keywords: vocal emotion recognition; multimodal learning; attention mechanism; deep learning.
DOI: 10.1504/IJICT.2026.154383
International Journal of Information and Communication Technology, 2026 Vol.27 No.69, pp.69 - 91
Received: 26 Feb 2026
Accepted: 02 Apr 2026
Published online: 25 Jun 2026 *


