Open Access Article

Title: Deep learning-based multimodal vocal emotion recognition

Authors: Lei Zhang

Addresses: Henan Academy of Drama, Henan University, Zhengzhou, 451464, China

Abstract: In the field of intelligent interaction and mental health screening, accurately identifying the emotions in singing is crucial. However, traditional methods rely solely on voice features, which are prone to misjudgement when recognising complex emotions (such as 'sarcastic' or 'mixed emotions'). To address this, this paper proposes a multimodal deep learning model that integrates singing and lyrics text, achieving deep collaboration of the two types of information through a cross-modal attention mechanism. Experiments show that the sentiment recognition accuracy of the model proposed in this paper reaches 81.3%, which is significantly higher than that of the model using only voice (accuracy 72.1%) and the early fusion method (accuracy 78.5%). Its comprehensive discrimination ability is also superior to the comparison baseline. This confirms that multimodal fusion can more comprehensively capture emotional cues, providing a reliable solution for achieving more refined human-computer interaction.

Keywords: vocal emotion recognition; multimodal learning; attention mechanism; deep learning.

DOI: 10.1504/IJICT.2026.154383

International Journal of Information and Communication Technology, 2026 Vol.27 No.69, pp.69 - 91

Received: 26 Feb 2026
Accepted: 02 Apr 2026

Published online: 25 Jun 2026 *