Open Access Article

Title: A weighted multimodal fusion and deep learning framework for quantitative evaluation of music performance effects

Authors: Weiwei Yan; Dan Ma; Yuan Luo; Maohua Cheng

Addresses: Guangxi Science and Technology Normal University, Laibin 546199, China ' Guangxi Science and Technology Normal University, Laibin 546199, China ' Guangxi Science and Technology Normal University, Laibin 546199, China ' Guangxi Science and Technology Normal University, Laibin 546199, China

Abstract: Music performance evaluation has long relied on subjective judgement, causing inconsistency and poor reproducibility. This study presents a quantitative framework using weighted multimodal fusion and deep learning, integrating audio, video, and motion data to assess technical and expressive performance. It adopts convolutional networks for feature extraction, an LSTM module for temporal motion modelling, and weighted fusion for joint representation learning. Experimental results show that the proposed multimodal model achieves the best performance among all comparison models, with MAE of 1.76, RMSE of 2.48, R2 of 0.95, Pearson r of 0.96, and F1-score of 97.1%. The overall correlation between system scores and expert ratings reaches 0.88, with rhythm control and pitch stability exceeding 0.89. The study demonstrates that weighted multimodal fusion can effectively improve the consistency, stability, and interpretability of music performance evaluation, providing a reproducible analytical pathway for quantitative assessment of artistic performance.

Keywords: music performance effect; multimodal feature fusion; deep learning model; intelligent evaluation system.

DOI: 10.1504/IJRIS.2026.154113

International Journal of Reasoning-based Intelligent Systems, 2026 Vol.18 No.15, pp.15 - 30

Received: 05 Feb 2026
Accepted: 21 Apr 2026

Published online: 12 Jun 2026 *