Open Access Article

Title: Collaborative optimisation of emotion regulation and audio synthesis based on PerformanceNet and multi-emotional music generation model

Authors: Li Chai

Addresses: Normal Department, Dazhou Vocational and Technical College, Dazhou, 635001, China

Abstract: In response to the three major challenges in AI music generation - limited chord representation, monotonous emotions, and low audio fidelity - this research proposes a novel end-to-end framework termed PEMF that integrates PerformanceNet with a multi-emotion music generation model. The core innovations include a structured four-dimensional chord encoding method using root, third, fifth, and crown notes to expand harmonic diversity to 60 chord types, a dual-encoding transformer architecture that independently processes melody and chord streams for superior structural coherence, and a fine-grained emotion regulation mechanism mapping pitch histograms and rhythm density parameters to Russell's two-dimensional emotion space for continuous control. For audio synthesis, an asymmetric U-net structure combined with a multi-band residual learning mechanism and a flooding loss strategy significantly enhances spectral fidelity and training stability. Experimental results demonstrate that PEMF achieves a chord vocabulary coverage near 1.0, an emotion recognition accuracy of 92.3% significantly outperforming symbolic transformer's 78.6%, a high-frequency energy retention rate of 89.1%, and a Fréchet audio distance of 0.5. System performance shows a 36.9% improvement in emotional consistency and a 64.2% reduction in latency compared to staged training, validating its efficacy in practical applications like music therapy and film scoring.

Keywords: PerformanceNet; multi-emotional music generation model; emotional regulation; audio synthesis.

DOI: 10.1504/IJICT.2026.153801

International Journal of Information and Communication Technology, 2026 Vol.27 No.57, pp.51 - 75

Received: 28 Oct 2025
Accepted: 04 Feb 2026

Published online: 26 May 2026 *