Title: Adaptive music generation by integrating improved VAE and improved GMVAE
Authors: Zhaoqing Ning
Addresses: Academy of Arts, Shangluo University, Shangluo, 726000, China
Abstract: This paper innovatively proposes an adaptive target-music generation model: it employs a controllable variational autoencoder (C-VAE) to construct decoupled structure/control latent variables, incorporates transformer-XL for modelling long-term dependencies, and combines a semantically guided modified variational autoencoder (S-GMVAE) to embed mode-emotion relationships into the latent space for controllable generation. On the MAESTRO and LMD datasets, the model achieves F1 = 93.76% and style matching = 91.84%. It maintains coherence = 90.16% even at 30% missing notes while exhibiting the lowest generation latency. Subjective evaluations reveal melody fluency, emotional authenticity, and semantic consistency all exceeding 4.6/5. Compared to PRNN, POP909-BART, MTR-VAE, and others, the model excels in both accuracy and real-time performance. Results from the experiment demonstrate that the proposed framework offers significant advantages in emotion-controlled style transfer and robust generation under missing information, providing effective support for intelligent composition, emotional soundtrack creation, and human-computer interaction music systems.
Keywords: music generation; adaptive; variational autoencoder; VAE; guided modified variational autoencoder; GMVAE; transformer.
DOI: 10.1504/IJRIS.2026.153460
International Journal of Reasoning-based Intelligent Systems, 2026 Vol.18 No.12, pp.18 - 37
Received: 26 Dec 2025
Accepted: 02 Feb 2026
Published online: 08 May 2026 *


