Open Access Article

Title: Development of deep learning combined with spectrum analysis to separate human voice from musical instrument sound in music production

Authors: Fei Xue

Addresses: School of Music, Nanjing Xiaozhuang University, Nanjing, 211171, China

Abstract: This study proposes a deep learning-based system for separating human voice from musical instrument sound in music production. A time-frequency fusion framework is constructed by integrating constitutional neural networks with long-short-term memory networks to model overlapping spectral and temporal features. An attention mechanism and multi-scale spectral analysis enhance stability under reverberation and multi-source interference. Experiments on MUSDB18 and MIR-1K show that the method achieves an SNR of 14.2 dB, a PESQ of 3.12, and a speech intelligibility score of 0.91, outperforming traditional approaches by about 35%. With GPU acceleration, system latency remains below 85 ms, meeting real-time monitoring requirements. The results demonstrate that deep learning-driven time-frequency fusion improves separation fidelity and provides an effective technical basis for music production, speech enhancement, and intelligent audio engineering.

Keywords: deep learning; spectrum analysis; sound source separation; vocal extraction; music making.

DOI: 10.1504/IJRIS.2026.152552

International Journal of Reasoning-based Intelligent Systems, 2026 Vol.18 No.10, pp.17 - 30

Received: 04 Nov 2025
Accepted: 26 Dec 2025

Published online: 26 Mar 2026 *