Open Access Article

Title: Timbre interpretable representation modelling method integrating self-supervised learning and music theory descriptors

Authors: Yousheng Cui; Jiaying He

Addresses: Sofia National Academy of Music, Sophia 1505, Bulgaria ' Faculty of Physics and Information Engineering, Zhaotong University, Zhaotong, Yunnan, 657000, China

Abstract: To enhance the semantic interpretability and structural generalisation of timbre representation, this paper presents an interpretable timbre modelling method combining self-supervised learning and music theory descriptors, applicable to music generation, audio retrieval and intelligent arrangement. We build a multi-task self-supervised framework with joint mask prediction and contrastive learning for hierarchical multi-scale time-frequency feature modelling. Tonality, rhythmic tension, harmonic stability and other descriptors are adopted, along with a prior-guided fusion mechanism to strengthen latent space semantic controllability. Comprehensive objective and subjective evaluations are conducted via dedicated metrics and a human perception system. Experiments on timbre classification, clustering and style transfer demonstrate that our method outperforms BEATs and SimCLR-Audio with notable gains in accuracy, F1-score, ARI and other indicators. Interpretability metrics are boosted by over 50%, and subjective perception scores are also improved. The method excels in timbre modelling and semantic mapping, offering a viable solution for understanding-driven intelligent music processing.

Keywords: timbre representation modelling; self-supervised learning; music theory descriptor; interpretable learning.

DOI: 10.1504/IJICT.2026.153317

International Journal of Information and Communication Technology, 2026 Vol.27 No.40, pp.26 - 48

Received: 30 Oct 2025
Accepted: 04 Dec 2025

Published online: 01 May 2026 *