Title: Enhancing digital art style recognition via a hybrid vision transformer and lightweight CNN with attention mechanisms
Authors: Ying Zhang; Fangzheng Lv
Addresses: Information Engineering Department, Yantai Vocational College, Yantai, 264670, China ' Department of Art Design and Education, Yantai Vocational College, Yantai, 264670, China
Abstract: Current methods for art style recognition often struggle to capture local details and balance global and texture features, leading to vague style representation during multi-scale fusion. To address limitations in capturing local details and balancing global and texture features in digital art style classification, we propose a hybrid model based on vision transformer and lightweight CNN. The model adopts a multi-scale attention-weighted fusion strategy and multi-task learning to optimise classification and style reconstruction simultaneously. Experimental results show the proposed model achieves a Kappa coefficient of 0.96, significantly outperforming baselines (0.78 and 0.74), with a classification accuracy of 96.3% and an F1 score of 97.8%. These findings demonstrate the model's strong performance in digital art style recognition, promoting intelligent applications in cultural and creative industries. The proposed model significantly improves accuracy and feature representation in digital art style recognition, supporting the intelligent development of the cultural and creative industry and advancing deep learning applications in art analysis.
Keywords: vision transformer; ViT; lightweight convolutional neural network; attention mechanism; digital art; style recognition.
DOI: 10.1504/IJICA.2025.149778
International Journal of Innovative Computing and Applications, 2025 Vol.15 No.4, pp.236 - 245
Received: 23 May 2025
Accepted: 18 Aug 2025
Published online: 12 Nov 2025 *