Title: Enhancing digital art style recognition via a hybrid vision transformer and lightweight CNN with attention mechanisms

Authors: Ying Zhang; Fangzheng Lv

Addresses: Information Engineering Department, Yantai Vocational College, Yantai, 264670, China ' Department of Art Design and Education, Yantai Vocational College, Yantai, 264670, China

Abstract: Current methods for art style recognition often struggle to capture local details and balance global and texture features, leading to vague style representation during multi-scale fusion. To address limitations in capturing local details and balancing global and texture features in digital art style classification, we propose a hybrid model based on vision transformer and lightweight CNN. The model adopts a multi-scale attention-weighted fusion strategy and multi-task learning to optimise classification and style reconstruction simultaneously. Experimental results show the proposed model achieves a Kappa coefficient of 0.96, significantly outperforming baselines (0.78 and 0.74), with a classification accuracy of 96.3% and an F1 score of 97.8%. These findings demonstrate the model's strong performance in digital art style recognition, promoting intelligent applications in cultural and creative industries. The proposed model significantly improves accuracy and feature representation in digital art style recognition, supporting the intelligent development of the cultural and creative industry and advancing deep learning applications in art analysis.

Keywords: vision transformer; ViT; lightweight convolutional neural network; attention mechanism; digital art; style recognition.

DOI: 10.1504/IJICA.2025.149778

International Journal of Innovative Computing and Applications, 2025 Vol.15 No.4, pp.236 - 245

Received: 23 May 2025
Accepted: 18 Aug 2025

Published online: 12 Nov 2025 *

Full-text access for editors Full-text access for subscribers Purchase this article Comment on this article