Title: YOLOv8-enhanced temporal modelling for semantic classification of ethically sensitive video content
Authors: Chelsi Sen; Kuldeep Kumar Yogi
Addresses: Department of Computer Science and Engineering, Banasthali Vidyapith, Niwai, Jaipur, Rajasthan – 304022, India ' Department of Computer Science and Engineering, Banasthali Vidyapith, Niwai, Jaipur, Rajasthan – 304022, India
Abstract: This paper presents a smart, efficient, and interpretable deep learning framework for content-based video classification, specifically targeting context-rich categories such as vulgar, kissing, accidental, abusive, and fighting scenes. The proposed approach employs YOLOv8-based object detection to guide an attention-weighted keyframe selection mechanism, thereby ensuring that only the most informative frames are retained for the processing. These selected keyframes were passed through a hybrid deep architecture combining ResNet50 for spatial feature extraction and Bi-LSTM for temporal sequence modelling. Extensive experiments were conducted on benchmark datasets, including UCF101, NudeNet, CADP, and a custom dataset curated for rare categories. An accuracy of 92.34% and an F1-score of 91.12% were achieved by the proposed model, performing better than the traditional models of 3D CNN, Conv-LSTM, and ResNet50 with dense frame sampling by 4-8%. Class AUC values of more than 0.94 were obtained in the analysis of the proposed model.
Keywords: attention mechanism; classification framework; deep learning; DL; LSTM; learning methods; neural networks; sensitive content; YOLO.
DOI: 10.1504/IJICA.2026.154204
International Journal of Innovative Computing and Applications, 2026 Vol.15 No.5, pp.300 - 310
Received: 08 Dec 2025
Accepted: 27 Jan 2026
Published online: 16 Jun 2026 *