Open Access Article

Title: Multimodal emotion feature extraction and information fusion methods for video content

Authors: Jianxiao Ma

Addresses: School of Intelligent Manufacturing and Electrical Engineering, Nanyang Normal University, Nanyang, 473061, China; Collaborative Innovation Center of Intelligent Explosion-Proof Equipment, Henan Province, Henan, 473000, China

Abstract: Multimodal sentiment analysis remains challenging due to the difficulty of fusing heterogeneous data like facial expressions, speech, and pose. Unlike unimodal analysis, it is often hindered by problems of information redundancy, heterogeneity, and complex temporal dynamics. This paper proposes MERCAT, a Transformer-based model featuring a cross-modal self-attention mechanism to capture deep correlations between different modalities. This design enables highly efficient and context-aware inter-modal fusion. Extensive experiments on multiple benchmarks show that MERCAT achieves excellent performance. It notably excels in emotion classification, significantly improving accuracy and F1-score over strong baselines, and in emotion intensity prediction, where it substantially reduces error and improves correlation. The study conclusively verifies the efficacy of the cross-modal attention mechanism for information fusion, providing a robust and effective solution for advancing multimodal sentiment analysis.

Keywords: multimodal sentiment analysis; sentiment feature extraction; information fusion; transformer architecture; cross-modal self-attention mechanism; sentiment intensity prediction.

DOI: 10.1504/IJICT.2026.152548

International Journal of Information and Communication Technology, 2026 Vol.27 No.28, pp.43 - 59

Received: 26 Sep 2025
Accepted: 22 Oct 2025

Published online: 26 Mar 2026 *