Title: Interactive AI virtual human synthesis technology based on multimedia visual images
Authors: Huichao Guo; Yongjian Hua; Ming Lei; Runhua Huang
Addresses: School of Safety Science and Emergency Management, Wuhan University of Technology, Wuhan, 430000, Hubei, China; Zhejiang Business College, Hangzhou, 310000, Zhejiang, China ' Zhejiang Business College, Hangzhou, 310000, Zhejiang, China ' Zhejiang Business College, Hangzhou, 310000, Zhejiang, China ' The Chinese University of Hong Kong, Shenzhen, Shenzhen, 518100, Guangdong, China; Zhejiang Business College, Hangzhou, 310000, Zhejiang, China
Abstract: In the specific implementation, this paper first developed a multimodal feature-extraction process to extract semantic and rhythmic information from speech. It combined the emotional tags in the text with facial keypoint information from diverse visual inputs (e.g., static photos, dynamic video sequences, and 3D face models) to form a unified feature vector. Secondly, this paper designed MAM to dynamically adjust the weights of audio, text, and visual modalities and to synchronise expression changes with speech rhythm through real-time speech stream analysis and frame-level feature alignment. The dynamic weight adjustment in MAM operates along both the temporal dimension, aligning frame-level features across modalities, and the channel dimension, reweighting feature maps to emphasise modality-specific contributions to expression synthesis. This synthesis technology introduces a DM based on the U-Net structure, uses fused features as conditional inputs, and drives the progressive generation of virtual human image sequences.
Keywords: multimedia visual images; diffusion model; DM; multimodal attention mechanism; MAM; virtual human synthesis; emotion perception.
DOI: 10.1504/IJDMB.2026.154769
International Journal of Data Mining and Bioinformatics, 2026 Vol.30 No.7, pp.26 - 50
Received: 30 Dec 2025
Accepted: 13 Apr 2026
Published online: 13 Jul 2026 *


