Open Access Article

Title: Cross-modal understanding-driven reconstruction of style coherence in AIGC-generated artistic images

Authors: Luhe Zhao; Wei Zhou; Wenjin He; Li Tan

Addresses: School of Design and Art, University of South China, Hengyang, 421001, China ' School of Design and Art, University of South China, Hengyang, 421001, China ' School of Design and Art, University of South China, Hengyang, 421001, China ' School of Design and Art, University of South China, Hengyang, 421001, China

Abstract: Creative teams increasingly ask generative systems to deliver series of images that look like they belong together. To stop style drift across scenes and tiles, we propose art coherence reconstruction, a cross-modal, set-aware pipeline for artistic image generation. In our scheme, first, we parse text and references to disentangle content from style; then, we fuse statistics and descriptors into a compact style code backed by a small memory; finally, we apply hierarchical modulation and set-level losses to steer colour, stroke, and grain consistently. On storyboards, character sheets, and mosaics, ArtCoRe lifts the coherence score from 0.56 to 0.71, halves drift from 0.22 to 0.10, raises alignment from 0.62 to 0.69, and cuts macro FID from 56.2 to 48.7. The result is stable palettes, cleaner seams, and predictable control for real projects.

Keywords: AIGC; cross-modal understanding; style coherence; diffusion models; hierarchical modulation; style memory; set-level evaluation.

DOI: 10.1504/IJICT.2026.152656

International Journal of Information and Communication Technology, 2026 Vol.27 No.30, pp.93 - 107

Received: 10 Dec 2025
Accepted: 12 Jan 2026

Published online: 01 Apr 2026 *