Title: A consistent simulation model for environmental art design generation driven by multimodal transformer
Authors: Li Ren
Addresses: College of Urban and Environmental Sciences, Henan Light Industry Vocational College, Zhengzhou, 450008, China
Abstract: Aiming at semantic disconnection and visual distortion between generated results and actual scenes in environmental art design, this paper proposes a consistent generation and simulation model based on multimodal transformer. Traditional methods have limitations in coordinating complex elements and ensuring spatial logic, hindering design implementation. By integrating multi-source information including text, sketches, and scene images, an end-to-end generation-simulation framework achieves consistent mapping from concept to high-fidelity visual output. Using the public dataset MIT ade20k, results show the model achieves significant improvements in visual fidelity (area under the curve 0.92, an increase of 8.2%) and user preference (normalised discounted cumulative gain @10 an increase of 15.7%), with all key indicators being statistically significant (p < 0.01). This confirms the model's effectiveness in enhancing automation and usability of environmental art design.
Keywords: multimodal transformer; environmental art; generative model; consistency simulation.
DOI: 10.1504/IJICT.2026.152918
International Journal of Information and Communication Technology, 2026 Vol.27 No.33, pp.69 - 88
Received: 25 Dec 2025
Accepted: 22 Jan 2026
Published online: 14 Apr 2026 *


