Open Access Article

Title: A consistent simulation model for environmental art design generation driven by multimodal transformer

Authors: Li Ren

Addresses: College of Urban and Environmental Sciences, Henan Light Industry Vocational College, Zhengzhou, 450008, China

Abstract: Aiming at semantic disconnection and visual distortion between generated results and actual scenes in environmental art design, this paper proposes a consistent generation and simulation model based on multimodal transformer. Traditional methods have limitations in coordinating complex elements and ensuring spatial logic, hindering design implementation. By integrating multi-source information including text, sketches, and scene images, an end-to-end generation-simulation framework achieves consistent mapping from concept to high-fidelity visual output. Using the public dataset MIT ade20k, results show the model achieves significant improvements in visual fidelity (area under the curve 0.92, an increase of 8.2%) and user preference (normalised discounted cumulative gain @10 an increase of 15.7%), with all key indicators being statistically significant (p < 0.01). This confirms the model's effectiveness in enhancing automation and usability of environmental art design.

Keywords: multimodal transformer; environmental art; generative model; consistency simulation.

DOI: 10.1504/IJICT.2026.152918

International Journal of Information and Communication Technology, 2026 Vol.27 No.33, pp.69 - 88

Received: 25 Dec 2025
Accepted: 22 Jan 2026

Published online: 14 Apr 2026 *