Open Access Article

Title: A computational model for urban spatial aesthetics based on multimodal graph convolutional networks

Authors: Yiran Fan

Addresses: College of Digital Technology and Design, Chongqing Polytechnic University of Electronic Technology, Chongqing, 401331, China

Abstract: In urban spatial visual aesthetics, street scene elements are not independent; their spatial arrangement and semantic associations shape aesthetic experience. Yet existing deep learning models directly extract global features, ignoring this structured relationship and causing evaluation bias for complex scenes. This paper proposes a global visual-semantic graph convolutional network that constructs scene graphs from semantic segmentation, aggregates spatial proximity and semantic complementarity via graph convolution, and fuses with global visual features. On a large-scale multi-city street view dataset, the model achieves 81.23% average accuracy, 2.8 percentage points above the optimal ensemble model; visual entropy correlates at 0.84 with aesthetic score, and spatial order deviation drops to 0.107, indicating it captures scene complexity and spatial hierarchy more closely to human judgment.

Keywords: visual aesthetics; graph convolutional network; street view; multimodal fusion.

DOI: 10.1504/IJRIS.2026.155780

International Journal of Reasoning-based Intelligent Systems, 2026 Vol.18 No.20, pp.39 - 53

Received: 19 May 2026
Accepted: 22 Jun 2026

Published online: 13 Aug 2026 *