Title: Deep learning-driven literary evaluation: fusing visualised representations and text embeddings
Authors: Junyong Cui
Addresses: Department of Discipline Inspection, Zhengzhou Vocational College of Industrial Safety, Zhengzhou, 451100, China
Abstract: Visualisations and their captions are central to data analysis, yet most models still process images and text separately and lose important cross-modal cues. To address this limitation, this paper proposes a fusion framework that jointly represents visualised elements and descriptive language within a shared transformer architecture. First, charts are decomposed into typed visual elements and encoded structurally; then, captions and analytic comments are contextualised and softly grounded to these elements; finally, a gated fusion module and multimodal transformer layers integrate both streams for retrieval, question answering and summary generation. Experiments on three figure-caption benchmarks show that the proposed model outperforms strong visual-only, text-only and dual-encoder baselines, raising top-10 accuracy by up to 2.6 percentage points and BLEU-4 scores by up to 2.8 points while reducing variance across runs.
Keywords: visualised data; text embeddings; multimodal fusion; transformer models; figure understanding; information retrieval.
DOI: 10.1504/IJICT.2026.153385
International Journal of Information and Communication Technology, 2026 Vol.27 No.43, pp.84 - 102
Received: 30 Dec 2025
Accepted: 31 Jan 2026
Published online: 06 May 2026 *


