Title: Intelligent assessment and teaching optimisation of college Russian translation based on transformer multimodal representation learning
Authors: Hongxue Yin; Yongping Wang
Addresses: School of Foreign Languages, Inner Mongolia University of Science and Technology, Inner Mongolia Baotou, 014010, China ' School of Automation and Electrical Engineering, Inner Mongolia University of Science and Technology, Inner Mongolia Baotou, 014010, China
Abstract: This study develops transformer-based multimodal intelligent evaluation model for college Russian translation instruction, tackling frequent pragmatic failures and context deficiency that lead to lagged feedback. It leverages XLM-RoBERTa to capture Russian's intricate morphological and syntactic features, adopts ViT for global visual context of accompanying images, and uses multi-head cross-attention (MCA) to deeply integrate and calibrate textual-visual semantics. Experiments on the improved multimodal corpus based on Wikipedia-based image text (WIT) show that the model's Pearson correlation coefficient to gauge the consistency of scoring is as high as 0.835, and the error diagnosis reaches 89.4%. In identifying high-order pragmatic inconsistency errors, compared with the ResNet+XLM-R (naive fusion) baseline model, its F1-score significantly improved to 0.87 (p < 0.01), with statistical significance verified. Highly consistent with expert scores in real teaching, the model proves valuable for accurate teacher feedback and cultivating students' text-image integrated translation thinking.
Keywords: multimodal representation learning; Russian translation teaching; translation intelligence evaluation; XLM-RoBERTa; vision transformer.
DOI: 10.1504/IJICT.2026.154378
International Journal of Information and Communication Technology, 2026 Vol.27 No.69, pp.1 - 20
Received: 26 Dec 2025
Accepted: 10 Feb 2026
Published online: 25 Jun 2026 *


