Open Access Article

Title: Intelligent assessment and teaching optimisation of college Russian translation based on transformer multimodal representation learning

Authors: Hongxue Yin; Yongping Wang

Addresses: School of Foreign Languages, Inner Mongolia University of Science and Technology, Inner Mongolia Baotou, 014010, China ' School of Automation and Electrical Engineering, Inner Mongolia University of Science and Technology, Inner Mongolia Baotou, 014010, China

Abstract: This study develops transformer-based multimodal intelligent evaluation model for college Russian translation instruction, tackling frequent pragmatic failures and context deficiency that lead to lagged feedback. It leverages XLM-RoBERTa to capture Russian's intricate morphological and syntactic features, adopts ViT for global visual context of accompanying images, and uses multi-head cross-attention (MCA) to deeply integrate and calibrate textual-visual semantics. Experiments on the improved multimodal corpus based on Wikipedia-based image text (WIT) show that the model's Pearson correlation coefficient to gauge the consistency of scoring is as high as 0.835, and the error diagnosis reaches 89.4%. In identifying high-order pragmatic inconsistency errors, compared with the ResNet+XLM-R (naive fusion) baseline model, its F1-score significantly improved to 0.87 (p < 0.01), with statistical significance verified. Highly consistent with expert scores in real teaching, the model proves valuable for accurate teacher feedback and cultivating students' text-image integrated translation thinking.

Keywords: multimodal representation learning; Russian translation teaching; translation intelligence evaluation; XLM-RoBERTa; vision transformer.

DOI: 10.1504/IJICT.2026.154378

International Journal of Information and Communication Technology, 2026 Vol.27 No.69, pp.1 - 20

Received: 26 Dec 2025
Accepted: 10 Feb 2026

Published online: 25 Jun 2026 *