Open Access Article

Title: Optimising English translation vocabulary selection based on corpus statistics and probabilistic modelling

Authors: Xiaojing Shen; Luping Zhang

Addresses: Basic Courses Teaching Department, Xinxiang Vocational and Technical College, Xinxiang, 453000, China ' Basic Courses Teaching Department, Xinxiang Vocational and Technical College, Xinxiang, 453000, China

Abstract: Lexical ambiguity is a core challenge in machine translation, such as translating 'apple' as either 'fruit' or 'Apple Inc.' depending on context. While existing neural machine translation models produce fluent output, they often exhibit bias in selecting specialised terminology and low-frequency words. To address this issue, this study innovatively combines statistical patterns from large-scale corpora with the probabilistic modelling capabilities of neural networks to construct a lexical selection optimisation framework. Experiments on the publicly available workshop on machine translation English-German translation dataset demonstrate that this approach improves the bilingual evaluation understudy score from 31.2 to 33.3 while significantly reducing the translation error rate from 52.1% to 49.8%. This confirms that integrating statistical prior knowledge effectively enhances machine translation accuracy and lexical consistency.

Keywords: machine translation; lexical choice corpus statistics; probabilistic modelling.

DOI: 10.1504/IJICT.2026.153514

International Journal of Information and Communication Technology, 2026 Vol.27 No.45, pp.62 - 79

Received: 16 Dec 2025
Accepted: 20 Jan 2026

Published online: 12 May 2026 *