Title: Deep learning-based automatic labelling of English syntactic variation and cross-dialect comparison
Authors: Yanli Jia; Xinhua Yuan
Addresses: School of Foreign Languages, Huainan Normal University, Huainan, 232038, China ' School of International Studies, Wenzhou Business College, Wenzhou, 325035, China
Abstract: With the growing demand for precise cross-dialect syntactic analysis, this work proposes an end-to-end framework that automatically labels English syntactic variants and quantifies their distribution across dialects. The approach integrates parser-generated silver annotations, a human-audited gold subset, and a dual-head neural model combining CRF-based sequence tagging and span classification. Domain adaptation with gradient reversal, moment matching, and supervised contrastive learning enhances robustness to dialectal shift, while probability calibration ensures accurate rate estimation. Evaluations on multi-source corpora covering American, British, Australian, and Indian English show that the proposed model improves out-of-dialect macro-F1 by 6.9 points over a strong RoBERTa baseline, reduces domain divergence in encoder space by over 55%, and recovers stable, interpretable contrasts for variants such as that-complementiser drop, particle movement, and dative alternation.
Keywords: syntactic variation; automatic labelling; cross-dialect analysis; CRF; contrastive learning.
DOI: 10.1504/IJICT.2026.152528
International Journal of Information and Communication Technology, 2026 Vol.27 No.26, pp.65 - 83
Received: 27 Oct 2025
Accepted: 08 Dec 2025
Published online: 25 Mar 2026 *


