Open Access Article

Title: Online oral English teaching system based on speech recognition technology and machine learning

Authors: Qingdi Si

Addresses: Puyang Medical College, Puyang, 457000, China

Abstract: To enhance the intelligence of pronunciation error detection and feedback precision in online oral English teaching, this study designs a system combining speech recognition and machine learning. Its core detection module uses the MFCCDBN model with feature fusion, and builds an SVMbased multiclassifier. Experimental data comes from the CSTR VCTK Corpus and the speech accent archive, containing 1,610 expertannotated phoneme samples. The model yields high accuracy for both samplesufficient and smallsample error types. Compared with LDASVM and Wav2Vec2.0SVM, it outperforms them in accuracy and standard error. Results prove the fusion model's stronger robustness and efficiency with limited data, offering a practical technical approach to boost learners' crosscultural communication competence.

Keywords: oral English teaching; pronunciation detection; Mel-frequency cepstral coefficient; MFCC; deep belief network; DBN; speech recognition; machine learning.

DOI: 10.1504/IJICT.2026.152857

International Journal of Information and Communication Technology, 2026 Vol.27 No.31, pp.22 - 43

Received: 19 Nov 2025
Accepted: 31 Dec 2025

Published online: 13 Apr 2026 *