Title: English digital transformation algorithm for distributed big data based on Spark
Authors: Xiaochao Yao
Addresses: Hainan Vocational University of Science and Technology, 18 Qiongshan Avenue, Meilan District, Haikou City, 570100, Hainan Province, China
Abstract: In the era of big data, the main challenge of digital transformation lies in the inability of traditional text clustering algorithms to efficiently process large-scale English data, resulting in low adaptability and delayed information extraction that hinder intelligent decision-making. This study was conducted to enhance the efficiency and interpretability of digital transformation in English text analysis. A distributed framework based on Apache Spark and an improved K-means algorithm is proposed to overcome the scalability and accuracy limitations of traditional methods. This method combines the density peak and maximum and minimum criteria, and significantly improves the efficiency and accuracy of clustering by accurately selecting the initial clustering centre and optimising the calculation process. Experimental results show that the improved algorithm has a clustering accuracy of 10.53% higher than that of the traditional algorithm on multiple datasets, and shows higher stability and performance when processing large-scale data. In summary, the English digital transformation algorithm based on Spark shows superior performance than traditional methods in a big data environment and has strong practical value.
Keywords: big data; distributed computing; Apache Spark; digital transformation algorithms in English.
DOI: 10.1504/IJIIDS.2026.156312
International Journal of Intelligent Information and Database Systems, 2026 Vol.18 No.8, pp.1 - 28
Received: 03 Jul 2025
Accepted: 24 Dec 2025
Published online: 10 Sep 2026 *


