Title: Clustering news articles using efficient similarity measure and N-grams

Authors: Desmond Bala Bisandu; Rajesh Prasad; Musa Muhammad Liman

Addresses: School of IT and Computing, American University of Nigeria, Yola, Adamawa State, Nigeria ' School of IT and Computing, American University of Nigeria, Yola, Adamawa State, Nigeria ' School of IT and Computing, American University of Nigeria, Yola, Adamawa State, Nigeria

Abstract: The rapid progress of information technology and web makes it easier to store huge amount of collected textual information, e.g., blogs, news articles, e-mail messages, reviews and forum postings. The growing size of textual dataset with high-dimensions and natural language pose a big challenge making it hard for such information to be categorised efficiently. Document clustering is an automatic unsupervised machine learning technique that aimed at grouping related set of items into clusters or subsets. The target is creating clusters with high internal coherence, but different from each other substantially. This paper presents a new document clustering technique using N-grams and efficient similarity measure known as 'improved sqrt-cosine similarity measure'. Comprehensive experiments are conducted to evaluate our proposed clustering technique and compared with an existing method. The results of the experiments show that our proposed clustering technique outperforms the existing techniques.

Keywords: information retrieval; clustering; similarity measures; data mining; N-grams.

DOI: 10.1504/IJKEDM.2018.095525

International Journal of Knowledge Engineering and Data Mining, 2018 Vol.5 No.4, pp.333 - 348

Received: 30 Apr 2018
Accepted: 22 Jul 2018

Published online: 08 Oct 2018 *

Full-text access for editors Full-text access for subscribers Purchase this article Comment on this article