Authors: Desmond Bala Bisandu; Rajesh Prasad; Musa Muhammad Liman
Addresses: School of IT and Computing, American University of Nigeria, Yola, Adamawa State, Nigeria ' School of IT and Computing, American University of Nigeria, Yola, Adamawa State, Nigeria ' School of IT and Computing, American University of Nigeria, Yola, Adamawa State, Nigeria
Abstract: The rapid progress of information technology and web makes it easier to store huge amount of collected textual information, e.g., blogs, news articles, e-mail messages, reviews and forum postings. The growing size of textual dataset with high-dimensions and natural language pose a big challenge making it hard for such information to be categorised efficiently. Document clustering is an automatic unsupervised machine learning technique that aimed at grouping related set of items into clusters or subsets. The target is creating clusters with high internal coherence, but different from each other substantially. This paper presents a new document clustering technique using N-grams and efficient similarity measure known as 'improved sqrt-cosine similarity measure'. Comprehensive experiments are conducted to evaluate our proposed clustering technique and compared with an existing method. The results of the experiments show that our proposed clustering technique outperforms the existing techniques.
Keywords: information retrieval; clustering; similarity measures; data mining; N-grams.
International Journal of Knowledge Engineering and Data Mining, 2018 Vol.5 No.4, pp.333 - 348
Available online: 24 Sep 2018 *Full-text access for editors Access for subscribers Purchase this article Comment on this article