Title: A lexical-semantics-based method for multi label text categorisation using word net

Authors: Rajni Jindal; Shweta Taneja

Addresses: Computer Engineering Department, Delhi Technological University, New Delhi – 110042, India ' Computer Engineering Department, Delhi Technological University, New Delhi – 110042, India

Abstract: Text categorisation is an upcoming area in the field of text mining. The text documents possess huge number of features due to their unstructured nature. In this paper, an algorithm for multi label categorisation of text documents based on the concepts of lexical and semantics using word net (MC-LSW) is proposed. The proposed algorithm is based on the concepts of lexical (tokens) and semantics of a language. It aims at minimising the number of tokens used for categorising text documents. MC-LSW uses word net to extract the semantic information of tokens. The proposed algorithm is implemented and tested on five datasets of text domain and is compared with the existing multi label categorisation algorithms. The proposed algorithm (MC-LSW) shows more efficient and promising results in terms of space and time complexity than the existing methods. Accuracy and precision measures have been improved by the proposed algorithm as well as hamming loss has been reduced.

Keywords: multi label text categorisation; lexical analysis; semantic analysis; word net.

DOI: 10.1504/IJDMMM.2017.088412

International Journal of Data Mining, Modelling and Management, 2017 Vol.9 No.4, pp.340 - 360

Received: 22 Mar 2016
Accepted: 16 Feb 2017

Published online: 06 Dec 2017 *

Full-text access for editors Full-text access for subscribers Purchase this article Comment on this article