Authors: Rajni Jindal; Shweta Taneja
Addresses: Computer Engineering Department, Delhi Technological University, New Delhi – 110042, India ' Computer Engineering Department, Delhi Technological University, New Delhi – 110042, India
Abstract: Text categorisation is an upcoming area in the field of text mining. The text documents possess huge number of features due to their unstructured nature. In this paper, an algorithm for multi label categorisation of text documents based on the concepts of lexical and semantics using word net (MC-LSW) is proposed. The proposed algorithm is based on the concepts of lexical (tokens) and semantics of a language. It aims at minimising the number of tokens used for categorising text documents. MC-LSW uses word net to extract the semantic information of tokens. The proposed algorithm is implemented and tested on five datasets of text domain and is compared with the existing multi label categorisation algorithms. The proposed algorithm (MC-LSW) shows more efficient and promising results in terms of space and time complexity than the existing methods. Accuracy and precision measures have been improved by the proposed algorithm as well as hamming loss has been reduced.
Keywords: multi label text categorisation; lexical analysis; semantic analysis; word net.
International Journal of Data Mining, Modelling and Management, 2017 Vol.9 No.4, pp.340 - 360
Received: 22 Mar 2016
Accepted: 16 Feb 2017
Published online: 01 Dec 2017 *