Int. J. of Data Mining, Modelling and Management   »   2017 Vol.9, No.4

 

 

Title: A lexical-semantics-based method for multi label text categorisation using word net

 

Authors: Rajni Jindal; Shweta Taneja

 

Addresses:
Computer Engineering Department, Delhi Technological University, New Delhi – 110042, India
Computer Engineering Department, Delhi Technological University, New Delhi – 110042, India

 

Abstract: Text categorisation is an upcoming area in the field of text mining. The text documents possess huge number of features due to their unstructured nature. In this paper, an algorithm for multi label categorisation of text documents based on the concepts of lexical and semantics using word net (MC-LSW) is proposed. The proposed algorithm is based on the concepts of lexical (tokens) and semantics of a language. It aims at minimising the number of tokens used for categorising text documents. MC-LSW uses word net to extract the semantic information of tokens. The proposed algorithm is implemented and tested on five datasets of text domain and is compared with the existing multi label categorisation algorithms. The proposed algorithm (MC-LSW) shows more efficient and promising results in terms of space and time complexity than the existing methods. Accuracy and precision measures have been improved by the proposed algorithm as well as hamming loss has been reduced.

 

Keywords: multi label text categorisation; lexical analysis; semantic analysis; word net.

 

DOI: 10.1504/IJDMMM.2017.10009450

 

Int. J. of Data Mining, Modelling and Management, 2017 Vol.9, No.4, pp.340 - 360

 

Date of acceptance: 16 Feb 2017
Available online: 01 Dec 2017

 

 

Editors Full text accessAccess for SubscribersPurchase this articleComment on this article