Text classification using scores based k-NN approach and term to category relevance weighting scheme Online publication date: Wed, 10-Aug-2016
by Ahmed Ben Afia; Hamid Amiri
International Journal of Signal and Imaging Systems Engineering (IJSISE), Vol. 9, No. 4/5, 2016
Abstract: Text categorisation is the task of deciding whether a document belongs to a set of pre-specified classes of documents. To reach this goal, a TC system must include two basic stages. First stage consists on features extraction using a term weighting scheme. Second stage is the classification using a machine learning algorithm. After proposing, a new term to category relevance weighting scheme, called TF.IDF.TCR, we focus on finding a new algorithm to perform classification step. Results of our experiments, in which we use many classifiers, show promising performances. On the other hand, using relevance to category to improve the term's discriminating power appears to be inapplicable when classifying an unlabelled document. As a solution, we propose a k-NN based approach using scores calculating in order to resolve the problem of unknown category.
Existing subscribers:
Go to Inderscience Online Journals to access the Full Text of this article.
If you are not a subscriber and you just want to read the full contents of this article, buy online access here.Complimentary Subscribers, Editors or Members of the Editorial Board of the International Journal of Signal and Imaging Systems Engineering (IJSISE):
Login with your Inderscience username and password:
Want to subscribe?
A subscription gives you complete access to all articles in the current issue, as well as to all articles in the previous three years (where applicable). See our Orders page to subscribe.
If you still need assistance, please email subs@inderscience.com