Title: Data mining classification techniques - comparison for better accuracy in prediction of cardiovascular disease

Authors: Richa Sharma; Shailendra Narayan Singh; Sujata Khatri

Addresses: ASET, Amity University, Noida, India ' ASET, Amity University, Noida, India ' Deen Dayal Upadhyay College, Delhi University, India

Abstract: Cardiovascular disease is a broad term which includes stroke or any disorder in the cardiovascular system that has the heart at its centre. This disease is a critical cause of mortality every year across the globe. Data mining utilises a variety of techniques and algorithms that could help to draw some interesting conclusions about cardiovascular disease. Data mining in healthcare can assist in predicting disease. This study aims to gain knowledge from a heart disease dataset and analyse several data mining classification techniques seeking improved accuracy and a lesser error rate in the results. The data set for the experiment is chosen from the UCI machine learning repository database. The dataset is analysed using two different data mining tools, i.e., WEKA and Tanagra. The analysis was done using the 10 fold cross validation technique. The results show that the Naive Bayes algorithm and the C-PLS algorithm outperform others with an accuracy of 83.71% and 84.44% respectively.

Keywords: data mining; classification techniques; machine learning tools; cardiovascular disease; KNN; Naïve Bayes; C-PLS; decision tree.

DOI: 10.1504/IJDATS.2019.103756

International Journal of Data Analysis Techniques and Strategies, 2019 Vol.11 No.4, pp.356 - 373

Received: 17 Feb 2017
Accepted: 18 Oct 2017

Published online: 27 Nov 2019 *

Full-text access for editors Full-text access for subscribers Purchase this article Comment on this article