Title: Data mining classification techniques - comparison for better accuracy in prediction of cardiovascular disease
Authors: Richa Sharma; Shailendra Narayan Singh; Sujata Khatri
Addresses: ASET, Amity University, Noida, India ' ASET, Amity University, Noida, India ' Deen Dayal Upadhyay College, Delhi University, India
Abstract: Cardiovascular disease is a broad term which includes stroke or any disorder in the cardiovascular system that has the heart at its centre. This disease is a critical cause of mortality every year across the globe. Data mining utilises a variety of techniques and algorithms that could help to draw some interesting conclusions about cardiovascular disease. Data mining in healthcare can assist in predicting disease. This study aims to gain knowledge from a heart disease dataset and analyse several data mining classification techniques seeking improved accuracy and a lesser error rate in the results. The data set for the experiment is chosen from the UCI machine learning repository database. The dataset is analysed using two different data mining tools, i.e., WEKA and Tanagra. The analysis was done using the 10 fold cross validation technique. The results show that the Naive Bayes algorithm and the C-PLS algorithm outperform others with an accuracy of 83.71% and 84.44% respectively.
Keywords: data mining; classification techniques; machine learning tools; cardiovascular disease; KNN; Naïve Bayes; C-PLS; decision tree.
DOI: 10.1504/IJDATS.2019.103756
International Journal of Data Analysis Techniques and Strategies, 2019 Vol.11 No.4, pp.356 - 373
Received: 17 Feb 2017
Accepted: 18 Oct 2017
Published online: 27 Nov 2019 *