Comparison of Support Vector Machine (SVM) and Naïve Bayes Algorithms in Diabetes Disease Classification

Main Article Content

Anita Desiani
Novi Rustiana Dewi
Muhammad Arhami
Dina Suzzete Sitorus
Suristhia Rahmadita

Abstract

High levels of sugar in the blood can cause diabetes. The longer people are unable to control glucose in their blood, the more complications it can cause, other diseases and even death. Early detection of diabetes is needed, one way is by carrying out data mining classification. Data mining classification in this research uses two algorithms, namely SVM (Support Vector Machine) and Naïve Bayes. This research compares the two algorithms using two methods, namely training split and k-fold cross validation which aims to get the best classification results in detecting diabetes. The best classification results are determined by calculating the average value of precision, recall and accuracy. Based on this research, the SVM algorithm with split percentage training produces average values for precision, recall and accuracy, namely 77%, 71.5%, 77.27%, while the SVM algorithm with k-fold cross validation produces average values for precision, recall , and accuracy is 77%, 72.5%, 71%. The Naïve Bayes algorithm with the split percentage training method produces average values for precision, recall and accuracy, namely 75.5%, 74.5%, 79%, while the Naïve Bayes algorithm with k-fold cross validation produces average values for precision, recall, and accuracy of 75.5%, 74.5%, 75%. The best classification result in detecting diabetes is the Naïve Bayes algorithm, the split percentage method, which provides the best accuracy, precision and recall values above 74%.

Downloads

Download data is not yet available.

Article Details

Section
Articles

References

[1] C. S. Sari, “Evaluasi Pemberian Informasi Obat Insulin Pada Pasien Rawat Jalan Rumah Sakit Pku Muhammadiyah Sekapuk,” Universitas Muhammadiyah Gresik, 2020.

[2] U. Hasanah, “Insulin Sebagai Pengatur kadar Gula Darah,” J. Kel. Sehat Sejah., vol. 11, no. 22, pp. 42–49, 2013.

[3] D. N. Anisa and Jumanto, “Klasifikasi Penyakit Diabetes Menggunakan Algoritma Naive Bayes,” Din. Inform., vol. 14, no. 1, pp. 33–42, 2022.

[4] E. D. Nurcahya, “Klasifikasi Penyakit Ayam Menggunakan Metode Support Vector Machine,” VOLT J. Ilm. Pendidik. Tek. Elektro, vol. 2, no. 1, p. 45, 2017.

[5] Kementrian Kesehatan Republik Indonesia, “Berat Badan Ideal Bantu Cegah Timbulnya Diabetes,” Biro Komunikasi dan Pelayanan Masyarakat, 2021.

[6] Badan Penelitian dan Pengembangan Kesehatan, “Laporan Riskesdas 2018 Nasional,” Lembaga Penerbit Balitbangkes. 2018.

[7] W. Hoaxiang and S. Smys, “Big Data Analysis and Perturbation Using Data Mining Algorithm,” J. Soft Comput. Paradig., vol. 3, no. 1, pp. 19–28, 2021.

[8] D. P. Utomo and M. Mesran, “Analisis Komparasi Metode Klasifikasi Data Mining dan Reduksi Atribut Pada Data Set Penyakit Jantung,” J. Media Inform. Budidarma, vol. 4, no. 2, p. 437, 2020.

[9] A. Damuri, U. Riyanto, H. Rusdianto, and M. Aminudin, “Implementasi Data Mining dengan Algoritma Naïve Bayes Untuk Klasifikasi Kelayakan Penerima Bantuan Sembako,” JURIKOM (Jurnal Ris. Komputer), vol. 8, no. 6, p. 219, 2021.

[10] A. Darmawan, N. Kustian, and W. Rahayu, “Implementasi Data Mining Menggunakan Model SVM Untuk Prediksi Kepuasan Pengunjung Taman Tabebuya,” STRING (Satuan Tulisan Ris. dan Inov. Teknol., vol. 2, no. 3, pp. 299–307, 2018.

[11] O. Arifin and T. B. Sasongko, “Analisa Perbandingan Tingkat Performansi Metode Support Vector Machine dan Naive Bayes Classifier untuk Klasifikasi Jalur Minat SMA,” Semin. Nas. Teknol. Inf. dan Multimed. 2018, pp. 67–72, 2018.

[12] A. Budianto, R. Ariyuana, and D. Maryono, “Perbandingan K-Nearest Neighbor (KNN) Dan Support Vector Machine (SVM) Dalam Pengenalan Karakter Plat Kendaraan Bermotor,” J. Ilm. Pendidik. Tek. dan Kejuru., vol. 11, no. 1, pp. 27–35, 2019.

[13] A. T. Novarina, E. Santoso, and Indriati, “Sistem Pakar Diagnosis Penyakit Hepatitis Menggunakan Metode Dempster Shafer,” J. Pengemb. Teknol. Inf. dan Ilmu Komput., vol. 2, no. 6, pp. 2252–2258, 2018.

[14] S. Sayed, M. Nassef, A. Badr, and I. Farag, “A Nested Genetic Algorithm For Feature Selection In High-Dimensional Cancer Microarray Datasets,” Expert Syst. Appl., vol. 121, pp. 233–243, 2019.

[15] H. Hermanto, A. Mustopa, A. Y. Kuntoro, and others, “Algoritma Klasifikasi Naive Bayes Dan Support Vector Machine Dalam Layanan Komplain Mahasiswa,” JITK (Jurnal Ilmu Pengetah. Dan Teknol. Komputer), vol. 5, no. 2, pp. 211–220, 2020.

[16] P. M. N. Dharmapatni and N. L. P. Merawati, “Penerapan Algoritma Support Vector Machine Dalam Sentimen Analisis Terkait Kenaikan Tarif BPJS Kesehatan,” J. Bumigora Inf. Technol., vol. 2, no. 2, pp. 105–112, 2020.

[17] R. Venkatesh, C. Balasubramanian, and M. Kaliappan, “Development of Big Data Predictive Analytics Model for Disease Prediction using Machine learning Technique,” J. Med. Syst., vol. 43, no. 8, 2019.

[18] N. Salmi and Z. Rustam, “Naïve Bayes Classifier Models for Predicting the Colon Cancer,” IOP Conf. Ser. Mater. Sci. Eng., vol. 546, no. 5, 2019.

[19] D. Cahya Putri Buani, “Penerapan Algoritma Naïve Bayes dengan Seleksi Fitur Algoritma Genetika Untuk Prediksi Gagal Jantung,” EVOLUSI J. Sains dan Manaj., vol. 9, no. 2, pp. 43–48, 2021.

[20] A. Desiani, “Perbandingan Implementasi Algoritma Naïve Bayes dan K-Nearest Neighbor Pada Klasifikasi Penyakit Hati,” Simkom, vol. 7, no. 2, pp. 104–110, 2022.

[21] D. A. Nasution, H. H. Khotimah, and N. Chamidah, “Perbandingan Normalisasi Data untuk Klasifikasi Wine Menggunakan Algoritma K-NN,” CESS (Journal Comput. Eng. Syst. Sci., vol. 4, no. 1, pp. 78–82, 2019.

[22] I. M. Parapat, M. T. Furqon, and Sutrisno, “Penerapan Metode Support Vector Machine (SVM) Pada Klasifikasi Penyimpangan Tumbuh Kembang Anak,” J. Pengemb. Teknol. Inf. dan Ilmu Komput., vol. 2, no. 10, pp. 3165–3166, 2018, [Online]. Available: http://j-ptiik.ub.ac.id

[23] D. A. Anggoro and N. D. Kurnia, “Comparison of Accuracy Level of Support Vector Machine (SVM) and K-Nearest Neighbors (KNN) Algorithms In Predicting Heart Disease,” Int. J., vol. 8, no. 5, pp. 1689–1694, 2020.

[24] P. T. D. P. Putu, M. F. Zambak, Suwarno, and P. Harahap, “Analisa Radiasi Sinar Matahari Terhadap Panel Surya 50 WP,” RELE (Rekayasa Elektr. dan Energi) J. Tek. Elektro, vol. 4, no. 1, pp. 48–54, 2021.

[25] A. Indriani, “Klasifikasi Data Forum dengan menggunakan Metode Naïve Bayes Classifier,” Semin. Nas. Apl. Teknol. Inf., pp. 1–10, 2014.

[26] B. P. Pratiwi, A. S. Handayani, and S. Sarjana, “Pengukuran Kinerja Sistem Kualitas Udara dengan Teknologi WSN Menggunakan Confusion Matrix,” J. Inform. Upgris, vol. 6, no. 2, Jan. 2021.

[27] E. Prasetyo, Data Mining: Konsep dan Aplikasi menggunakan MATLAB. Yogyakarta: ANDI, 2012.