The Impact of Balancing Techniques and Feature Selection on Machine Learning Models for Diabetes Detection

dc.contributor.authorSinap, Vahid
dc.date.accessioned2026-08-12T15:06:04Z
dc.date.issued2025
dc.departmentFırat Üniversitesi
dc.description.abstractThe detection of diabetes is crucial for effective management and prevention of the disease, which poses significant health risks globally. This study introduces a novel approach to diabetes detection by combining advanced data balancing techniques and feature selection methods, including Lasso (L1) regularization, to enhance the performance of predictive models in imbalanced datasets. Techniques such as Random Under Sampling (RUS), Adaptive Synthetic Sampling (ADASYN), and Synthetic Minority Over-sampling Technique (SMOTE) were employed alongside models including Random Forest (RF), CatBoost (CB), Extreme Gradient Boosting (XGB), K-Nearest Neighbors (KNN), Gaussian Naive Bayes (GNB), Logistic Regression (LR), and Gradient Boosting (GB) to assess their impact on model accuracy and generalization capabilities. The findings reveal that the RF model achieved the highest accuracy of 93.25% when utilizing the SMOTE technique, underscoring the importance of appropriate data handling strategies in improving predictive outcomes. Furthermore, when all features were utilized without selection, the RF model attained an accuracy of 95.31%, indicating the model’s capacity to capture complex patterns when feature richness is maximized. The comprehensive methodology used in the study achieved a higher accuracy in diabetes detection than research in the literature and provided important outputs for developing reliable prediction models in healthcare.
dc.description.abstractDiyabet, küresel ölçekte önemli sağlık riskleri oluşturmaktadır. Diyabetin tespiti, hastalığın etkili yönetimi ve önlenmesi için büyük önem taşımaktadır. Bu çalışma, dengesiz veri setlerinde diyabet tespiti için çeşitli dengeleme tekniklerini ve Lasso (L1) düzenlemesi de dahil olmak üzere özellik seçim yöntemlerini birleştirerek diyabet tespitine yeni bir yaklaşım getirmektedir. Çalışmada, Random Under Sampling (RUS), Adaptive Synthetic Sampling (ADASYN) ve Synthetic Minority Over-sampling Technique (SMOTE) gibi teknikler, Random Forest (RF), CatBoost (CB), Extreme Gradient Boosting (XGB), K-En Yakın Komşu (KNN), Gaussian Naive Bayes (GNB), Lojistik Regresyon (LR) ve Gradient Boosting (GB) modelleri ile kullanılarak bu tekniklerin model doğruluğu ve genelleme yetenekleri üzerindeki etkileri değerlendirilmiştir. Bulgular, SMOTE tekniği kullanıldığında RF modelinin %93,25 ile en yüksek doğruluğa ulaştığını göstermektedir, bu da uygun veri işleme stratejilerinin tahmin sonuçlarını iyileştirmede önemini vurgulamaktadır. Ayrıca, özellik seçimi yapılmaksızın tüm özellikler kullanıldığında, RF modeli %95,31 doğruluk elde etmiş ve bu da özellik zenginliği maksimize edildiğinde modelin karmaşık desenleri yakalama kapasitesini ortaya koymaktadır. Araştırmada kullanılan kapsamlı metodoloji, diyabet tespitinde literatürdeki araştırmalardan yüksek bir doğruluğa ulaşmış ve sağlık hizmetlerinde güvenilir tahmin modelleri geliştirmek için önemli çıktılar sağlamıştır.
dc.identifier.doi10.35234/fumbd.1556260
dc.identifier.endpage320
dc.identifier.issn1308-9072
dc.identifier.issue1
dc.identifier.startpage303
dc.identifier.urihttps://doi.org/10.35234/fumbd.1556260
dc.identifier.urihttps://hdl.handle.net/11508/27700
dc.identifier.volume37
dc.language.isoen
dc.publisherFırat University
dc.publisherFırat Üniversitesi
dc.relation.ispartofFırat University Journal of Engineering Science
dc.relation.ispartofFırat Üniversitesi Mühendislik Bilimleri Dergisi
dc.relation.publicationcategoryMakale - Ulusal Hakemli Dergi - Kurum Öğretim Elemanı
dc.rightsinfo:eu-repo/semantics/openAccess
dc.snmzKA_DergiPark_20260511
dc.subjectMachine Learning (Other)
dc.subjectMakine Öğrenme (Diğer)
dc.titleThe Impact of Balancing Techniques and Feature Selection on Machine Learning Models for Diabetes Detection
dc.title.alternativeDengeleme Tekniklerinin ve Özellik Seçiminin Diyabet Tespitinde Makine Öğrenmesi Modelleri Üzerindeki Etkisi
dc.typeArticle

Dosyalar