Comparative evaluation of machine learning models for predicting Cimbex quadrimaculata population density across multiple problem formulations

dc.contributor.authorGural, Yunus
dc.date.accessioned2026-08-12T17:28:44Z
dc.date.issued2026
dc.departmentFırat Üniversitesi
dc.description.abstractThe high variability and nonlinear relationships between environmental variables (such as temperature, relative humidity, and altitude) in ecological datasets prevent classical statistical models from obtaining accurate predictions. This study aimed to compare and investigate the performance of AI-based machine learning methods in analyzing complex ecological data structures. An agricultural dataset containing meteorological and vegetation variables was used as the representative case study. This dataset is based on population observations of Cimbex quadrimaculata in Diyarbak & imath;r (E & gbreve;il) and Elaz & imath;& gbreve; (Keban) provinces in T & uuml;rkiye between 2020 and 2022. Three different modeling approaches (binary classification, multiclass classification, and regression) were applied to the same data. This three-approach design enabled a systematic comparison of model performance, generalizability, and explainability on the same dataset using different definitions of the target variable. For classification tasks, the model performance was evaluated using accuracy, F1 score, and AUC metrics under a stratified 10-fold cross-validation scheme. Regression models, on the other hand, were assessed within a nested cross-validation framework using R & sup2;, root mean square error (RMSE), mean absolute error (MAE). Ensemble-based boosting AI algorithms (Gradient Boosting, XGBoost, and LightGBM) demonstrated high accuracy and generalizability in characterizing the highly nonlinear relationships, nested effects, and non-additive interactions among multiple variables. Furthermore, the SHAP analysis improved the interpretability of the models and revealed that temperature- and humidity-related variables were consistently among the most influential predictors in the model predictions. Comparative performance evaluations of machine learning models showed that Gradient Boosting (94.3% accuracy, 0.983 AUC) and XGBoost (84.6% accuracy) were the strongest predictors in binary classification scenarios and overall analyses, respectively. In regression analyses, LightGBM and Random Forest algorithms stood out with cross-validation performances of approximately R & sup2; approximate to 0.73. In particular, the success of ensemble-based learning methods in capturing multidimensional relationships in ecological datasets explains the high predictive accuracy and robustness of these models across complex ecological data structures.
dc.description.sponsorshipFirat oeniversitesi [FF.25.51]
dc.description.sponsorshipThis study was supported by Firat University under the project number FF.25.51.This research was supported by Firat University, grant number FF.25.51. The funders had no role in study design, data collection and analysis, or preparation of the manuscript.
dc.identifier.doi10.1371/journal.pone.0346494
dc.identifier.issn1932-6203
dc.identifier.issue4
dc.identifier.pmid41931572
dc.identifier.scopus2-s2.0-105034959465
dc.identifier.scopusqualityQ1
dc.identifier.urihttps://doi.org/10.1371/journal.pone.0346494
dc.identifier.urihttps://hdl.handle.net/11508/55426
dc.identifier.volume21
dc.identifier.wosWOS:001732903500005
dc.identifier.wosqualityQ2
dc.indekslendigikaynakWeb of Science
dc.indekslendigikaynakScopus
dc.indekslendigikaynakPubMed
dc.language.isoen
dc.publisherPublic Library Science
dc.relation.ispartofPlos One
dc.relation.publicationcategoryMakale - Uluslararası Hakemli Dergi - Kurum Öğretim Elemanı
dc.rightsinfo:eu-repo/semantics/openAccess
dc.snmzKA_WoS_20260511
dc.subjectRandom Forests
dc.subjectNatural Enemies
dc.subjectClassification
dc.subjectAlgorithms
dc.titleComparative evaluation of machine learning models for predicting Cimbex quadrimaculata population density across multiple problem formulations
dc.typeArticle

Dosyalar