Tweet Verileri İçin Metin Sınıflandırmasında Gelişmiş Makine Öğrenmesi Modelleri: CatBoost ve LightGBM ile Performans Karşılaştırması

dc.contributor.authorEşidir, Kamil Abdullah
dc.date.accessioned2026-08-12T15:02:35Z
dc.date.issued2025
dc.departmentFırat Üniversitesi
dc.description.abstractÇalışmada, sosyal medya tabanlı metinlerden oluşan ikili sınıflandırma problemi kapsamında duygu analizi gerçekleştirilmiştir. Analiz sürecinde, metinler dil ön işleme adımlarından geçirilmiş ve cümle düzeyinde çok dilli BERT modeli kullanılarak vektörleştirilmiştir. Dengesiz sınıf dağılımı problemi ise SMOTE (Synthetic Minority Over-sampling Technique) yöntemi ile dengelenmiştir. Sınıflandırmada LightGBM ve CatBoost makine öğrenmesi modelleri tercih edilmiştir. Modellere beş katlı çapraz doğrulama uygulanarak, doğruluk, F1 skoru, duyarlılık, özgüllük ve ROC-AUC gibi çeşitli performans metrikleri hesaplanmıştır. Görsel analizlerde metin uzunluğu, kelime sayısı ve kelime bulutu benzeri yapısal dağılımlar incelenmiştir. Elde edilen sonuçlara göre her iki model de yüksek sınıflandırma başarısı göstermiştir. CatBoost doğruluk (%87,4), F1 skoru (0,763), hassasiyet (0,737) ve duyarlılık (0,793) ölçütlerinde LightGBM’ye kıyasla tutarlı bir üstünlük sağlamıştır. Pozitif sınıfı daha başarılı tanıması ve dengeli genel performansı ile öne çıkmıştır. İki modelin ROC-AUC değeri ise eşit (0,926) bulunmuş ve sınıflar arası ayrım gücünün yüksek olduğu anlaşılmıştır. Elde edilen sonuçlar, gelişmiş vektörleştirme tekniklerinin makine öğrenmesi modelleri ile bütünleştiğinde duygu analizinde etkili çıktılar üretebildiğini ortaya koymaktadır.
dc.description.abstractIn this study, sentiment analysis was performed for a binary classification problem consisting of social media-based texts. In the analysis process, the texts were subjected to language preprocessing steps and vectorized using the multilingual BERT model at the sentence level. The unbalanced class distribution problem was balanced with the SMOTE (Synthetic Minority Over-sampling Technique) method. LightGBM and CatBoost machine learning models were preferred for classification. Five-fold cross-validation was applied to the models and various performance metrics such as accuracy, F1 score, sensitivity, specificity and ROC-AUC were calculated. In visual analysis, text length, word count and word cloud-like structural distributions were analyzed. According to the results, both models showed high classification performance. CatBoost consistently outperformed LightGBM in accuracy (87.4%), F1 score (0.763), precision (0.737) and sensitivity (0.793). It stood out with its better recognition of the positive class and balanced overall performance. The ROC-AUC value of the two models was equal (0.926), indicating high discrimination power between classes. The results show that advanced vectorization techniques can produce effective outputs in sentiment analysis when integrated with machine learning models.
dc.identifier.doi10.54525/bbmd.1677261
dc.identifier.endpage123
dc.identifier.issn3023-7459
dc.identifier.issue2
dc.identifier.startpage112
dc.identifier.urihttps://doi.org/10.54525/bbmd.1677261
dc.identifier.urihttps://hdl.handle.net/11508/26510
dc.identifier.volume18
dc.language.isotr
dc.publisherAkademik Bilişim Vakfı
dc.relation.ispartofBilgisayar Bilimleri ve Mühendisliği Dergisi
dc.relation.publicationcategoryMakale - Ulusal Hakemli Dergi - Kurum Öğretim Elemanı
dc.rightsinfo:eu-repo/semantics/openAccess
dc.snmzKA_DergiPark_20260511
dc.subjectBusiness Process Management
dc.subjectİş Süreçleri Yönetimi
dc.subjectDecision Support and Group Support Systems
dc.subjectKarar Desteği ve Grup Destek Sistemleri
dc.subjectManagement Information Systems
dc.subjectYönetim Bilişim Sistemleri
dc.titleTweet Verileri İçin Metin Sınıflandırmasında Gelişmiş Makine Öğrenmesi Modelleri: CatBoost ve LightGBM ile Performans Karşılaştırması
dc.title.alternativeAdvanced Machine Learning Models for Text Classification of Tweet Data: Performance Comparison with CatBoost and LightGBM
dc.typeArticle

Dosyalar