Comparison of pre-trained language models in terms of carbon emissions, time and accuracy in multi-label text classification using AutoML

dc.contributor.authorSavci, Pinar
dc.contributor.authorDas, Bihter
dc.date.accessioned2026-08-12T16:57:55Z
dc.date.issued2023
dc.departmentFırat Üniversitesi
dc.description.abstractSince Turkish is an agglutinative language and contains reduplication, idiom, and metaphor words, Turkish texts are sources of information with extremely rich meanings. For this reason, the processing and classification of Turkish texts according to their characteristics is both timeconsuming and difficult. In this study, the performances of pre-trained language models for multi-text classification using Autotrain were compared in a 250 K Turkish dataset that we created. The results showed that the BERTurk (uncased, 128 k) language model on the dataset showed higher accuracy performance with a training time of 66 min compared to the other models and the CO2 emission was quite low. The ConvBERTurk mC4 (uncased) model is also the best-performing second language model. As a result of this study, we have provided a deeper understanding of the capabilities of pre-trained language models for Turkish on machine learning.
dc.identifier.doi10.1016/j.heliyon.2023.e15670
dc.identifier.issn2405-8440
dc.identifier.issue5
dc.identifier.orcid0000-0002-2498-3297
dc.identifier.pmid37187909
dc.identifier.scopus2-s2.0-85153857146
dc.identifier.scopusqualityQ1
dc.identifier.urihttps://doi.org/10.1016/j.heliyon.2023.e15670
dc.identifier.urihttps://hdl.handle.net/11508/46648
dc.identifier.volume9
dc.identifier.wosWOS:001029523500001
dc.identifier.wosqualityN/A
dc.indekslendigikaynakWeb of Science
dc.indekslendigikaynakScopus
dc.indekslendigikaynakPubMed
dc.language.isoen
dc.publisherCell Press
dc.relation.ispartofHeliyon
dc.relation.publicationcategoryMakale - Uluslararası Hakemli Dergi - Kurum Öğretim Elemanı
dc.rightsinfo:eu-repo/semantics/openAccess
dc.snmzKA_WoS_20260511
dc.subjectNatural language processing
dc.subjectAutotrain
dc.subjectMulti-text classification
dc.subjectPre-trained language models
dc.subjectArtificial intelligence
dc.subjectAutoNLP
dc.titleComparison of pre-trained language models in terms of carbon emissions, time and accuracy in multi-label text classification using AutoML
dc.typeArticle

Dosyalar