Named Entity Recognition for Historical Texts in Turkish

dc.contributor.authorŞeker, Miraytu?
dc.contributor.authorDemir, Muhammed Abdullah
dc.contributor.authorArzu, Mehmet
dc.contributor.authorKaya, Mahmut
dc.date.accessioned2026-08-12T16:08:02Z
dc.date.issued2025
dc.departmentFırat Üniversitesi
dc.description10th International Conference on Computer Science and Engineering, UBMK 2025 -- 17 September 2025 through 21 September 2025 -- Istanbul -- 214243
dc.description.abstractNamed Entity Recognition (NER) is used in the subfield of Natural Language Processing (NLP) to assign entity names in unstructured text to categories determined according to their relevance. This study aims to create a balanced and high-quality dataset for NER tasks on Turkish historical texts and to compare and evaluate classical machine learning approaches and different deep learning models. A dataset of 12,455 words was prepared from Turkish historical documents found on Wikipedia and various other sources. Six main entity types (PER, LOC, ORG, DATE, TITLE, EVENT) were considered in the dataset; all words were manually labeled in a context-sensitive manner according to the BILOU (Beginning, Inside, Last, Outside, Unit) labeling scheme. During the data preprocessing stage, sentence and word segmentation, removal of unnecessary characters, and appropriate labeling of words that can have multiple classes depending on the context were performed. In the experimental section, comprehensive comparisons were made between three different Transformer-based BERTurk-based transformer models (BERTurk (cased, 32k), BERTurk (cased, 128k), BERTurk (uncased, 128k)) and the traditional CRF model. The models were evaluated using basic metrics such as accuracy, precision, sensitivity, and F1 score. The results showed that the highest performance was achieved with the BERTurk (cased, 128k) model, with 94.10% accuracy and 93.88% F1 score. Transformer-based models generally outperformed the traditional CRF algorithm by a significant margin. © 2025 IEEE.
dc.identifier.doi10.1109/UBMK67458.2025.11206819
dc.identifier.endpage627
dc.identifier.issn2521-1641
dc.identifier.issue2025
dc.identifier.scopus2-s2.0-105030815363
dc.identifier.scopusqualityN/A
dc.identifier.startpage622
dc.identifier.urihttps://doi.org/10.1109/UBMK67458.2025.11206819
dc.identifier.urihttps://hdl.handle.net/11508/41012
dc.indekslendigikaynakScopus
dc.language.isoen
dc.publisherInstitute of Electrical and Electronics Engineers Inc.
dc.relation.ispartofInternational Conference on Computer Science and Engineering, UBMK
dc.relation.publicationcategoryKonferans Öğesi - Uluslararası - Kurum Öğretim Elemanı
dc.rightsinfo:eu-repo/semantics/closedAccess
dc.snmzKA_Scopus_20260511
dc.subjectBERTurk; Named Entity Recognition; Natural Language Processing; Transformer
dc.titleNamed Entity Recognition for Historical Texts in Turkish
dc.typeConference Object

Dosyalar