Named Entity Recognition for Historical Texts in Turkish
| dc.contributor.author | Şeker, Miraytu? | |
| dc.contributor.author | Demir, Muhammed Abdullah | |
| dc.contributor.author | Arzu, Mehmet | |
| dc.contributor.author | Kaya, Mahmut | |
| dc.date.accessioned | 2026-08-12T16:08:02Z | |
| dc.date.issued | 2025 | |
| dc.department | Fırat Üniversitesi | |
| dc.description | 10th International Conference on Computer Science and Engineering, UBMK 2025 -- 17 September 2025 through 21 September 2025 -- Istanbul -- 214243 | |
| dc.description.abstract | Named Entity Recognition (NER) is used in the subfield of Natural Language Processing (NLP) to assign entity names in unstructured text to categories determined according to their relevance. This study aims to create a balanced and high-quality dataset for NER tasks on Turkish historical texts and to compare and evaluate classical machine learning approaches and different deep learning models. A dataset of 12,455 words was prepared from Turkish historical documents found on Wikipedia and various other sources. Six main entity types (PER, LOC, ORG, DATE, TITLE, EVENT) were considered in the dataset; all words were manually labeled in a context-sensitive manner according to the BILOU (Beginning, Inside, Last, Outside, Unit) labeling scheme. During the data preprocessing stage, sentence and word segmentation, removal of unnecessary characters, and appropriate labeling of words that can have multiple classes depending on the context were performed. In the experimental section, comprehensive comparisons were made between three different Transformer-based BERTurk-based transformer models (BERTurk (cased, 32k), BERTurk (cased, 128k), BERTurk (uncased, 128k)) and the traditional CRF model. The models were evaluated using basic metrics such as accuracy, precision, sensitivity, and F1 score. The results showed that the highest performance was achieved with the BERTurk (cased, 128k) model, with 94.10% accuracy and 93.88% F1 score. Transformer-based models generally outperformed the traditional CRF algorithm by a significant margin. © 2025 IEEE. | |
| dc.identifier.doi | 10.1109/UBMK67458.2025.11206819 | |
| dc.identifier.endpage | 627 | |
| dc.identifier.issn | 2521-1641 | |
| dc.identifier.issue | 2025 | |
| dc.identifier.scopus | 2-s2.0-105030815363 | |
| dc.identifier.scopusquality | N/A | |
| dc.identifier.startpage | 622 | |
| dc.identifier.uri | https://doi.org/10.1109/UBMK67458.2025.11206819 | |
| dc.identifier.uri | https://hdl.handle.net/11508/41012 | |
| dc.indekslendigikaynak | Scopus | |
| dc.language.iso | en | |
| dc.publisher | Institute of Electrical and Electronics Engineers Inc. | |
| dc.relation.ispartof | International Conference on Computer Science and Engineering, UBMK | |
| dc.relation.publicationcategory | Konferans Öğesi - Uluslararası - Kurum Öğretim Elemanı | |
| dc.rights | info:eu-repo/semantics/closedAccess | |
| dc.snmz | KA_Scopus_20260511 | |
| dc.subject | BERTurk; Named Entity Recognition; Natural Language Processing; Transformer | |
| dc.title | Named Entity Recognition for Historical Texts in Turkish | |
| dc.type | Conference Object |







