Developing Question-Answering Models in Low-Resource Languages: A Case Study on Turkish Medical Texts Using Transformer-Based Approaches

dc.contributor.authorIncidelen, Mert
dc.contributor.authorAydogan, Murat
dc.date.accessioned2026-08-12T16:09:10Z
dc.date.issued2024
dc.departmentFırat Üniversitesi
dc.description8th International Artificial Intelligence and Data Processing Symposium, IDAP 2024 -- 21 September 2024 through 22 September 2024 -- Malatya -- 203423
dc.description.abstractIn this study, transformer-based pre-trained language models were fine-tuned using medical texts for question-answering (QA) tasks in Turkish, a low-resource language. Variations of the BERTurk pre-trained language model created using large Turkish corpus were used for QA tasks. The study presents a medical Turkish QA dataset created using Turkish Wikipedia and medical theses located in the Thesis Center of the Council of Higher Education in Turkey. This dataset, containing a total of 8200 question-answer pairs, is used to fine-tune the BERTurk model. The performance of the models was evaluated by Exact Match (EM) and F1 score. The BERTurk (cased, 32k) model achieved an EM of 51.097 and an F1 score of 74.148, while the BERTurk (cased, 128 k) model achieved an EM of 55.121 and an F1 score of 77.187. The results show that pre-trained language models can be successfully used for question-answer tasks in low-resource languages such as Turkish. This study lays an important foundation for Turkish medical text processing and automatic QA tasks and sheds light on future research in this field. © 2024 IEEE.
dc.identifier.doi10.1109/IDAP64064.2024.10711128
dc.identifier.isbn979-833153149-2
dc.identifier.scopus2-s2.0-85207860765
dc.identifier.scopusqualityN/A
dc.identifier.urihttps://doi.org/10.1109/IDAP64064.2024.10711128
dc.identifier.urihttps://hdl.handle.net/11508/41625
dc.indekslendigikaynakScopus
dc.language.isoen
dc.publisherInstitute of Electrical and Electronics Engineers Inc.
dc.relation.ispartof8th International Artificial Intelligence and Data Processing Symposium, IDAP 2024
dc.relation.publicationcategoryKonferans Öğesi - Uluslararası - Kurum Öğretim Elemanı
dc.rightsinfo:eu-repo/semantics/closedAccess
dc.snmzKA_Scopus_20260511
dc.subjectBERTurk; Medical Domain; Natural Language Processing; Question-Answering
dc.titleDeveloping Question-Answering Models in Low-Resource Languages: A Case Study on Turkish Medical Texts Using Transformer-Based Approaches
dc.typeConference Object

Dosyalar