Derin öğrenme yöntemleriyle otomatik olarak üretilmiş sahte metin özetlerinin tespiti
| dc.contributor.advisor | Özbay, Erdal | |
| dc.contributor.author | Öner, İsmail | |
| dc.date.accessioned | 2026-09-08T07:02:15Z | |
| dc.date.issued | 2026 | |
| dc.department | FÜ, Fen Bilimleri Enstitüsü, Bilgisayar Mühendisliği Anabilim Dalı | |
| dc.description.abstract | Bu tez çalışmasında, akademik metinler bağlamında yapay zeka üretimi metinlerin tespit edilmesine yönelik XLM-RoBERTa tabanlı bir derin öğrenme yaklaşımı önerilmiştir. Büyük dil modellerinin insan yazımına oldukça yakın, tutarlı ve biçimsel açıdan düzenli metinler üretebilmesi, özellikle akademik yazımda insan ve yapay zeka üretimi içeriklerin ayırt edilmesini zorlaştırmaktadır. Bu çalışma, söz konusu probleme karşı farklı üretici model mimarilerine daha dayanıklı ve yüzeysel veri artefaktlarına bağımlılığı azaltılmış bir tespit sistemi geliştirmeyi amaçlamaktadır. Çalışma kapsamında insan yazımı akademik metinler ile altı farklı büyük dil modeli tarafından üretilmiş sentetik metinlerden oluşan dengeli bir veri yapısı oluşturulmuştur. Veri seti A ve veri seti B'nin birleştirilmesiyle toplam 63.000 örnekten oluşan Büyük Karışım veri seti elde edilmiştir. Model, bu veri seti üzerinde ince ayar sürecinden geçirilmiş ve eğitimden bağımsız olarak hazırlanan 1.200 örneklik Sıfır-Kalıntı test seti üzerinde değerlendirilmiştir. Test setinde yüzeysel biçimsel kalıntıların etkisi azaltılarak modelin gerçek ayırt edici kapasitesi analiz edilmiştir. Deneysel sonuçlar, önerilen modelin Sıfır-Kalıntı test seti üzerinde yüksek ve kararlı performans sergilediğini göstermiştir. Model, ham metin koşulunda %93.41 doğruluk, %92.35 kesinlik, %94.66 duyarlılık ve %93.41 F1-skoru elde etmiştir. Ayrıca 0.9810 AUC değeri, modelin insan yazımı ve yapay zeka üretimi akademik metinleri güçlü biçimde ayırt edebildiğini göstermektedir. | |
| dc.description.abstract | This thesis proposes an XLM-RoBERTa-based deep learning approach for detecting AI-generated texts in the academic writing domain. The increasing ability of large language models to generate fluent, coherent, and formally structured texts has made it difficult to distinguish human-written academic texts from AI-generated content. This study aims to develop a detection system that is more robust against different generator model architectures and less dependent on superficial dataset artifacts. In this study, a balanced data structure was created using human-written academic texts and synthetic texts generated by six different large language models. By combining Dataset A and Dataset B, a Large Mixture dataset containing 63,000 samples was obtained. The proposed model was fine-tuned on this dataset and evaluated on an independent Zero-Artifact test set consisting of 1,200 samples. In the test set, the effects of superficial formatting artifacts were reduced in order to analyze the model's actual discriminative capacity. The experimental results show that the proposed model achieved high and stable performance on the Zero-Artifact test set. Under the raw text condition, the model obtained 93.41% accuracy, 92.35% precision, 94.66% recall, and 93.41% F1-score. In addition, the AUC value of 0.9810 indicates that the model can strongly distinguish between human-written and AI-generated academic texts. Overall, the findings suggest that multi-LLM training data and artifact-controlled evaluation improve the reliability of AI-generated text detection in academic contexts. | |
| dc.identifier.citation | ÖNER, İ. (2026). Derin öğrenme yöntemleriyle otomatik olarak üretilmiş sahte metin özetlerinin tespiti (Tez No. 1013891) [Yüksek lisans tezi, FIRAT ÜNİVERSİTESİ]. | |
| dc.identifier.uri | https://tez.yok.gov.tr/UlusalTezMerkezi/TezGoster?key=5T1_CZ5-UGb9QCmoURec4FGG4KtKl3PqN8IWmeYl_AP1a99czCBJzIbzv_oDerPu | |
| dc.identifier.uri | https://hdl.handle.net/11508/64570 | |
| dc.identifier.yoktezid | 1013891 | |
| dc.institutionauthor | Öner, İsmail | |
| dc.language.iso | tr | |
| dc.publisher | Fırat Üniveristesi | |
| dc.relation.publicationcategory | Tez | |
| dc.rights | info:eu-repo/semantics/openAccess | |
| dc.snmz | KA_TEZ_20250903 | |
| dc.subject | Bilgisayar Mühendisliği Bilimleri-Bilgisayar ve Kontrol | |
| dc.title | Derin öğrenme yöntemleriyle otomatik olarak üretilmiş sahte metin özetlerinin tespiti | |
| dc.title.alternative | Detection of automatically generated fake text summaries using deep learning methods | |
| dc.type | Master Thesis |







