DiagPat: An Explainable Language Detection Model Using EEG Signals

dc.contributor.authorKeles, Tugce
dc.contributor.authorYildirim, Kubra
dc.contributor.authorTanko, Dahiru
dc.contributor.authorTas, Suat
dc.contributor.authorTasci, Irem
dc.contributor.authorTasci, Burak
dc.contributor.authorDogan, Sengul
dc.date.accessioned2026-09-08T07:11:34Z
dc.date.issued2026
dc.departmentFırat Üniveristesi
dc.description.abstractElectroencephalography (EEG) offers a non-invasive and cost-effective means of probing brain activity during language processing; however, prior EEG-based language studies have been limited by small datasets, a predominant focus on native-speaker or speech-unit recognition rather than direct language detection, evaluation on only a small number of experimental settings, and frequent reliance on computationally intensive deep learning models with limited interpretability. The proposed feature engineering models classifies EEG segments by language and task mode. The languages are Arabic and Turkish. The modes are reading and listening. In this study, a signal refers to one fixed-length multi-channel EEG segment (14 channels & times; 15 s at 128 Hz). A channel refers to one electrode time series within that segment. To address these gaps, we curated a new EEG language detection dataset from 346 participants (98 Arabic and 248 Turkish) recorded in reading and listening modes, yielding 6364 EEG segments. Using this dataset, we proposed DiagPat, an explainable feature engineering (XFE) model that extracts transition table-based features from both EEG channels and signals through diagonal pattern analysis. The model combines DiagPat feature extraction with iterative neighborhood component analysis (INCA) for feature selection, at algorithm-based k-nearest neighbors (tkNN) classifier for prediction, and the Directed Lobish (DLob) symbolic language for explainability. We evaluated the framework across nine classification cases covering language detection, mode detection, and mixed multi-class settings. The proposed DiagPat-driven XFE model achieved more than 90% accuracy in all cases, with accuracies ranging from 92.14% to 99.35%, and generated case-specific cortical connectome diagrams that supported the interpretable characterization of language- and mode-related brain activity. Subject-independent results were also reported using leave-one-subject-out cross-validation (LOSO CV), where LOSO accuracies ranged from 29.75% to 83.50%. Thus, the 10-fold CV results show segment-level performance, whereas the LOSO results show subject-level generalization. Balanced accuracy and macro-F1 are also reported. These findings indicate that DiagPat provides an accurate, lightweight, and explainable framework for EEG-based language detection.
dc.description.sponsorshipScientific and Technological Research Council of Turkey (TUBITAK) [123E129] -- This work is supported by the 123E129 project fund provided by the Scientific and Technological Research Council of Turkey (TUBITAK).
dc.identifier.doi10.3390/s26113363
dc.identifier.issn1424-8220
dc.identifier.issue11
dc.identifier.pmid42280883
dc.identifier.scopus2-s2.0-105041555067
dc.identifier.scopusqualityQ1
dc.identifier.urihttps://doi.org/10.3390/s26113363
dc.identifier.urihttps://hdl.handle.net/11508/65082
dc.identifier.volume26
dc.identifier.wosWOS:001790239300001
dc.identifier.wosqualityQ2
dc.indekslendigikaynakWeb of Science
dc.indekslendigikaynakScopus
dc.indekslendigikaynakPubMed
dc.language.isoen
dc.publisherMdpi
dc.relation.ispartofSensors
dc.relation.publicationcategoryMakale - Uluslararası Hakemli Dergi - Kurum Öğretim Elemanı
dc.rightsinfo:eu-repo/semantics/openAccess
dc.snmzKA_WOS_20250903
dc.subjectDiagpat
dc.subjectEeg Language Detection
dc.subjectCognitive Science
dc.subjectNeuroscience
dc.subjectExplainable Feature Engineering
dc.titleDiagPat: An Explainable Language Detection Model Using EEG Signals
dc.typeArticle

Dosyalar