A Comparative Investigation of Cepstral Feature Extraction Methods for Deepfake Speech Detection

dc.contributor.authorAkinci, Nida
dc.contributor.authorOzbay, Erdal
dc.date.accessioned2026-09-08T07:11:54Z
dc.date.issued2026
dc.departmentFırat Üniveristesi
dc.description.abstractThe widespread adoption of voice-based authentication systems has been accompanied by an escalating threat from deep learning-based synthetic speech generation techniques. This study presents a comparative and experimental investigation of cepstral feature extraction methods for deepfake speech detection. Specifically, Mel-Frequency Cepstral Coefficients (MFCC), Linear-Frequency Cepstral Coefficients (LFCC), and Constant-Q Cepstral Coefficients (CQCC) are systematically evaluated with respect to their frequency scaling characteristics, spectral resolution properties, and capacity to capture artifacts specific to synthetic speech production. Experiments were conducted on 5571 audio samples drawn from the ASVspoof 2021 Logical Access evaluation partition, with all methods assessed under identical classification conditions using a linear Support Vector Machine. Results indicate that CQCC attains the highest numerical performance, achieving 83.59% accuracy, 89.15% ROC-AUC, and 15.83% Equal Error Rate (EER); however, the performance difference between MFCC and CQCC does not reach statistical significance (p = 0.202). Five-fold cross-validation corroborates this finding (CQCC: 87.89% +/- 0.81%). McNemar's test confirms that the performance difference between LFCC and CQCC is statistically significant (p = 0.036). A fine-grained attack-wise analysis across 13 spoofing systems reveals that no single feature representation consistently outperforms the others across all attack types; CQCC achieves the highest accuracy on 6 out of 13 systems, while MFCC remains competitive on several attack categories. The overall findings indicate that deepfake detection performance is highly sensitive not only to the classifier architecture but also to the choice of frequency scale, cepstral transformation design, and data conditions. Empirical motivation is provided that multi-feature strategies integrating complementary frequency representations may offer more robust and generalizable detection solutions.
dc.description.sponsorshipThis research received no external funding.
dc.identifier.doi10.3390/app16136707
dc.identifier.issn2076-3417
dc.identifier.issue13
dc.identifier.scopus2-s2.0-105044567281
dc.identifier.scopusqualityQ1
dc.identifier.urihttps://doi.org/10.3390/app16136707
dc.identifier.urihttps://hdl.handle.net/11508/65208
dc.identifier.volume16
dc.identifier.wosWOS:001818121800001
dc.identifier.wosqualityQ2
dc.indekslendigikaynakWeb of Science
dc.indekslendigikaynakScopus
dc.language.isoen
dc.publisherMdpi
dc.relation.ispartofApplied Sciences-Basel
dc.relation.publicationcategoryMakale - Uluslararası Hakemli Dergi - Kurum Öğretim Elemanı
dc.rightsinfo:eu-repo/semantics/openAccess
dc.snmzKA_WOS_20250903
dc.subjectDeepfake Speech Detection
dc.subjectAnti-Spoofing Countermeasure
dc.subjectConstant-Q Transform
dc.subjectCepstral Analysis
dc.titleA Comparative Investigation of Cepstral Feature Extraction Methods for Deepfake Speech Detection
dc.typeArticle

Dosyalar