An integrated explainable framework for multimodal deepfake detection across image, audio, and video data

dc.contributor.authorArmagan, Senanur
dc.contributor.authorGundogan, Esra
dc.contributor.authorKaya, Mehmet
dc.contributor.authorAlhajj, Reda
dc.date.accessioned2026-09-08T07:13:42Z
dc.date.issued2026
dc.departmentFırat Üniveristesi
dc.description.abstractThe rapid advancement of deepfake generation techniques has introduced significant challenges for digital forensics, particularly due to the increasing realism and multimodal nature of manipulated content. Existing detection approaches are predominantly unimodal and lack interpretability, limiting their effectiveness in real-world scenarios. To address these limitations, this study proposes a unified and explainable deepfake detection framework that operates across image, video, audio, and multimodal data. The proposed framework integrates modality-specific deep learning architectures, InceptionV3 for images, DenseNet169-Bidirectional Long Short-Term Memory (BiLSTM) for video, and Convolutional Neural Network (CNN)-BiLSTM with Extreme Gradient Boosting (XGBoost) for audio, along with a cross-attention-based fusion model for multimodal analysis. A key contribution of this work is a modality-adaptive explainable artificial intelligence (XAI) strategy that systematically combines Shapley Additive Explanations (SHAP), Local Interpretable Model-Agnostic Explanations (LIME), Layer-wise Relevance Propagation (LRP), and Gradient-weighted Class Activation Mapping (Grad-CAM) to generate consistent visual and textual explanations across different data types. Extensive experiments conducted on benchmark datasets, including 140 K-Faces, Celeb-DF-V2, In-The-Wild, and LAV-DF, demonstrate the effectiveness of the proposed approach, achieving accuracy rates of 99% for images, 96% for videos, 98% for audio, and 99% for multimodal data. Furthermore, the framework provides interpretable insights into model decisions through an interactive interface, enhancing transparency and usability in practical applications. The results indicate that the proposed system not only improves detection performance but also addresses the critical need for unified and interpretable deepfake analysis in multimodal environments.
dc.description.sponsorshipFirat University Scientific Research Projects Unit (FUBAP) [MF.24.119] -- This research was supported by the Firat University Scientific Research Projects Unit (FUBAP) under Grant MF.24.119.
dc.identifier.doi10.1016/j.asej.2026.104233
dc.identifier.issn2090-4479
dc.identifier.issn2090-4495
dc.identifier.issue7
dc.identifier.scopus2-s2.0-105038838478
dc.identifier.scopusqualityQ1
dc.identifier.urihttps://doi.org/10.1016/j.asej.2026.104233
dc.identifier.urihttps://hdl.handle.net/11508/65540
dc.identifier.volume17
dc.identifier.wosWOS:001766032100001
dc.identifier.wosqualityQ1
dc.indekslendigikaynakWeb of Science
dc.indekslendigikaynakScopus
dc.language.isoen
dc.publisherElsevier
dc.relation.ispartofAin Shams Engineering Journal
dc.relation.publicationcategoryMakale - Uluslararası Hakemli Dergi - Kurum Öğretim Elemanı
dc.rightsinfo:eu-repo/semantics/openAccess
dc.snmzKA_WOS_20250903
dc.subjectMultimodal Deepfake Detection
dc.subjectXai
dc.subjectShap
dc.subjectLime
dc.subjectLrp
dc.subjectGrad-Cam
dc.subjectBilstm
dc.subjectDensenet169
dc.subjectInceptionv3
dc.subjectHybrid Model
dc.subjectCross Attention
dc.titleAn integrated explainable framework for multimodal deepfake detection across image, audio, and video data
dc.typeArticle

Dosyalar