RE-LIG: A Faithfulness-Driven Layer Integrated Gradients Framework for Explainable Medical Visual Question Answering

dc.contributor.authorBalik, Esra
dc.contributor.authorAygun, Irfan
dc.contributor.authorKaya, Mehmet
dc.date.accessioned2026-09-08T07:13:57Z
dc.date.issued2026
dc.departmentFırat Üniveristesi
dc.description.abstractMedical Visual Question Answering (Med-VQA) systems have the potential to support medical image interpretation and clinical decision-making processes. However, the black-box nature of existing systems and low-resolution constraints limit the transparency of model decisions, hindering clinical applicability. This work proposes a high-resolution holistic framework called robust and efficient layer-integrated gradients (RE-LIG) to enhance reliability and explainability in Med-VQA systems. The proposed architecture is built upon three key components: (1) high-resolution visual encoding: the PubMedCLIP encoder is scaled to high-resolution using dynamic positional embedding interpolation to capture fine details. (2) Multimodal semantic fusion: clinical questions solved by BioLinkBERT and visual features obtained by PubMedCLIP are aligned through a coattention mechanism. (3) Explainability framework: to counter the noisy nature of classical gradient methods, the RE-LIG algorithm, which combines noise tunneling and layer-based integration strategies, has been integrated into the system. Extensive experiments conducted on the SLAKE dataset demonstrate the proposed framework's success in primarily increasing model faithfulness. Quantitative analyses demonstrate that the RE-LIG method achieves a + 28.9% higher explanation fidelity (RE-LIG AOPC = 0.3180 vs. Vanilla IG = 0.2467, Bootstrap 95% CI [0.262-0.375], Wilcoxon p < 0.001) compared to standard gradient approaches. While achieving this gain in explainability, competitive performance with state-of-the-art (SOTA) models was achieved without compromising diagnostic performance (80.77% overall accuracy, 87.61% closed-ended, and 77.34% open-ended performance). Ablation studies confirm that the integrated noise reduction mechanisms shift the model's focus from background noise to actual pathological boundaries. The findings demonstrate that explainability is not merely a visual aid for clinical confidence but a measurable and verifiable requirement.
dc.description.sponsorshipFimath;rat University -- Trkiye Bilimsel ve Teknolojik Arascedil;timath;rma Kurumu [125E165] -- Open access funding provided by the Scientific and Technological Research Council of Turkiye (TUB & Idot;TAK).
dc.identifier.doi10.1007/s10278-026-02079-8
dc.identifier.issn2948-2925
dc.identifier.issn2948-2933
dc.identifier.pmid42373891
dc.identifier.scopus2-s2.0-105043502158
dc.identifier.scopusqualityN/A
dc.identifier.urihttps://doi.org/10.1007/s10278-026-02079-8
dc.identifier.urihttps://hdl.handle.net/11508/65636
dc.identifier.wosWOS:001806982800001
dc.identifier.wosqualityQ2
dc.indekslendigikaynakWeb of Science
dc.indekslendigikaynakScopus
dc.indekslendigikaynakPubMed
dc.language.isoen
dc.publisherSpringer
dc.relation.ispartofJournal of Imaging Informatics in Medicine
dc.relation.publicationcategoryMakale - Uluslararası Hakemli Dergi - Kurum Öğretim Elemanı
dc.rightsinfo:eu-repo/semantics/openAccess
dc.snmzKA_WOS_20250903
dc.subjectMedical Vqa
dc.subjectExplainable Artificial Intelligence (Xai)
dc.subjectVision Transformers
dc.subjectRe-Lig
dc.subjectMultimodal Learning
dc.titleRE-LIG: A Faithfulness-Driven Layer Integrated Gradients Framework for Explainable Medical Visual Question Answering
dc.typeArticle

Dosyalar