Deep Learning-Based Visual Question Answering for Medical Imaging: Insights from the PathVQA Dataset

dc.contributor.authorBalik, Esra
dc.contributor.authorKaya, Mehmet
dc.date.accessioned2026-08-12T16:08:43Z
dc.date.issued2024
dc.departmentFırat Üniversitesi
dc.description2024 International Conference on Decision Aid Sciences and Applications, DASA 2024 -- 11 December 2024 through 12 December 2024 -- Manama -- 206116
dc.description.abstractExtracting information from medical data and making the correct diagnosis based on this information is critical in medical decision support systems. Interpretation of complex medical images requires a deep medical knowledge. Traditional methods have performance limitations in such data-intensive analyses and are insufficient in terms of obtaining accurate results. In this study, we present an effective solution to obtain fast and meaningful information from medical images and develop a Medical Visual Question Answering (MedVQA) system. With this system, it is aimed to produce correct answers by making sense of the questions directed to the contents of medical images by using deep learning-based models trained on the Path Vqadataset. In the training process of the model, transfer learning techniques and Attention mechanisms are used to better analyze the complex structure of medical terminology. With this proposed solution, it is aimed to accelerate the diagnostic processes of healthcare professionals and to provide automatic understanding of visual information. As a result of the experiments, it was observed that the MedVQA model developed as a result of the training on free-form question-answer pairs on the PathVQA dataset performed 58.63% with BERT+VGGI9 and 60.130/0 with BioBERT+VGGI9. For the question-answer pairs in the Yes-No form, 91.80% and 92.17% accuracy was obtained, respectively. Our study showed a significant performance improvement compared to existing approaches and demonstrated that MedVQA models can be effectively used for deep learning-based extraction of meaningful information from medical images and extraction of features of textual data with natural language processing models. Our study aims to contribute to healthcare professionals to obtain more accurate and faster solutions by using decision support systems with an innovative approach in the field of medical image analysis. © 2024 IEEE.
dc.identifier.doi10.1109/DASA63652.2024.10836414
dc.identifier.isbn979-835036910-6
dc.identifier.scopus2-s2.0-85217216078
dc.identifier.scopusqualityN/A
dc.identifier.urihttps://doi.org/10.1109/DASA63652.2024.10836414
dc.identifier.urihttps://hdl.handle.net/11508/41384
dc.indekslendigikaynakScopus
dc.language.isoen
dc.publisherInstitute of Electrical and Electronics Engineers Inc.
dc.relation.ispartof2024 International Conference on Decision Aid Sciences and Applications, DASA 2024
dc.relation.publicationcategoryKonferans Öğesi - Uluslararası - Kurum Öğretim Elemanı
dc.rightsinfo:eu-repo/semantics/closedAccess
dc.snmzKA_Scopus_20260511
dc.subjectBERT; BioBERT; Medical Visual Question Answering; VGG19
dc.titleDeep Learning-Based Visual Question Answering for Medical Imaging: Insights from the PathVQA Dataset
dc.typeConference Object

Dosyalar