Deep Learning-Based Visual Question Answering for Medical Imaging: Insights from the PathVQA Dataset
| dc.contributor.author | Balik, Esra | |
| dc.contributor.author | Kaya, Mehmet | |
| dc.date.accessioned | 2026-08-12T16:08:43Z | |
| dc.date.issued | 2024 | |
| dc.department | Fırat Üniversitesi | |
| dc.description | 2024 International Conference on Decision Aid Sciences and Applications, DASA 2024 -- 11 December 2024 through 12 December 2024 -- Manama -- 206116 | |
| dc.description.abstract | Extracting information from medical data and making the correct diagnosis based on this information is critical in medical decision support systems. Interpretation of complex medical images requires a deep medical knowledge. Traditional methods have performance limitations in such data-intensive analyses and are insufficient in terms of obtaining accurate results. In this study, we present an effective solution to obtain fast and meaningful information from medical images and develop a Medical Visual Question Answering (MedVQA) system. With this system, it is aimed to produce correct answers by making sense of the questions directed to the contents of medical images by using deep learning-based models trained on the Path Vqadataset. In the training process of the model, transfer learning techniques and Attention mechanisms are used to better analyze the complex structure of medical terminology. With this proposed solution, it is aimed to accelerate the diagnostic processes of healthcare professionals and to provide automatic understanding of visual information. As a result of the experiments, it was observed that the MedVQA model developed as a result of the training on free-form question-answer pairs on the PathVQA dataset performed 58.63% with BERT+VGGI9 and 60.130/0 with BioBERT+VGGI9. For the question-answer pairs in the Yes-No form, 91.80% and 92.17% accuracy was obtained, respectively. Our study showed a significant performance improvement compared to existing approaches and demonstrated that MedVQA models can be effectively used for deep learning-based extraction of meaningful information from medical images and extraction of features of textual data with natural language processing models. Our study aims to contribute to healthcare professionals to obtain more accurate and faster solutions by using decision support systems with an innovative approach in the field of medical image analysis. © 2024 IEEE. | |
| dc.identifier.doi | 10.1109/DASA63652.2024.10836414 | |
| dc.identifier.isbn | 979-835036910-6 | |
| dc.identifier.scopus | 2-s2.0-85217216078 | |
| dc.identifier.scopusquality | N/A | |
| dc.identifier.uri | https://doi.org/10.1109/DASA63652.2024.10836414 | |
| dc.identifier.uri | https://hdl.handle.net/11508/41384 | |
| dc.indekslendigikaynak | Scopus | |
| dc.language.iso | en | |
| dc.publisher | Institute of Electrical and Electronics Engineers Inc. | |
| dc.relation.ispartof | 2024 International Conference on Decision Aid Sciences and Applications, DASA 2024 | |
| dc.relation.publicationcategory | Konferans Öğesi - Uluslararası - Kurum Öğretim Elemanı | |
| dc.rights | info:eu-repo/semantics/closedAccess | |
| dc.snmz | KA_Scopus_20260511 | |
| dc.subject | BERT; BioBERT; Medical Visual Question Answering; VGG19 | |
| dc.title | Deep Learning-Based Visual Question Answering for Medical Imaging: Insights from the PathVQA Dataset | |
| dc.type | Conference Object |







