Medical Report Generation from Medical Images Using Vision Transformer and Bart Deep Learning Architectures

dc.contributor.authorUcan, Murat
dc.contributor.authorKaya, Buket
dc.contributor.authorKaya, Mehmet
dc.contributor.authorAlhajj, Reda
dc.date.accessioned2026-08-12T16:58:19Z
dc.date.issued2025
dc.departmentFırat Üniversitesi
dc.description16th International Conference on Social Networks Analysis and Mining -- SEP 02-05, 2024 -- Rende, ITALY
dc.description.abstractGenerating medical reports from medical images using traditional methods is a time-consuming process that is prone to human error and requires experience. Failure to generate fast reports from medical images delays the treatment of patients, and misdiagnosis can lead to adverse conditions that can cause the death of patients. The main objective of this study is to develop a high-performance deep learning model that can autonomously generate medical reports from medical images. The proposed model consists of a Vision Transformer (ViT) encoder and a Bidirectional Autoregressive Transformer (BART) decoder. Training and testing on the model was conducted using images and reports from the Indiana University Chest X-Ray dataset. The developed model is analyzed with measurable parameters and then compared with its competitors in the literature using the same dataset. The proposed Vi-Ba architecture achieved success scores of 0.150, 0.154, 0.274 in bleu-4, meteor and rouge word matching evaluation metrics, respectively. The Vi-Ba model achieved high reporting performance compared to the studies reviewed in the literature. The results show that the proposed architecture can be used by specialized doctors in hospitals to diagnose diseases faster and more accurately. In this way, misdiagnosis and treatments will be reduced and human life will be protected.
dc.description.sponsorshipScientific and Technological Research Council of Turkey (TUBITAK) [123E171]
dc.description.sponsorshipThis research was supported by the Scientific and Technological Research Council of Turkey (TUBITAK) under Grant No 123E171.
dc.identifier.doi10.1007/978-3-031-78554-2_17
dc.identifier.endpage267
dc.identifier.isbn978-3-031-78553-5
dc.identifier.isbn978-3-031-78554-2
dc.identifier.issn0302-9743
dc.identifier.issn1611-3349
dc.identifier.orcid0000-0003-2995-8282
dc.identifier.orcid0000-0001-9219-2262
dc.identifier.orcid0000-0001-9505-181X
dc.identifier.scopus2-s2.0-85218451982
dc.identifier.scopusqualityQ3
dc.identifier.startpage257
dc.identifier.urihttps://doi.org/10.1007/978-3-031-78554-2_17
dc.identifier.urihttps://hdl.handle.net/11508/46812
dc.identifier.volume15214
dc.identifier.wosWOS:001447243400017
dc.identifier.wosqualityN/A
dc.indekslendigikaynakWeb of Science
dc.indekslendigikaynakScopus
dc.language.isoen
dc.publisherSpringer International Publishing Ag
dc.relation.ispartofSocial Networks Analysis and Mining, Asonam 2024, Pt Iv
dc.relation.publicationcategoryKonferans Öğesi - Uluslararası - Kurum Öğretim Elemanı
dc.rightsinfo:eu-repo/semantics/closedAccess
dc.snmzKA_WoS_20260511
dc.subjectDeep Learning
dc.subjectVision Transformer
dc.subjectViT
dc.subjectBidirectional Autoregressive Transformer
dc.subjectBART
dc.subjectMedical Report Generation
dc.subjectChest X-rays
dc.titleMedical Report Generation from Medical Images Using Vision Transformer and Bart Deep Learning Architectures
dc.typeConference Object

Dosyalar