From Pixels to Paragraphs: Exploring Enhanced Image-to-Text Generation using Inception v3 and Attention Mechanisms

dc.contributor.authorKaraca, Zeynep
dc.contributor.authorDas, Bihter
dc.date.accessioned2026-08-12T15:33:11Z
dc.date.issued2023
dc.departmentFırat Üniversitesi
dc.description.abstractProcessing visual data and converting it into text plays a crucial role in fields like information retrieval and data analysis in the digital world. At this juncture, the \"image-to-text\" transformation, which bridges the gap between visual and textual data, has garnered significant interest from researchers and industry experts. This article presents a study on generating text from images. The study aims to measure the contribution of adding an attention mechanism to the encoder-decoder-based Inception v3 deep learning architecture for image-to-text generation. In the model, the Inception v3 model is trained on the Flickr8k dataset to extract image features. The encoder-decoder structure with an attention mechanism is employed for next-word prediction, and the model is trained on the train images of the Flickr8k dataset for performance evaluation. Experimental results demonstrate the model's satisfactory ability to accurately perceive objects in images.
dc.identifier.doi10.24012/dumf.1340656
dc.identifier.endpage610
dc.identifier.issn1309-8640
dc.identifier.issn2146-4391
dc.identifier.issue4
dc.identifier.startpage603
dc.identifier.trdizinid1276050
dc.identifier.urihttps://doi.org/10.24012/dumf.1340656
dc.identifier.urihttps://search.trdizin.gov.tr/tr/yayin/detay/1276050
dc.identifier.urihttps://hdl.handle.net/11508/33740
dc.identifier.volume14
dc.indekslendigikaynakTR-Dizin
dc.language.isoen
dc.relation.ispartofDicle Üniversitesi Mühendislik Fakültesi Mühendislik Dergisi
dc.relation.publicationcategoryMakale - Ulusal Hakemli Dergi - Kurum Öğretim Elemanı
dc.relation.tubitakinfo:eu-repo/grantAgreement/TUBITAK//
dc.rightsinfo:eu-repo/semantics/openAccess
dc.snmzKA_TR-Dizin_20260511
dc.subjectAttention Mechanisms
dc.subjectInception v3 Model
dc.subjectTextual Content Extraction
dc.subjectImage-to-Text Generation
dc.titleFrom Pixels to Paragraphs: Exploring Enhanced Image-to-Text Generation using Inception v3 and Attention Mechanisms
dc.typeArticle

Dosyalar