Attention-Based CNN-RNN Arabic Text Recognition from Natural Scene Images

dc.contributor.authorButt, Hanan
dc.contributor.authorRaza, Muhammad Raheel
dc.contributor.authorRamzan, Muhammad Javed
dc.contributor.authorAli, Muhammad Junaid
dc.contributor.authorHaris, Muhammad
dc.date.accessioned2026-08-12T18:07:24Z
dc.date.issued2021
dc.departmentFırat Üniversitesi
dc.description.abstractAccording to statistics, there are 422 million speakers of the Arabic language. Islam is the second-largest religion in the world, and its followers constitute approximately 25% of the world's population. Since the Holy Quran is in Arabic, nearly all Muslims understand the Arabic language per some analytical information. Many countries have Arabic as their native and official language as well. In recent years, the number of internet users speaking the Arabic language has been increased, but there is very little work on it due to some complications. It is challenging to build a robust recognition system (RS) for cursive nature languages such as Arabic. These challenges become more complex if there are variations in text size, fonts, colors, orientation, lighting conditions, noise within a dataset, etc. To deal with them, deep learning models show noticeable results on data modeling and can handle large datasets. Convolutional neural networks (CNNs) and recurrent neural networks (RNNs) can select good features and follow the sequential data learning technique. These two neural networks offer impressive results in many research areas such as text recognition, voice recognition, several tasks of Natural Language Processing (NLP), and others. This paper presents a CNN-RNN model with an attention mechanism for Arabic image text recognition. The model takes an input image and generates feature sequences through a CNN. These sequences are transferred to a bidirectional RNN to obtain feature sequences in order. The bidirectional RNN can miss some preprocessing of text segmentation. Therefore, a bidirectional RNN with an attention mechanism is used to generate output, enabling the model to select relevant information from the feature sequences. An attention mechanism implements end-to-end training through a standard backpropagation algorithm.
dc.identifier.doi10.3390/forecast3030033
dc.identifier.endpage540
dc.identifier.issn2571-9394
dc.identifier.issue3
dc.identifier.orcid0000-0003-0208-9419
dc.identifier.orcid0000-0002-6305-2583
dc.identifier.orcid0000-0002-9523-5800
dc.identifier.orcid0000-0002-2885-3617
dc.identifier.scopus2-s2.0-85123177230
dc.identifier.scopusqualityQ1
dc.identifier.startpage520
dc.identifier.urihttps://doi.org/10.3390/forecast3030033
dc.identifier.urihttps://hdl.handle.net/11508/62682
dc.identifier.volume3
dc.identifier.wosWOS:000700733800001
dc.identifier.wosqualityQ1
dc.indekslendigikaynakWeb of Science
dc.indekslendigikaynakScopus
dc.language.isoen
dc.publisherMdpi
dc.relation.ispartofForecasting
dc.relation.publicationcategoryMakale - Uluslararası Hakemli Dergi - Kurum Öğretim Elemanı
dc.rightsinfo:eu-repo/semantics/openAccess
dc.snmzKA_WoS_20260511
dc.subjectimage text recognition
dc.subjectdeep learning
dc.subjectrecurrent neural networks (RNNs)
dc.subjectconvolutional neural networks (CNNs)
dc.subjectbidirectional RNN
dc.subjectattention mechanism
dc.subjecttext segmentation
dc.subjectnatural scene images
dc.titleAttention-Based CNN-RNN Arabic Text Recognition from Natural Scene Images
dc.typeArticle

Dosyalar