Attention guided 3D CNN-LSTM model for accurate speech based emotion recognition

dc.contributor.authorAtila, Orhan
dc.contributor.authorSengur, Abdulkadir
dc.date.accessioned2026-08-12T18:06:55Z
dc.date.issued2021
dc.departmentFırat Üniversitesi
dc.description.abstractIn this paper, a novel approach, which is based on attention guided 3D convolutional neural networks (CNN)-long short-term memory (LSTM) model, is proposed for speech based emotion recognition. The proposed attention guided 3D CNN-LSTM model is trained in end-to-end fashion. The input speech signals are initially resampled and pre-processed for noise removing and emphasizing the high frequencies. Then, spectrogram, Mel-frequency cepstral coefficient (MFCC), cochleagram and fractal dimension methods are used to convert the input speech signals into the speech images. The obtained images are concatenated into four-dimensional volumes and used as input to the developed 28 layered attention integrated 3D CNN-LSTM model. In the 3D CNN-LSTM model, there are six 3D convolutional layers, two batch normalization (BN) layers, five Rectified Linear Unit (ReLu) layers, three 3D max pooling layers, one attention, one LSTM, one flatten and one dropout layers, and two fully connected layers. The attention layer is connected to the 3D convolution layers. Three datasets namely Ryerson Audio-Visual Database of Emotional Speech (RAVDESS), RML and SAVEE are used in the experimental works. Besides, the mixture of these datasets is also used in the experimental works. Classification accuracy, sensitivity, specificity and F1-score are used for evaluation of the developed method. The obtained results are also compared with some of the recently published results and it is seen that the proposed method outperforms the compared methods. (C) 2021 Elsevier Ltd. All rights reserved.
dc.identifier.doi10.1016/j.apacoust.2021.108260
dc.identifier.issn0003-682X
dc.identifier.issn1872-910X
dc.identifier.orcid0000-0001-7211-913X
dc.identifier.scopus2-s2.0-85109219667
dc.identifier.scopusqualityQ1
dc.identifier.urihttps://doi.org/10.1016/j.apacoust.2021.108260
dc.identifier.urihttps://hdl.handle.net/11508/62505
dc.identifier.volume182
dc.identifier.wosWOS:000687528600045
dc.identifier.wosqualityQ1
dc.indekslendigikaynakWeb of Science
dc.indekslendigikaynakScopus
dc.language.isoen
dc.publisherElsevier Sci Ltd
dc.relation.ispartofApplied Acoustics
dc.relation.publicationcategoryMakale - Uluslararası Hakemli Dergi - Kurum Öğretim Elemanı
dc.rightsinfo:eu-repo/semantics/closedAccess
dc.snmzKA_WoS_20260511
dc.subjectSpeech emotion recognition
dc.subjectAttention
dc.subject3D CNN-LSTM model
dc.titleAttention guided 3D CNN-LSTM model for accurate speech based emotion recognition
dc.typeArticle

Dosyalar