In-depth investigation of speech emotion recognition studies from past to present -The importance of emotion recognition from speech signal for AI-

dc.contributor.authorSonmez, Yesimim uLGEN
dc.contributor.authorVarol, Asaf
dc.date.accessioned2026-08-12T17:38:47Z
dc.date.issued2024
dc.departmentFırat Üniversitesi
dc.description.abstractIn the super smart society (Society 5.0), new and rapid methods are needed for speech recognition, emotion recognition, and speech emotion recognition areas to maximize human-machine or human-computer interaction and collaboration. Speech signal contains much information about the speaker, such as age, sex, ethnicity, health condition, emotion, and thoughts. The field of study which analyzes the mood of the person from the speech is called speech emotion recognition (SER). Classifying the emotions from the speech data is a complicated problem for artificial intelligence, and its sub-discipline, machine learning. Because it is hard to analyze the speech signal which contains various frequencies and characteristics. Speech data are digitized with signal processing methods and speech features are obtained. These features vary depending on the emotions such as sadness, fear, anger, happiness, boredom, confusion, etc. Even though different methods have been developed for determining the audio properties and emotion recognition, the success rate varies depending on the languages, cultures, emotions, and data sets. In speech emotion recognition, there is a need for new methods which can be applied in data sets with different sizes, which will increase classification success, in which best properties can be obtained, and which are affordable. The success rates are affected by many factors such as the methods used, lack of speech emotion datasets, the homogeneity of the database, the difficulty of the language (linguistic differences), the noise in audio data and the length of the audio data. Within the scope of this study, studies on emotion recognition from speech signals from past to present have been analyzed in detail. In this study, classification studies based on a discrete emotion model using speech data belonging to the Berlin emotional database (EMODB), Italian emotional speech database (EMOVO), The Surrey audio-visual expressed emotion database (SAVEE), Ryerson Audio-Visual Database of Emotional Speech and Song Database (RAVDESS), which are mostly independent of the speaker and content, are examined. The results of both classical classifiers and deep learning methods are compared. Deep learning results are more successful, but classical classification is more important in determining the defining features of speech, song or voice. So It develops feature extraction stage. This study will be able to contribute to the literature and help the researchers in the SER field.
dc.description.sponsorshipThe authors declare the following financial interests/personal re-lationships which may be considered as potential competing interests: Another publication related to the subject of this study was encour-aged by TUBITAK since it was published in IEEE Access 2021. If there are other authors, they declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper.
dc.identifier.doi10.1016/j.iswa.2024.200351
dc.identifier.issn2667-3053
dc.identifier.orcid0000-0002-2090-0263
dc.identifier.orcid0000-0003-1606-4079
dc.identifier.scopus2-s2.0-85187954118
dc.identifier.scopusqualityQ1
dc.identifier.urihttps://doi.org/10.1016/j.iswa.2024.200351
dc.identifier.urihttps://hdl.handle.net/11508/58566
dc.identifier.volume22
dc.identifier.wosWOS:001306336500001
dc.identifier.wosqualityQ2
dc.indekslendigikaynakWeb of Science
dc.indekslendigikaynakScopus
dc.language.isoen
dc.publisherElsevier
dc.relation.ispartofIntelligent Systems with Applications
dc.relation.publicationcategoryMakale - Uluslararası Hakemli Dergi - Kurum Öğretim Elemanı
dc.rightsinfo:eu-repo/semantics/openAccess
dc.snmzKA_WoS_20260511
dc.subjectArtificial intelligence
dc.subjectMachine-learning methods
dc.subjectSpeech emotion recognition
dc.titleIn-depth investigation of speech emotion recognition studies from past to present -The importance of emotion recognition from speech signal for AI-
dc.typeArticle

Dosyalar