Enhancing speech emotion recognition through deep learning and handcrafted feature fusion

dc.contributor.authorEris, Fatma Gunes
dc.contributor.authorAkbal, Erhan
dc.date.accessioned2026-08-12T18:10:39Z
dc.date.issued2024
dc.departmentFırat Üniversitesi
dc.description.abstractIn this paper, we introduce an innovative investigation in speech emotion recognition (SER). The proposed model combines deep learning-based and handcrafted audio features to achieve optimal accuracy. The proposed model employs an iterative feature selection and majority voting pipeline to obtain better results by fusing deep learning-based and handcrafted features. The wav2vec2 model and the openSmile audio processing library are used in order to extract audio features from audio data. Then the feature selection and majority voting techniques are used to identify the optimal feature selection methods for diverse feature sets and to combine their strengths. The experiments are performed using a diverse and extensive corpus to ensure the robustness of the proposed method. In the construction of this multi-corpus dataset, we used four well-known benchmark datasets, namely Ravdess, Savee, Crema-D, and Tess. All of these datasets are publicly available. These datasets are combined on six common emotions: sadness, happiness, fear, anger, surprise, and disgust. The resultant dataset comprises 11,511 samples across these categories. The proposed method has been shown to achieve results comparable to those reported in the existing literature. The experimental results indicate that the proposed pipeline leads to a 3 % improvement in classification accuracy. The highest achieved accuracy on the multi-corpus dataset is 92.55 %.
dc.identifier.doi10.1016/j.apacoust.2024.110070
dc.identifier.issn0003-682X
dc.identifier.issn1872-910X
dc.identifier.orcid0000-0002-5257-7560
dc.identifier.orcid0000-0002-6048-6060
dc.identifier.scopus2-s2.0-85192722463
dc.identifier.scopusqualityQ1
dc.identifier.urihttps://doi.org/10.1016/j.apacoust.2024.110070
dc.identifier.urihttps://hdl.handle.net/11508/63359
dc.identifier.volume222
dc.identifier.wosWOS:001241408800001
dc.identifier.wosqualityQ1
dc.indekslendigikaynakWeb of Science
dc.indekslendigikaynakScopus
dc.language.isoen
dc.publisherElsevier Sci Ltd
dc.relation.ispartofApplied Acoustics
dc.relation.publicationcategoryMakale - Uluslararası Hakemli Dergi - Kurum Öğretim Elemanı
dc.rightsinfo:eu-repo/semantics/closedAccess
dc.snmzKA_WoS_20260511
dc.subjectAcoustic feature extraction
dc.subjectFeature engineering
dc.subjectEmotion recognition
dc.subjectDeep learning
dc.subjectFeature fusion
dc.titleEnhancing speech emotion recognition through deep learning and handcrafted feature fusion
dc.typeArticle

Dosyalar