A Speech Emotion Recognition Model Based on Multi-Level Local Binary and Local Ternary Patterns

dc.contributor.authorSonmez, Yesim Ulgen
dc.contributor.authorVarol, Asaf
dc.date.accessioned2026-08-12T17:35:55Z
dc.date.issued2020
dc.departmentFırat Üniversitesi
dc.description.abstractInterpreting a speech signal is quite challenging because it consists of different frequencies and features that vary according to emotions. Although different algorithms are being developed in the speech emotion recognition (SER) domain, the success rates vary according to the spoken languages, emotions, and databases. In this study, a new lightweight effective SER method has been developed that has low computational complexity. This method, called 1BTPDN, is applied on RAVDESS, EMO-DB, SAVEE, and EMOVO databases. First, low-pass filter coefficients are obtained by applying a one-dimensional discrete wavelet transform on the raw audio data. The features are extracted by applying textural analysis methods, a one-dimensional local binary pattern, and a one-dimensional local ternary pattern to each filter. Using neighborhood component analysis, the most dominant 1024 features are selected from 7680 features while the other features are discarded. These 1024 features are selected as the input of the classifier which is a third-degree polynomial kernel-based support vector machine. The success rates of the 1BTPDN reached 95.16% 89.16%, 76.67%, and 74.31%; in the RAVDESS, EMO-DB, SAVEE, and EMOVO databases, respectively. The recognition rates are higher compared to many textural, acoustic, and deep learning state-of-the-art SER methods.
dc.identifier.doi10.1109/ACCESS.2020.3031763
dc.identifier.endpage190796
dc.identifier.issn2169-3536
dc.identifier.orcid0000-0003-1606-4079
dc.identifier.orcid0000-0002-2090-0263
dc.identifier.scopus2-s2.0-85102863932
dc.identifier.scopusqualityQ1
dc.identifier.startpage190784
dc.identifier.urihttps://doi.org/10.1109/ACCESS.2020.3031763
dc.identifier.urihttps://hdl.handle.net/11508/57733
dc.identifier.volume8
dc.identifier.wosWOS:000584839800001
dc.identifier.wosqualityQ2
dc.indekslendigikaynakWeb of Science
dc.indekslendigikaynakScopus
dc.language.isoen
dc.publisherIeee-Inst Electrical Electronics Engineers Inc
dc.relation.ispartofIeee Access
dc.relation.publicationcategoryMakale - Uluslararası Hakemli Dergi - Kurum Öğretim Elemanı
dc.rightsinfo:eu-repo/semantics/openAccess
dc.snmzKA_WoS_20260511
dc.subjectFeature extraction
dc.subjectTime-frequency analysis
dc.subjectClassification algorithms
dc.subjectDatabases
dc.subjectTransforms
dc.subjectSupport vector machines
dc.subjectDiscrete wavelet transform
dc.subjectlocal binary pattern
dc.subjectlocal ternary pattern
dc.subjectneighborhood component analysis
dc.subjectspeech emotion recognition
dc.titleA Speech Emotion Recognition Model Based on Multi-Level Local Binary and Local Ternary Patterns
dc.typeArticle

Dosyalar