A new pyramidal concatenated CNN approach for environmental sound classification

dc.contributor.authorDemir, Fatih
dc.contributor.authorTurkoglu, Muammer
dc.contributor.authorAslan, Muzaffer
dc.contributor.authorSengur, Abdulkadir
dc.date.accessioned2026-08-12T17:50:28Z
dc.date.issued2020
dc.departmentFırat Üniversitesi
dc.description.abstractRecently, there has been an incremental interest on Environmental Sound Classification (ESC), which is an important topic of the non-speech audio classification task. A novel approach, which is based on deep Convolutional Neural Networks (CNN), is proposed in this study. The proposed approach covers a bunch of stages such as pre-processing, deep learning based feature extraction, feature concatenation, feature reduction and classification, respectively. In the first stage, the input sound signals are denoised and are converted into sound images by using the Sort Time Fourier Transform (STFT) method. After sound images are formed, pre-trained CNN models are used for deep feature extraction. In this stage, VGG16, VGG19 and DenseNet201 models are considered. The feature extraction is performed in a pyramidal fashion which makes the dimension of the feature vector quite large. For both dimension reduction and the determination of the most efficient features, a feature selection mechanism is considered after feature concatenation stage. In the last stage of the proposed method, a Support Vector Machines (SVM) classifier is used. The efficiency of the proposed method is calculated on various ESC datasets such as ESC 10, ESC 50 and UrbanSound8K, respectively. The experimental works show that the proposed method produced 94.8%, 81.4% and 78.14% accuracy scores for ESC-10, ESC-50 and UrbanSound8K datasets. The obtained results are also compared with the state-of-the art methods achievements. (C) 2020 Elsevier Ltd. All rights reserved.
dc.identifier.doi10.1016/j.apacoust.2020.107520
dc.identifier.issn0003-682X
dc.identifier.issn1872-910X
dc.identifier.orcid0000-0003-3210-3664
dc.identifier.orcid0000-0003-1614-2639
dc.identifier.orcid0000-0002-2418-9472
dc.identifier.scopus2-s2.0-85088033071
dc.identifier.scopusqualityQ1
dc.identifier.urihttps://doi.org/10.1016/j.apacoust.2020.107520
dc.identifier.urihttps://hdl.handle.net/11508/62232
dc.identifier.volume170
dc.identifier.wosWOS:000565374000032
dc.identifier.wosqualityQ1
dc.indekslendigikaynakWeb of Science
dc.indekslendigikaynakScopus
dc.language.isoen
dc.publisherElsevier Sci Ltd
dc.relation.ispartofApplied Acoustics
dc.relation.publicationcategoryMakale - Uluslararası Hakemli Dergi - Kurum Öğretim Elemanı
dc.rightsinfo:eu-repo/semantics/closedAccess
dc.snmzKA_WoS_20260511
dc.subjectSound classification
dc.subjectDeep learning
dc.subjectSVM
dc.subjectSTFT
dc.subjectCNN
dc.titleA new pyramidal concatenated CNN approach for environmental sound classification
dc.typeArticle

Dosyalar