DEVELOPMENT OF AN IOT BASED ENVIRONMENTAL SOUND EVENT RECOGNITION METHODS USING WIRELESS ACOUSTIC SENSOR NETWORKS IN SMART CITIES
| dc.contributor.advisor | YAMAN, ORHAN | |
| dc.contributor.author | ALI, YUSUF YAU | |
| dc.date.accessioned | 2026-08-12T10:09:50Z | |
| dc.date.issued | 2025 | |
| dc.department | FÜ, Fen Bilimleri Enstitüsü, Adli Bilişim Mühendisliği Anabilim Dalı | |
| dc.description.abstract | Bu çalışma, ESP32 cihazıyla kaydedilen özel bir veri seti ve kamuya açık ESC-10 veri setini kullanarak çeşitli yaklaşımları değerlendirmek suretiyle çevresel ses sınıflandırmasını geliştirmeyi amaçlamaktadır. Özel veri seti, "havaalanı, plaj, uçak, otoyol, lobi, konak, ofis ve restoran" olmak üzere sekiz kategori içerirken, ESC-10 veri seti 10 sınıfta toplam 400 kısa kayıttan oluşmaktadır. Birinci yöntem, transfer öğrenmesini kullanarak, önceden eğitilmiş DarkNet53 model ile ses kayıtlarının Mel-spektrogram ve gri tonlamalı görüntülere dönüştürülmesi yoluyla özellik çıkarma potansiyelini ortaya koymuştur. ESC-10 veri setinde %97.24 doğruluk oranı elde eden bu model, önceden eğitilmiş konvolüsyonel sinir ağlarının (CNN) nispeten küçük veri setlerini işleme, hesaplama karmaşıklığını azaltma ve rekabetçi sonuçlar elde etme konusundaki adaptasyon yeteneğini ve verimliliğini göstermiştir. İkinci yöntem, özel veri seti üzerinde Gelişmiş Sinir Ağı Mimarileri olan Uzun Kısa Süreli Bellek (LSTM), Çift Yönlü Uzun Kısa Süreli Bellek (Bi-LSTM) ve CNN modellerini değerlendirmiştir. Bi-LSTM üstün bir performans sergileyerek bağlamsal öğrenme ve zamansal bağımlılıkları yakalama konusundaki güçlü yeteneğini ortaya koymuştur. Bu yöntemle Bi-LSTM, %99.4'lük en yüksek eğitim doğruluğu ve %97'lik test doğruluğuna ulaşmıştır. Bu sonuçlar, ses verilerinin özel özelliklerine uygun esnek model mimarilerinin önemini vurgulamaktadır. Son yöntem, Mel Frekans Kepstral Katsayıları (MFCC) ve Destek Vektör Makineleri (SVM) uygulamasına odaklanmış, özellik çıkarma ve segmentasyon süresinin etkisini incelemiştir. Bu yöntem, 10 saniyelik segmentasyon ile istatistiksel MFCC özellikleri kullanılarak %97.39'luk maksimum test doğruluğuna ulaşmıştır. Bulgular, sınıflandırma görevlerini yönetmede ve doğruluğu artırmada özellik çıkarma ve uzatılmış segmentasyonun önemini ortaya koymaktadır. Bu araştırma, farklı metodolojilerin entegrasyonu yoluyla pratik uygulamalar için kesin ve güvenilir çevresel ses sınıflandırma sistemlerinin oluşturulmasını büyük ölçüde geliştirmektedir. | |
| dc.description.abstract | This study aims to enhance environmental sound classification by evaluating several approaches with a custom dataset recoded with Esp32 device and the publicly accessible ESC-10 dataset. The custom dataset has eight categories: airport, beach, aircraft, highway, lobby, lodge, office, and restaurant, whereas the ESC-10 dataset consists of 400 brief recordings across 10 classes. The first method, used transfer learning with the pre-trained DarkNet53 model, the potential of converting sound recordings into Mel-spectrogram and gray-scale images for feature extraction. Achieving a 97.24% accuracy on the ESC-10 dataset, this model outshined the adaptability and efficiency of pre-trained convolutional neural networks (CNNs) in handling relatively small datasets, reducing computational complexity, and achieving competitive results. The second method showed the advanced neural network architectures, considering Long Short-Term Memory (LSTM) Bidirectional Long Short-Term Memory (Bi-LSTM) and CNN models, on the custom dataset. Bi-LSTM achieved superior performance, showing its super capability for contextual learning and capturing temporal dependencies. These results highlight the importance of flexible model architectures to the specific characteristics of the sound data, with Bi-LSTM achieving the highest training accuracy of 99.4% and test accuracy of 97%. The final method emphasized the application of Mel-Frequency Cepstral Coefficients (MFCC) in conjunction with Support Vector Machines (SVM), concentrating on the impact of feature extraction and segmentation duration. This method achieved a peak test accuracy of 97.39% by employing statistical MFCC features with a 10-second segmentation. The findings underscore the significance of feature extraction and extended segmentation in managing classification jobs and enhancing accuracy. This research greatly enhances the creation of precise and dependable environmental sound classification systems for practical applications through the integration of different methodologies. | |
| dc.identifier.citation | YAU ALI, Y. (2025). Development of an IoT based environmental sound event recognition methods using wireless acoustic sensor networks in smart cities (Tez No. 919321) [Yüksek lisans tezi, Fırat Üniversitesi]. | |
| dc.identifier.uri | https://tez.yok.gov.tr/UlusalTezMerkezi/TezGoster?key=E_eEUHQic_C-LvhxNQn1W4U-OyHrh5glE3zhK1KcG_7IJCceoNcMR5x2A7PZ5W3s | |
| dc.identifier.uri | https://hdl.handle.net/11508/22528 | |
| dc.identifier.yoktezid | 919321 | |
| dc.language.iso | en | |
| dc.publisher | Fırat Üniveristesi | |
| dc.relation.publicationcategory | Tez | |
| dc.rights | info:eu-repo/semantics/openAccess | |
| dc.snmz | KA_TEZ_20260511 | |
| dc.subject | Bilim ve Teknoloji | |
| dc.title | DEVELOPMENT OF AN IOT BASED ENVIRONMENTAL SOUND EVENT RECOGNITION METHODS USING WIRELESS ACOUSTIC SENSOR NETWORKS IN SMART CITIES | |
| dc.title.alternative | Akıllı şehirlerde kablosuz akustik sensör ağları kullanarak çevresel ses olayı tanıma yöntemlerinin IoT tabanlı geliştirilmesi | |
| dc.type | Master Thesis |







