Vocal Fold Disorders Classification and Optimization of a Custom Video Laryngoscopy Dataset Through Structural Similarity Index and a Deep Learning-Based Approach

dc.contributor.authorEmre, Elif
dc.contributor.authorCetintas, Dilber
dc.contributor.authorYildirim, Muhammed
dc.contributor.authorEmre, Sadettin
dc.date.accessioned2026-08-12T17:42:34Z
dc.date.issued2025
dc.departmentFırat Üniversitesi
dc.description.abstractBackground/Objectives: Video laryngoscopy is one of the primary methods used by otolaryngologists for detecting and classifying laryngeal lesions. However, the diagnostic process of these images largely relies on clinicians' visual inspection, which can lead to overlooked small structural changes, delayed diagnosis, and interpretation errors. Methods: AI-based approaches are becoming increasingly critical for accelerating early-stage diagnosis and improving reliability. This study proposes a hybrid Convolutional Neural Network (CNN) architecture that eliminates repetitive and clinically insignificant frames from videos, utilizing only meaningful key frames. Video data from healthy individuals, patients with vocal fold nodules, and those with vocal fold polyps were summarized using three different threshold values with the Structural Similarity Index Measure (SSIM). Results: The resulting key frames were then classified using a hybrid CNN. Experimental findings demonstrate that selecting an appropriate threshold can significantly reduce the model's memory usage and processing load while maintaining accuracy. In particular, a threshold value of 0.90 provided richer information content thanks to the selection of a wider variety of frames, resulting in the highest success rate. Fine-tuning the last 20 layers of the MobileNetV2 and Xception backbones, combined with the fusion of extracted features, yielded an overall classification accuracy of 98%. Conclusions: The proposed approach provides a mechanism that eliminates unnecessary data and prioritizes only critical information in video-based diagnostic processes, thus helping physicians accelerate diagnostic decisions and reduce memory requirements.
dc.identifier.doi10.3390/jcm14196899
dc.identifier.issn2077-0383
dc.identifier.issue19
dc.identifier.orcid0000-0003-0710-2280
dc.identifier.orcid0000-0001-5659-3499
dc.identifier.orcid0000-0003-1866-4721
dc.identifier.pmid41095977
dc.identifier.scopus2-s2.0-105018849024
dc.identifier.scopusqualityQ1
dc.identifier.urihttps://doi.org/10.3390/jcm14196899
dc.identifier.urihttps://hdl.handle.net/11508/59787
dc.identifier.volume14
dc.identifier.wosWOS:001593667300001
dc.identifier.wosqualityQ1
dc.indekslendigikaynakWeb of Science
dc.indekslendigikaynakScopus
dc.indekslendigikaynakPubMed
dc.language.isoen
dc.publisherMdpi
dc.relation.ispartofJournal of Clinical Medicine
dc.relation.publicationcategoryMakale - Uluslararası Hakemli Dergi - Kurum Öğretim Elemanı
dc.rightsinfo:eu-repo/semantics/openAccess
dc.snmzKA_WoS_20260511
dc.subjectartificial intelligence
dc.subjectdysphonia
dc.subjectendoscopic video
dc.subjectlaryngoscopy
dc.subjectvocal cord polyps
dc.titleVocal Fold Disorders Classification and Optimization of a Custom Video Laryngoscopy Dataset Through Structural Similarity Index and a Deep Learning-Based Approach
dc.typeArticle

Dosyalar