Vocal Fold Disorders Classification and Optimization of a Custom Video Laryngoscopy Dataset Through Structural Similarity Index and a Deep Learning-Based Approach
| dc.contributor.author | Emre, Elif | |
| dc.contributor.author | Cetintas, Dilber | |
| dc.contributor.author | Yildirim, Muhammed | |
| dc.contributor.author | Emre, Sadettin | |
| dc.date.accessioned | 2026-08-12T17:42:34Z | |
| dc.date.issued | 2025 | |
| dc.department | Fırat Üniversitesi | |
| dc.description.abstract | Background/Objectives: Video laryngoscopy is one of the primary methods used by otolaryngologists for detecting and classifying laryngeal lesions. However, the diagnostic process of these images largely relies on clinicians' visual inspection, which can lead to overlooked small structural changes, delayed diagnosis, and interpretation errors. Methods: AI-based approaches are becoming increasingly critical for accelerating early-stage diagnosis and improving reliability. This study proposes a hybrid Convolutional Neural Network (CNN) architecture that eliminates repetitive and clinically insignificant frames from videos, utilizing only meaningful key frames. Video data from healthy individuals, patients with vocal fold nodules, and those with vocal fold polyps were summarized using three different threshold values with the Structural Similarity Index Measure (SSIM). Results: The resulting key frames were then classified using a hybrid CNN. Experimental findings demonstrate that selecting an appropriate threshold can significantly reduce the model's memory usage and processing load while maintaining accuracy. In particular, a threshold value of 0.90 provided richer information content thanks to the selection of a wider variety of frames, resulting in the highest success rate. Fine-tuning the last 20 layers of the MobileNetV2 and Xception backbones, combined with the fusion of extracted features, yielded an overall classification accuracy of 98%. Conclusions: The proposed approach provides a mechanism that eliminates unnecessary data and prioritizes only critical information in video-based diagnostic processes, thus helping physicians accelerate diagnostic decisions and reduce memory requirements. | |
| dc.identifier.doi | 10.3390/jcm14196899 | |
| dc.identifier.issn | 2077-0383 | |
| dc.identifier.issue | 19 | |
| dc.identifier.orcid | 0000-0003-0710-2280 | |
| dc.identifier.orcid | 0000-0001-5659-3499 | |
| dc.identifier.orcid | 0000-0003-1866-4721 | |
| dc.identifier.pmid | 41095977 | |
| dc.identifier.scopus | 2-s2.0-105018849024 | |
| dc.identifier.scopusquality | Q1 | |
| dc.identifier.uri | https://doi.org/10.3390/jcm14196899 | |
| dc.identifier.uri | https://hdl.handle.net/11508/59787 | |
| dc.identifier.volume | 14 | |
| dc.identifier.wos | WOS:001593667300001 | |
| dc.identifier.wosquality | Q1 | |
| dc.indekslendigikaynak | Web of Science | |
| dc.indekslendigikaynak | Scopus | |
| dc.indekslendigikaynak | PubMed | |
| dc.language.iso | en | |
| dc.publisher | Mdpi | |
| dc.relation.ispartof | Journal of Clinical Medicine | |
| dc.relation.publicationcategory | Makale - Uluslararası Hakemli Dergi - Kurum Öğretim Elemanı | |
| dc.rights | info:eu-repo/semantics/openAccess | |
| dc.snmz | KA_WoS_20260511 | |
| dc.subject | artificial intelligence | |
| dc.subject | dysphonia | |
| dc.subject | endoscopic video | |
| dc.subject | laryngoscopy | |
| dc.subject | vocal cord polyps | |
| dc.title | Vocal Fold Disorders Classification and Optimization of a Custom Video Laryngoscopy Dataset Through Structural Similarity Index and a Deep Learning-Based Approach | |
| dc.type | Article |







