Query-Driven Retinal Layer Segmentation in OCT Using Cross-Attentive Feature Learning

dc.contributor.authorSobahi, Nebras
dc.contributor.authorOzcelik, Salih Taha Alperen
dc.contributor.authorAtila, Orhan
dc.contributor.authorSengur, Abdulkadir
dc.contributor.authorAkpinar, Muhammed Halil
dc.date.accessioned2026-09-08T07:11:47Z
dc.date.issued2026
dc.departmentFırat Üniveristesi
dc.description.abstractBackground/Objectives: Retinal layer segmentation in optical coherence tomography (OCT) is essential for the diagnosis and monitoring of retinal diseases such as age-related macular degeneration (AMD) and diabetic macular edema (DME). Although deep learning methods have achieved strong performance, most rely on dense pixel-wise predictions and often struggle to preserve anatomical consistency, particularly in regions with low contrast or structural deformation. This study aims to address these limitations by introducing a query-based segmentation framework that explicitly models retinal layer structure. Methods: In this paper, we propose the RetiQueryNet architecture that employs encoding of retinal layers in the form of query embeddings with the use of cross attention to interact with pixel level features encoded by a transformer based encoder. The architecture integrates multi-scale features through a compact query-driven decoder with modest additional computational overhead. Normalization and resizing of OCT images preceded their usage as inputs, while the layer labels were converted to multi-class segmentation maps. In the training process, we used loss function with combination of cross entropy loss and Dice loss. Our model performance was compared with multiple state-of-the-art models such as U-Net, DeepLabV3, FPN, MANet and SegFormer, while performance metrics were Dice, IoU and mean surface distance (MSD). Results: RetiQueryNet was able to attain a mean Dice score of 0.934 +/- 0.0046 and outperformed all baseline models on the main performance measures. Improvements were particularly evident in challenging retinal layers such as IBRPE and OBRPE, where boundary ambiguity is high. It should be noted that RetiQueryNet had a relatively lower MSD value, meaning that the predicted boundaries were more accurate. Furthermore, visual observations suggest that the approach generated smooth and coherent segmentations. Conclusions: The findings demonstrate that query-based modeling offers a viable approach to pixel-wise segmentation. In particular, by making use of structural priors in the form of learnable queries, RetiQueryNet improves not only segmentation accuracy but also anatomical consistency. Query-based modeling appears to be an exciting area for retinal image segmentation that could potentially be applied to other applications in medical image segmentation.
dc.description.sponsorshipFirat University, Scientific Research Project Committee [TEKF.24.47] -- This research was funded by Firat University, Scientific Research Project Committee, under grant No. TEKF.24.47.
dc.identifier.doi10.3390/diagnostics16111697
dc.identifier.issn2075-4418
dc.identifier.issue11
dc.identifier.pmid42279565
dc.identifier.scopus2-s2.0-105041395761
dc.identifier.scopusqualityQ2
dc.identifier.urihttps://doi.org/10.3390/diagnostics16111697
dc.identifier.urihttps://hdl.handle.net/11508/65159
dc.identifier.volume16
dc.identifier.wosWOS:001790342800001
dc.identifier.wosqualityQ1
dc.indekslendigikaynakWeb of Science
dc.indekslendigikaynakScopus
dc.indekslendigikaynakPubMed
dc.language.isoen
dc.publisherMdpi
dc.relation.ispartofDiagnostics
dc.relation.publicationcategoryMakale - Uluslararası Hakemli Dergi - Kurum Öğretim Elemanı
dc.rightsinfo:eu-repo/semantics/openAccess
dc.snmzKA_WOS_20250903
dc.subjectOct
dc.subjectRetinal Layer Segmentation
dc.subjectTransformer
dc.subjectQuery-Based Learning
dc.subjectCross-Attention
dc.subjectMedical Image Segmentation
dc.titleQuery-Driven Retinal Layer Segmentation in OCT Using Cross-Attentive Feature Learning
dc.typeArticle

Dosyalar