Long-Tail Aware Cross-Modal Graph Attention Network for Fine-Grained Indoor 3D Semantic Segmentation of Point Clouds

dc.contributor.authorOzbay, Erdal
dc.contributor.authorOzbay, Feyza Altunbey
dc.date.accessioned2026-09-08T07:11:34Z
dc.date.issued2026
dc.departmentFırat Üniveristesi
dc.description.abstractHighlights What are the main findings? The proposed LT-CM-GACNet++ effectively improves semantic segmentation performance on long-tail distributed indoor point cloud data by leveraging cross-modal feature fusion. The integration of CMGA and prototype-based long-tail learning significantly enhances rare-class recognition while maintaining strong accuracy. What are the implications of the main findings? The results demonstrate that cross-modal learning combined with long-tail aware optimization provides a robust solution for fine-grained 3D indoor scene understanding. The proposed framework can be extended to other multimodal 3D vision tasks, offering a scalable approach for handling class imbalance in large-scale real-world datasets.Highlights What are the main findings? The proposed LT-CM-GACNet++ effectively improves semantic segmentation performance on long-tail distributed indoor point cloud data by leveraging cross-modal feature fusion. The integration of CMGA and prototype-based long-tail learning significantly enhances rare-class recognition while maintaining strong accuracy. What are the implications of the main findings? The results demonstrate that cross-modal learning combined with long-tail aware optimization provides a robust solution for fine-grained 3D indoor scene understanding. The proposed framework can be extended to other multimodal 3D vision tasks, offering a scalable approach for handling class imbalance in large-scale real-world datasets.Abstract Accurate and efficient semantic segmentation of point cloud data is critical in many application areas involving indoor scene understanding. In particular, fine-grained object categories, high data density, and class imbalance in high-resolution indoor datasets significantly limit class discrimination in 3D semantic segmentation. The multimodal data structure, high-fidelity geometry, and long-tail class distribution of the recently popular ScanNet++ dataset further exacerbate these challenges. This study proposes a novel Long-Tail Aware Cross-Modal Graph Attention Network (LT-CM-GACNet++) to address fine-grained 3D semantic segmentation under long-tail distributions. The proposed method integrates dynamic graph-based geometric feature extraction with a lightweight visual feature extractor based on MobileNetV3, enabling effective fusion of geometric and RGB-based information. The proposed Cross-Modal Graph Attention (CMGA) module facilitates adaptive information transfer between modalities, enabling more effective representation learning of both local and global contextual features. To mitigate the adverse effects of long-tail class distributions, prototype-based representation learning and a class frequency-aware loss function are jointly employed. This strategy improves the learning of rare classes while enhancing the discrimination between visually and geometrically similar categories. In the preprocessing stage, density-based sampling, normal vector estimation, and block-based fixed-size point cloud generation are applied to high-resolution mesh-derived data. The proposed model is evaluated on 50 scenes and 100 semantic classes selected from the ScanNet++ dataset. Experimental results demonstrate that the proposed method achieves significant improvements over existing approaches in terms of both overall segmentation performance and rare-class performance. In particular, notable gains are observed in mean Intersection over Union (mIoU) and rare-class mIoU metrics. These results highlight the effectiveness of cross-modal learning for high-resolution 3D scene segmentation under long-tail distributions.
dc.identifier.doi10.3390/s26113401
dc.identifier.issn1424-8220
dc.identifier.issue11
dc.identifier.orcid0000-0003-0629-6888
dc.identifier.pmid42280919
dc.identifier.scopus2-s2.0-105041388935
dc.identifier.scopusqualityQ1
dc.identifier.urihttps://doi.org/10.3390/s26113401
dc.identifier.urihttps://hdl.handle.net/11508/65081
dc.identifier.volume26
dc.identifier.wosWOS:001790268300001
dc.identifier.wosqualityQ2
dc.indekslendigikaynakWeb of Science
dc.indekslendigikaynakScopus
dc.indekslendigikaynakPubMed
dc.language.isoen
dc.publisherMdpi
dc.relation.ispartofSensors
dc.relation.publicationcategoryMakale - Uluslararası Hakemli Dergi - Kurum Öğretim Elemanı
dc.rightsinfo:eu-repo/semantics/openAccess
dc.snmzKA_WOS_20250903
dc.subjectCross-Modal Fusion
dc.subjectGraph Attention Networks
dc.subjectLong-Tail Learning
dc.subjectPoint Cloud
dc.subjectSemantic Segmentation
dc.titleLong-Tail Aware Cross-Modal Graph Attention Network for Fine-Grained Indoor 3D Semantic Segmentation of Point Clouds
dc.typeArticle

Dosyalar