Long-Tail Aware Cross-Modal Graph Attention Network for Fine-Grained Indoor 3D Semantic Segmentation of Point Clouds
| dc.contributor.author | Ozbay, Erdal | |
| dc.contributor.author | Ozbay, Feyza Altunbey | |
| dc.date.accessioned | 2026-09-08T07:11:34Z | |
| dc.date.issued | 2026 | |
| dc.department | Fırat Üniveristesi | |
| dc.description.abstract | Highlights What are the main findings? The proposed LT-CM-GACNet++ effectively improves semantic segmentation performance on long-tail distributed indoor point cloud data by leveraging cross-modal feature fusion. The integration of CMGA and prototype-based long-tail learning significantly enhances rare-class recognition while maintaining strong accuracy. What are the implications of the main findings? The results demonstrate that cross-modal learning combined with long-tail aware optimization provides a robust solution for fine-grained 3D indoor scene understanding. The proposed framework can be extended to other multimodal 3D vision tasks, offering a scalable approach for handling class imbalance in large-scale real-world datasets.Highlights What are the main findings? The proposed LT-CM-GACNet++ effectively improves semantic segmentation performance on long-tail distributed indoor point cloud data by leveraging cross-modal feature fusion. The integration of CMGA and prototype-based long-tail learning significantly enhances rare-class recognition while maintaining strong accuracy. What are the implications of the main findings? The results demonstrate that cross-modal learning combined with long-tail aware optimization provides a robust solution for fine-grained 3D indoor scene understanding. The proposed framework can be extended to other multimodal 3D vision tasks, offering a scalable approach for handling class imbalance in large-scale real-world datasets.Abstract Accurate and efficient semantic segmentation of point cloud data is critical in many application areas involving indoor scene understanding. In particular, fine-grained object categories, high data density, and class imbalance in high-resolution indoor datasets significantly limit class discrimination in 3D semantic segmentation. The multimodal data structure, high-fidelity geometry, and long-tail class distribution of the recently popular ScanNet++ dataset further exacerbate these challenges. This study proposes a novel Long-Tail Aware Cross-Modal Graph Attention Network (LT-CM-GACNet++) to address fine-grained 3D semantic segmentation under long-tail distributions. The proposed method integrates dynamic graph-based geometric feature extraction with a lightweight visual feature extractor based on MobileNetV3, enabling effective fusion of geometric and RGB-based information. The proposed Cross-Modal Graph Attention (CMGA) module facilitates adaptive information transfer between modalities, enabling more effective representation learning of both local and global contextual features. To mitigate the adverse effects of long-tail class distributions, prototype-based representation learning and a class frequency-aware loss function are jointly employed. This strategy improves the learning of rare classes while enhancing the discrimination between visually and geometrically similar categories. In the preprocessing stage, density-based sampling, normal vector estimation, and block-based fixed-size point cloud generation are applied to high-resolution mesh-derived data. The proposed model is evaluated on 50 scenes and 100 semantic classes selected from the ScanNet++ dataset. Experimental results demonstrate that the proposed method achieves significant improvements over existing approaches in terms of both overall segmentation performance and rare-class performance. In particular, notable gains are observed in mean Intersection over Union (mIoU) and rare-class mIoU metrics. These results highlight the effectiveness of cross-modal learning for high-resolution 3D scene segmentation under long-tail distributions. | |
| dc.identifier.doi | 10.3390/s26113401 | |
| dc.identifier.issn | 1424-8220 | |
| dc.identifier.issue | 11 | |
| dc.identifier.orcid | 0000-0003-0629-6888 | |
| dc.identifier.pmid | 42280919 | |
| dc.identifier.scopus | 2-s2.0-105041388935 | |
| dc.identifier.scopusquality | Q1 | |
| dc.identifier.uri | https://doi.org/10.3390/s26113401 | |
| dc.identifier.uri | https://hdl.handle.net/11508/65081 | |
| dc.identifier.volume | 26 | |
| dc.identifier.wos | WOS:001790268300001 | |
| dc.identifier.wosquality | Q2 | |
| dc.indekslendigikaynak | Web of Science | |
| dc.indekslendigikaynak | Scopus | |
| dc.indekslendigikaynak | PubMed | |
| dc.language.iso | en | |
| dc.publisher | Mdpi | |
| dc.relation.ispartof | Sensors | |
| dc.relation.publicationcategory | Makale - Uluslararası Hakemli Dergi - Kurum Öğretim Elemanı | |
| dc.rights | info:eu-repo/semantics/openAccess | |
| dc.snmz | KA_WOS_20250903 | |
| dc.subject | Cross-Modal Fusion | |
| dc.subject | Graph Attention Networks | |
| dc.subject | Long-Tail Learning | |
| dc.subject | Point Cloud | |
| dc.subject | Semantic Segmentation | |
| dc.title | Long-Tail Aware Cross-Modal Graph Attention Network for Fine-Grained Indoor 3D Semantic Segmentation of Point Clouds | |
| dc.type | Article |







