Skeleton-Based Activity Recognition for Children with Autism Using Graph Convolutional Networks

dc.contributor.authorAy, Betul
dc.contributor.authorOzturk, Mehmet Ata
dc.contributor.authorAydin, Galip
dc.date.accessioned2026-09-08T07:11:34Z
dc.date.issued2026
dc.departmentFırat Üniveristesi
dc.description.abstractMovement-based and physical activity programs are central tools in autism intervention, so recognizing the activities a child performs during therapy is valuable for objective progress tracking. Manual monitoring of these sessions is time-consuming and subjective, and raw videos raise privacy concerns because it shows identifiable children. We address autism therapeutic activity recognition from privacy-preserving 2D skeletons, and we focus on the practical difficulty of how several therapeutic activities differ only in subtle motion details. As a backbone, we adopt ProtoGCN, a graph convolutional network that represents each action as a combination of learnable motion prototypes. However, this contrastive backbone organizes all classes at once, so it does not enforce a margin between the few pairs that remain entangled after training. We therefore introduce a Refine-Confusable (RC) module, a training-only regularizer that pushes apart the empirically most-confused class pairs using a hinge-margin loss over momentum-updated class centroids. The module changes neither the backbone nor the inference cost. On the MMASD dataset, restricted to the ten-class 2D-skeleton configuration, the RC module improves the base model across random, session-independent, and subject-independent evaluation. The gain is largest on the strictest subject-independent split and a clip-level analysis confirms that this improvement is statistically significant. Under the protocol-matched holdout, the method reaches 96.30% accuracy with 0.959 macro-F1, surpassing recent 2D-skeleton baselines while keeping a lightweight and privacy-preserving modality. The improvements are modest, as expected on a small clinical dataset, and t-SNE and prototype visualizations show that the learned representation is discriminative and interpretable.
dc.description.sponsorshipThis research received no external funding.
dc.identifier.doi10.3390/s26144638
dc.identifier.issn1424-8220
dc.identifier.issue14
dc.identifier.pmid42515520
dc.identifier.scopus2-s2.0-105045932600
dc.identifier.scopusqualityQ1
dc.identifier.urihttps://doi.org/10.3390/s26144638
dc.identifier.urihttps://hdl.handle.net/11508/65079
dc.identifier.volume26
dc.identifier.wosWOS:001833475500001
dc.identifier.wosqualityQ2
dc.indekslendigikaynakWeb of Science
dc.indekslendigikaynakScopus
dc.indekslendigikaynakPubMed
dc.language.isoen
dc.publisherMdpi
dc.relation.ispartofSensors
dc.relation.publicationcategoryMakale - Uluslararası Hakemli Dergi - Kurum Öğretim Elemanı
dc.rightsinfo:eu-repo/semantics/openAccess
dc.snmzKA_WOS_20250903
dc.subjectAutism Spectrum Disorder
dc.subjectHuman Activity Recognition
dc.subject2D Pose Estimation
dc.subjectGraph Convolutional Network
dc.subjectSkeleton-Based Action Recognition
dc.subjectDeep Learning
dc.subjectPhysical Activity
dc.subjectPrivacy-Preserving Monitoring
dc.subjectVideo-Based Motion Analysis
dc.titleSkeleton-Based Activity Recognition for Children with Autism Using Graph Convolutional Networks
dc.typeArticle

Dosyalar