Advancing multimodal emotion analysis: a hybrid deep learning approach with intermediate fusion and multi-task learning

dc.contributor.authorKaya, Fatih
dc.contributor.authorKaraca, Yunus Emre
dc.contributor.authorAslan, Serpil
dc.contributor.authorYildirim, Muhammed
dc.date.accessioned2026-08-12T17:28:47Z
dc.date.issued2026
dc.departmentFırat Üniversitesi
dc.description.abstractEmotion analysis is a critical research domain focused on detecting the emotional states of individuals or communities across multiple data modalities, including text, images, and audio. While substantial progress has been made in unimodal (text-based) sentiment analysis, real-world scenarios often involve multimodal data, making integrated approaches essential for capturing contextual richness and improving predictive accuracy. This study introduces a hybrid deep learning model that combines text and visual features through an intermediate fusion mechanism and multi-task learning framework. Textual inputs are processed using RoBERTa and BiGRU layers, while visual inputs are analyzed through ViT and ResNet50 architectures enhanced by the Convolutional Block Attention Module (CBAM). The fused multimodal representations enable simultaneous and more robust emotion classification. Experimental results on the MVSA dataset demonstrate the superior performance of the proposed model, achieving 96.02% accuracy, 95.51% precision, 94.07% recall, and 94.73% F1-score, outperforming several state-of-the-art multimodal benchmarks. These findings underscore the model's methodological contributions and its strong potential for advancing the field of multimodal emotion analysis in both academic research and real-world applications.
dc.description.sponsorshipMalatya Turgut zal University
dc.description.sponsorshipOpen access funding provided by the Scientific and Technological Research Council of Turkiye (TUB & Idot;TAK).
dc.identifier.doi10.1007/s00371-026-04475-1
dc.identifier.issn0178-2789
dc.identifier.issn1432-2315
dc.identifier.issue6
dc.identifier.scopus2-s2.0-105036168269
dc.identifier.scopusqualityQ1
dc.identifier.urihttps://doi.org/10.1007/s00371-026-04475-1
dc.identifier.urihttps://hdl.handle.net/11508/55441
dc.identifier.volume42
dc.identifier.wosWOS:001746927200001
dc.identifier.wosqualityQ2
dc.indekslendigikaynakWeb of Science
dc.indekslendigikaynakScopus
dc.language.isoen
dc.publisherSpringer
dc.relation.ispartofVisual Computer
dc.relation.publicationcategoryMakale - Uluslararası Hakemli Dergi - Kurum Öğretim Elemanı
dc.rightsinfo:eu-repo/semantics/openAccess
dc.snmzKA_WoS_20260511
dc.subjectMultimodal emotion analysis
dc.subjectIntermediate fusion
dc.subjectMulti-task learning
dc.subjectVision transformer
dc.subjectSentiment classification
dc.titleAdvancing multimodal emotion analysis: a hybrid deep learning approach with intermediate fusion and multi-task learning
dc.typeArticle

Dosyalar