QUBVIS: query based multi-modal summarization system using CLIP based transformer and vision language models

dc.contributor.authorAltundogan, Turan Goktug
dc.contributor.authorKarakose, Mehmet
dc.date.accessioned2026-08-12T17:27:06Z
dc.date.issued2025
dc.departmentFırat Üniversitesi
dc.description.abstractIn this study, a new approach is proposed for user-interactive summarization of online videos. In the proposed approach, video-to-video summarization is performed with a very high success rate using a multimodal transformer architecture (QUBVIS) that also takes activity queries from the user as input, and the resulting summary video is subjected to captioning using a Vision Language Model with a GPT-2 decoder. The developed models are integrated with a Flask API and presented in a way that online video platforms can easily integrate into their systems. In addition, a simple web interface using this API is developed to provide API communication with the user. The performance evaluations of both models of the proposed method show our superiority over similar studies in the literature.
dc.description.sponsorshipTUBITAK (The Scientific and Technological Research Council of Turkey) [5220154]
dc.description.sponsorshipThis study was supported by the TUBITAK (The Scientific and Technological Research Council of Turkey) under Grant No: 5220154. This study was produced from Phd Thesis belong to Turan Goktug Altundogan.
dc.identifier.doi10.1016/j.softx.2025.102303
dc.identifier.issn2352-7110
dc.identifier.orcid0000-0002-3276-3788
dc.identifier.orcid0000-0002-8677-3105
dc.identifier.scopus2-s2.0-105012950251
dc.identifier.scopusqualityQ2
dc.identifier.urihttps://doi.org/10.1016/j.softx.2025.102303
dc.identifier.urihttps://hdl.handle.net/11508/55068
dc.identifier.volume31
dc.identifier.wosWOS:001548766400001
dc.identifier.wosqualityQ2
dc.indekslendigikaynakWeb of Science
dc.indekslendigikaynakScopus
dc.language.isoen
dc.publisherElsevier
dc.relation.ispartofSoftwarex
dc.relation.publicationcategoryMakale - Uluslararası Hakemli Dergi - Kurum Öğretim Elemanı
dc.rightsinfo:eu-repo/semantics/openAccess
dc.snmzKA_WoS_20260511
dc.subjectVideo summarization
dc.subjectQuery based summarization
dc.subjectVision language models
dc.subjectTransformers
dc.titleQUBVIS: query based multi-modal summarization system using CLIP based transformer and vision language models
dc.typeArticle

Dosyalar