QUBVIS: query based multi-modal summarization system using CLIP based transformer and vision language models
| dc.contributor.author | Altundogan, Turan Goktug | |
| dc.contributor.author | Karakose, Mehmet | |
| dc.date.accessioned | 2026-08-12T17:27:06Z | |
| dc.date.issued | 2025 | |
| dc.department | Fırat Üniversitesi | |
| dc.description.abstract | In this study, a new approach is proposed for user-interactive summarization of online videos. In the proposed approach, video-to-video summarization is performed with a very high success rate using a multimodal transformer architecture (QUBVIS) that also takes activity queries from the user as input, and the resulting summary video is subjected to captioning using a Vision Language Model with a GPT-2 decoder. The developed models are integrated with a Flask API and presented in a way that online video platforms can easily integrate into their systems. In addition, a simple web interface using this API is developed to provide API communication with the user. The performance evaluations of both models of the proposed method show our superiority over similar studies in the literature. | |
| dc.description.sponsorship | TUBITAK (The Scientific and Technological Research Council of Turkey) [5220154] | |
| dc.description.sponsorship | This study was supported by the TUBITAK (The Scientific and Technological Research Council of Turkey) under Grant No: 5220154. This study was produced from Phd Thesis belong to Turan Goktug Altundogan. | |
| dc.identifier.doi | 10.1016/j.softx.2025.102303 | |
| dc.identifier.issn | 2352-7110 | |
| dc.identifier.orcid | 0000-0002-3276-3788 | |
| dc.identifier.orcid | 0000-0002-8677-3105 | |
| dc.identifier.scopus | 2-s2.0-105012950251 | |
| dc.identifier.scopusquality | Q2 | |
| dc.identifier.uri | https://doi.org/10.1016/j.softx.2025.102303 | |
| dc.identifier.uri | https://hdl.handle.net/11508/55068 | |
| dc.identifier.volume | 31 | |
| dc.identifier.wos | WOS:001548766400001 | |
| dc.identifier.wosquality | Q2 | |
| dc.indekslendigikaynak | Web of Science | |
| dc.indekslendigikaynak | Scopus | |
| dc.language.iso | en | |
| dc.publisher | Elsevier | |
| dc.relation.ispartof | Softwarex | |
| dc.relation.publicationcategory | Makale - Uluslararası Hakemli Dergi - Kurum Öğretim Elemanı | |
| dc.rights | info:eu-repo/semantics/openAccess | |
| dc.snmz | KA_WoS_20260511 | |
| dc.subject | Video summarization | |
| dc.subject | Query based summarization | |
| dc.subject | Vision language models | |
| dc.subject | Transformers | |
| dc.title | QUBVIS: query based multi-modal summarization system using CLIP based transformer and vision language models | |
| dc.type | Article |







