TRANSFORMER TABANLI MODELLER KULLANARAK GÖRÜNTÜLERİ BETİMLEYEN RESİM ALTYAZILARININ ÜRETİLMESİ
| dc.contributor.advisor | DAŞ, BİHTER | |
| dc.contributor.author | KARACA, ZEYNEP | |
| dc.date.accessioned | 2026-08-12T10:16:27Z | |
| dc.date.issued | 2024 | |
| dc.department | FÜ, Fen Bilimleri Enstitüsü, Yazılım Mühendisliği Anabilim Dalı | |
| dc.description.abstract | Doğal dil işleme alanının oldukça zorlu bir görevi olan resim yazısı oluşturma işlemi, bilgisayar tarafından insan dil yapısına en uygun şekilde otomatik olarak altyazı üretilmesi olarak adlandırılmaktadır. Resim yazısı oluşturulurken en önemli hedef, görüntüyü en doğru ve en iyi şekilde açıklayan cümleler oluşturmaktadır. Bu doğrultuda transformer tabanlı dil modellerinin kullanılması, cümle performansını büyük ölçüde etkilemektedir. Bu tez çalışmasında, görüntülerden resim yazısı üretilmesinde kullanılan transformer tabanlı dil modellerinin karşılaştırmalı analizi, derin öğrenme tabanlı dil modelleri ve derin öğrenme yöntemlerinin performansı incelenmektedir. Belirlenen bu hedef kapsamında iki farklı veri kümesinde geliştirilen dört uygulamamız bulunmaktadır. Uygulamalarımızda MSCOCO ve Flickr8k veri kümeleri kullanılarak, görüntüleri açıklayan ingilizce dilinde cümle üretilmesi gerçekleştirilmektedir. | |
| dc.description.abstract | Image caption creation, which is a very challenging task in the field of natural language processing, is called the automatic production of subtitles by the computer in the most appropriate way to the human language structure. The most important goal when creating a caption is to create sentences that describe the image in the most accurate and best way. In this regard, the use of transformer-based language models greatly affects sentence performance. In this thesis, the comparative analysis of transformer-based language models used in generating captions from images, the performance of convolution-based language models, and deep learning methods are examined. Within the scope of this determined goal, we have four applications developed on two different data sets. In our applications, sentences in English language describing the images are produced by using MSCOCO and Flickr8k datasets. | |
| dc.identifier.citation | KARACA, Z. (2024). Transformer tabanlı modeller kullanarak görüntüleri betimleyen resim altyazılarının üretilmesi (Tez No. 885392) [Yüksek lisans tezi, Fırat Üniversitesi]. | |
| dc.identifier.uri | https://tez.yok.gov.tr/UlusalTezMerkezi/TezGoster?key=usXiZIM9Lp0wk-YzRoaT-3ttC6g-E42ANnbhKXjMm0VTyGBjsa49TycksJ0caw-n | |
| dc.identifier.uri | https://hdl.handle.net/11508/24099 | |
| dc.identifier.yoktezid | 885392 | |
| dc.language.iso | tr | |
| dc.publisher | Fırat Üniveristesi | |
| dc.relation.publicationcategory | Tez | |
| dc.rights | info:eu-repo/semantics/openAccess | |
| dc.snmz | KA_TEZ_20260511 | |
| dc.subject | Bilgisayar Mühendisliği Bilimleri-Bilgisayar ve Kontrol | |
| dc.title | TRANSFORMER TABANLI MODELLER KULLANARAK GÖRÜNTÜLERİ BETİMLEYEN RESİM ALTYAZILARININ ÜRETİLMESİ | |
| dc.title.alternative | Generating image captions describing images using transformer-based models | |
| dc.type | Master Thesis |







