Use of Large Language Models for Medical Synthetic Data Generation in Mental Illness

dc.contributor.authorAygün, İrfan
dc.contributor.authorKaya, Mehmet
dc.date.accessioned2026-08-12T16:09:06Z
dc.date.issued2023
dc.departmentFırat Üniversitesi
dc.description7th IET Smart Cities Symposium, SCS 2023 -- 3 December 2023 through 5 December 2023 -- Virtual, Online -- 199627
dc.description.abstractData quantity and quality are very important for the development of medical artificial intelligence research. Nowadays, thanks to easier access to data, studies in this field produce very successful results. However, many factors such as protection of patient rights in medical data and confidentiality of personal data prevent researchers from directly accessing the data. For this reason, synthetic data creation studies are often needed both to expand the training and test sets and to create sample cases to be used in the relevant field. In this study, various synthetic patient data are created to be presented to a language model that enables the detection of psychological disorders through patient text. Synthetic data sets were produced with 200 artificial patient data created with popular LLM examples ChatGPT and Google Bard. The quality of synthetic data was measured with the help of a pre-trained BERT model using these datasets. In the experiments, it was observed that chatbots that generate instant data, such as ChatGPT and Google Bard, produced successful results at rates of 89% and 86% with the language representation model. With the experimental results, it appears that LLM studies can provide more successful results than advanced language models in various medical text production tasks. © The Institution of Engineering & Technology 2023.
dc.description.sponsorshipTürkiye Bilimsel ve Teknolojik Araştırma Kurumu, TÜBİTAK, (122E437); Türkiye Bilimsel ve Teknolojik Araştırma Kurumu, TÜBİTAK
dc.identifier.doi10.1049/icp.2024.1033
dc.identifier.endpage656
dc.identifier.issn2732-4494
dc.identifier.issue44
dc.identifier.scopus2-s2.0-85194232069
dc.identifier.scopusqualityQ4
dc.identifier.startpage652
dc.identifier.urihttps://doi.org/10.1049/icp.2024.1033
dc.identifier.urihttps://hdl.handle.net/11508/41581
dc.identifier.volume2023
dc.indekslendigikaynakScopus
dc.language.isoen
dc.publisherInstitution of Engineering and Technology
dc.relation.ispartofIET Conference Proceedings
dc.relation.publicationcategoryKonferans Öğesi - Uluslararası - Kurum Öğretim Elemanı
dc.rightsinfo:eu-repo/semantics/closedAccess
dc.snmzKA_Scopus_20260511
dc.subjectChatGPT; Deep Learning; Google Bard; Large Language Models; Mental Illness
dc.titleUse of Large Language Models for Medical Synthetic Data Generation in Mental Illness
dc.typeConference Object

Dosyalar