Evaluation of accuracy, clinical reliability and readability of LLM-based chatbot responses in prosthetic dentistry FAQs
| dc.contributor.author | Tartuk, Bülent Kadir | |
| dc.contributor.author | Altıntaş, Eyyüp | |
| dc.date.accessioned | 2026-09-08T07:06:59Z | |
| dc.date.issued | 2026 | |
| dc.department | Fırat Üniveristesi | |
| dc.description.abstract | Aims: Evidence comparing multiple contemporary large language model (LLM)-based chatbots in prosthetic dentistry using multidimensional outcome measures remains limited. This study comparatively evaluated the responses generated by ChatGPT, Gemini, Copilot and DeepSeek to frequently asked questions (FAQs) related to prosthetic dentistry in terms of accuracy, clinical reliability and readability. Methods: Thirty-nine FAQs obtained from publicly available patient education resources were equally distributed across fixed, removable and implant-supported prosthesis categories (n=13 each). Questions were submitted in Turkish on the same day under standardized conditions to ChatGPT, Gemini, Copilot and DeepSeek Chatbots, all of which were accessed through their publicly available web interfaces. Responses generated in Turkish were independently scored by three prosthodontists using five-point Likert scales to assess accuracy and clinical reliability. Readability was assessed using the Ateşman and Bezirci-Yılmaz formulas. Inter-rater agreement was analyzed using the intraclass correlation coefficient (ICC). Repeated-measures comparisons were performed using the Friedman test, followed by Bonferroni-adjusted pairwise Wilcoxon signed-rank tests. Effect sizes were reported using Kendall’s W. Results: Inter-rater agreement was high for accuracy (ICC=0.86) and clinical reliability (ICC=0.83). Significant inter-system differences were observed in accuracy, clinical reliability, and readability outcomes (all p<0.001; Kendall’s W=0.31-0.46). ChatGPT demonstrated the highest accuracy and most favorable readability values, whereas Gemini showed the highest clinical reliability scores. Copilot and DeepSeek generally exhibited lower performances. Implant-related questions yielded significantly lower accuracy and reliability scores than fixed and removable prosthesis questions (p<0.05). Conclusion: LLM-based chatbots demonstrated heterogeneous performance in answering questions related to prosthetic dentistry. Although some systems may assist preliminary patient education, meaningful differences in clinical reliability and readability indicate that chatbot outputs should be interpreted cautiously and reviewed by dental professionals, particularly for implant-related topics. | |
| dc.identifier.doi | 10.38053/acmj.1905988 | |
| dc.identifier.endpage | 523 | |
| dc.identifier.issn | 2718-0115 | |
| dc.identifier.issue | 3 | |
| dc.identifier.startpage | 513 | |
| dc.identifier.trdizinid | 1424735 | |
| dc.identifier.uri | https://doi.org/10.38053/acmj.1905988 | |
| dc.identifier.uri | https://search.trdizin.gov.tr/tr/yayin/detay/1424735 | |
| dc.identifier.uri | https://hdl.handle.net/11508/64840 | |
| dc.identifier.volume | 8 | |
| dc.indekslendigikaynak | TR-Dizin | |
| dc.language.iso | en | |
| dc.relation.ispartof | Anatolian Current Medical Journal | |
| dc.relation.publicationcategory | Makale - Ulusal Hakemli Dergi - Kurum Öğretim Elemanı | |
| dc.rights | info:eu-repo/semantics/openAccess | |
| dc.snmz | KA_TR_20250903 | |
| dc.subject | Bilgisayar Bilimleri | |
| dc.subject | Yazılım Mühendisliği | |
| dc.subject | Diş Hekimliği | |
| dc.title | Evaluation of accuracy, clinical reliability and readability of LLM-based chatbot responses in prosthetic dentistry FAQs | |
| dc.type | Article |







