Can Human Preference Redefine AI Testing? A Side-by-Side Evaluation Platform for Large Language Models

dc.contributor.authorAslan, Yaprak
dc.contributor.authorKaymak, Asiye
dc.contributor.authorMangan, EbrarSena
dc.contributor.authorSelçuk, Ömer Faruk
dc.contributor.authorBaykara, Muhammet
dc.date.accessioned2026-09-08T07:08:34Z
dc.date.issued2026
dc.departmentFırat Üniveristesi
dc.description2nd International Symposium on AI-Driven Engineering Systems, ISADES 2026 -- 19 June 2026 through 20 June 2026 -- Hybrid, Mbale -- 226193
dc.description.abstractThe growing prevalence of large language models (LLMs) has made it necessary to measure the performance and output quality of these systems on an objective basis. The chief objective of this study is to develop a user-centric, interactive testing platform that evaluates LLM performance throughout various tasks. Thanks to the designed system architecture, a single input (prompt) provided by the user is simultaneously sent to multiple models in the background. The results obtained can be compared side-by-side. Data was collected from 35 volunteer users during the beta tests. It was determined that the ranking algorithm, powered by crowd-sourced votes, matches global LMSYS Chatbot Arena data at an 82.8% rate (Spearman correlation) within the scope of this pilot study. Given the initial sample size, these findings demonstrate promising local consistency rather than definitive global alignment. Additionally, the platform achieved a score of 81.93 on the System Usability Scale (SUS). Current methods generally focus on limited binary tests. This testing tool, however, combines multi-model comparison into a single interactive interface, supplying a complete evaluation perspective. © 2026 IEEE.
dc.identifier.doi10.1109/ISADES69945.2026.11608247
dc.identifier.isbn979-831953447-7
dc.identifier.scopus2-s2.0-105046130392
dc.identifier.scopusqualityN/A
dc.identifier.urihttps://doi.org/10.1109/ISADES69945.2026.11608247
dc.identifier.urihttps://hdl.handle.net/11508/64953
dc.indekslendigikaynakScopus
dc.language.isoen
dc.publisherInstitute of Electrical and Electronics Engineers Inc.
dc.relation.ispartofISADES 2026 - 2nd International Symposium on AI-Driven Engineering Systems, Proceedings
dc.relation.publicationcategoryKonferans Öğesi - Uluslararası - Kurum Öğretim Elemanı
dc.rightsinfo:eu-repo/semantics/closedAccess
dc.snmzKA_Scopus_20250903
dc.subjectAi Testing Processes
dc.subjectComparative Study
dc.subjectHuman-Preference-Based Testing
dc.subjectLarge Language Models
dc.subjectMulti-Model Evaluation
dc.subjectPromptbased Performance Measurement
dc.subjectSoftware Quality Assurance
dc.titleCan Human Preference Redefine AI Testing? A Side-by-Side Evaluation Platform for Large Language Models
dc.typeConference Object

Dosyalar