Can Human Preference Redefine AI Testing? A Side-by-Side Evaluation Platform for Large Language Models
| dc.contributor.author | Aslan, Yaprak | |
| dc.contributor.author | Kaymak, Asiye | |
| dc.contributor.author | Mangan, EbrarSena | |
| dc.contributor.author | Selçuk, Ömer Faruk | |
| dc.contributor.author | Baykara, Muhammet | |
| dc.date.accessioned | 2026-09-08T07:08:34Z | |
| dc.date.issued | 2026 | |
| dc.department | Fırat Üniveristesi | |
| dc.description | 2nd International Symposium on AI-Driven Engineering Systems, ISADES 2026 -- 19 June 2026 through 20 June 2026 -- Hybrid, Mbale -- 226193 | |
| dc.description.abstract | The growing prevalence of large language models (LLMs) has made it necessary to measure the performance and output quality of these systems on an objective basis. The chief objective of this study is to develop a user-centric, interactive testing platform that evaluates LLM performance throughout various tasks. Thanks to the designed system architecture, a single input (prompt) provided by the user is simultaneously sent to multiple models in the background. The results obtained can be compared side-by-side. Data was collected from 35 volunteer users during the beta tests. It was determined that the ranking algorithm, powered by crowd-sourced votes, matches global LMSYS Chatbot Arena data at an 82.8% rate (Spearman correlation) within the scope of this pilot study. Given the initial sample size, these findings demonstrate promising local consistency rather than definitive global alignment. Additionally, the platform achieved a score of 81.93 on the System Usability Scale (SUS). Current methods generally focus on limited binary tests. This testing tool, however, combines multi-model comparison into a single interactive interface, supplying a complete evaluation perspective. © 2026 IEEE. | |
| dc.identifier.doi | 10.1109/ISADES69945.2026.11608247 | |
| dc.identifier.isbn | 979-831953447-7 | |
| dc.identifier.scopus | 2-s2.0-105046130392 | |
| dc.identifier.scopusquality | N/A | |
| dc.identifier.uri | https://doi.org/10.1109/ISADES69945.2026.11608247 | |
| dc.identifier.uri | https://hdl.handle.net/11508/64953 | |
| dc.indekslendigikaynak | Scopus | |
| dc.language.iso | en | |
| dc.publisher | Institute of Electrical and Electronics Engineers Inc. | |
| dc.relation.ispartof | ISADES 2026 - 2nd International Symposium on AI-Driven Engineering Systems, Proceedings | |
| dc.relation.publicationcategory | Konferans Öğesi - Uluslararası - Kurum Öğretim Elemanı | |
| dc.rights | info:eu-repo/semantics/closedAccess | |
| dc.snmz | KA_Scopus_20250903 | |
| dc.subject | Ai Testing Processes | |
| dc.subject | Comparative Study | |
| dc.subject | Human-Preference-Based Testing | |
| dc.subject | Large Language Models | |
| dc.subject | Multi-Model Evaluation | |
| dc.subject | Promptbased Performance Measurement | |
| dc.subject | Software Quality Assurance | |
| dc.title | Can Human Preference Redefine AI Testing? A Side-by-Side Evaluation Platform for Large Language Models | |
| dc.type | Conference Object |







