Comparison of human-AI agreement in ASA scoring by gender and duration of clinical experience: a real-world study

dc.contributor.authorCatak, Tuba
dc.contributor.authorAksu, Ahmet
dc.contributor.authorSaltali, Ali Ozgul
dc.contributor.authorBerilgen, Busra
dc.date.accessioned2026-08-12T17:28:38Z
dc.date.issued2026
dc.departmentFırat Üniversitesi
dc.description.abstractBackground Accurate preoperative risk identification is critical for patient safety and postoperative outcomes. Anaesthesiologists make decisions on the basis of ASA classification and additional parameters. Artificial intelligence (Al)-based decision support may offer more objective judgments. Methods In this retrospective multi-rater study, four anaesthesiologists and an Al system independently evaluated 1,000 cases. ASA class, postoperative ICU requirement, anaesthesia preference, intraoperative risk prediction, and additional recommendations were assessed. Concordance was analysed using Krippendorff's alpha, Cohen's kappa, Gwet's AC2, and PABAK, with percentage agreement estimated by bootstrapping. Al-physician agreement was further examined using fixed-effects logistic regression including clinician sex and professional experience as covariates. Results Physician-physician agreement was generally good to excellent across outcomes, whereas physician-Al agreement was lower and variable when assessed using K, PABAK, Gwet's ACZ, and observed agreement (P.). The highest Al concordance was observed for intraoperative risk prediction and ICU requirement, while the lowest was for anaesthesia preference. Exploratory analyses suggested that Al-physician concordance may vary by clinician experience and sex; no significant effects of sex or experience were observed for intraoperative anaesthesia-related risk prediction. Conclusion Although Al shows high concordance with physician decisions in objective/algorithmic domains, concordance remains limited in contextual and experience-based domains (anaesthesia preference). The findings support positioning Al as a safe 'second eye/warning' tool within human-in-the-loop workflows, rather than as an independent authority. Prospective, externally validated studies are needed.
dc.identifier.doi10.1186/s12911-026-03399-z
dc.identifier.issn1472-6947
dc.identifier.issue1
dc.identifier.pmid41709240
dc.identifier.scopus2-s2.0-105032220696
dc.identifier.scopusqualityQ1
dc.identifier.urihttps://doi.org/10.1186/s12911-026-03399-z
dc.identifier.urihttps://hdl.handle.net/11508/55366
dc.identifier.volume26
dc.identifier.wosWOS:001712157300001
dc.identifier.wosqualityQ2
dc.indekslendigikaynakWeb of Science
dc.indekslendigikaynakScopus
dc.indekslendigikaynakPubMed
dc.language.isoen
dc.publisherBmc
dc.relation.ispartofBmc Medical Informatics and Decision Making
dc.relation.publicationcategoryMakale - Uluslararası Hakemli Dergi - Kurum Öğretim Elemanı
dc.rightsinfo:eu-repo/semantics/openAccess
dc.snmzKA_WoS_20260511
dc.subjectAgreement
dc.subjectAnaesthesia
dc.subjectArtificial intelligence
dc.subjectASA physical status
dc.subjectConcordance
dc.titleComparison of human-AI agreement in ASA scoring by gender and duration of clinical experience: a real-world study
dc.typeArticle

Dosyalar