Citation
Shabbir, Muhammad and Bandy, Altaf and Mehboob, Bushra and Mahboob, Usman and Liaqat, Ambreen and Almarshad, Feras and Adam, Siti Khadijah
(2026)
Quality of AI vs human-generated Single Best Answer questions: a systematic review and meta-analysis.
Medicni Perspektivi, 31 (2).
pp. 135-148.
ISSN 2307-0404; eISSN: 2786-4804
Abstract
Single Best Answer Questions (SBAs) are essential and resource-intensive assessment tools in health professions education. Artificial Intelligence (AI), such as large language models (LLMs), can automate the creation of SBAs; however, evidence comparing the quality of AI-generated and human-created items is still dispersed. The purpose of the study is to compare the psychometric quality, measured by difficulty and discrimination indices, of AIgenerated SBAs with those authored by humans in health professions education. The current study followed PRISMA guidelines. The search was conducted on Scopus, PubMed, and Google Scholar. Studies published through April 25th, 2025, and those that directly compared AI- and human-generated SBAs and reported the mean, standard deviation, and sample size for both difficulty and discrimination indices were included. Two reviewers independently extracted the data. Standardized mean differences (SMDs) were calculated and combined using random-effects models (Jamovi MAJOR module, version 2.6.44-06 March 2025). Heterogeneity and publication bias were assessed. Four studies met the inclusion criteria, providing eight comparison outcomes (4 for difficulty, 4 for discrimination). The combined analysis of both outcomes revealed no statistically significant difference, overall (SMD= -0.084, 95% CI: -0.65 to 0.49, p=0.773); however, the heterogeneity was very high (I²=92.7%). Separate analyses revealed that AI-generated questions were significantly easier than human-generated questions (SMD= +0.541, 95% CI: 0.17 to 0.91, p=0.004; I²=62.3%). Conversely, human-authored questions demonstrated significantly higher discrimination indices than AI-generated questions (SMD= -0.701, 95% CI: -1.33 to -0.08, p=0.028; I²=86.2%). No evidence of publication bias was found. AI-generated items tend to be easier, potentially aiding accessibility, whereas human-authored items currently exhibit superior discriminatory power, which is crucial for robust assessment. High heterogeneity underscores context dependency.
Download File
Additional Metadata
Actions (login required)
 |
View Item |