UPM Institutional Repository

Quality of AI vs human-generated Single Best Answer questions: a systematic review and meta-analysis


Citation

Shabbir, Muhammad and Bandy, Altaf and Mehboob, Bushra and Mahboob, Usman and Liaqat, Ambreen and Almarshad, Feras and Adam, Siti Khadijah (2026) Quality of AI vs human-generated Single Best Answer questions: a systematic review and meta-analysis. Medicni Perspektivi, 31 (2). pp. 135-148. ISSN 2307-0404; eISSN: 2786-4804

Abstract

Single Best Answer Questions (SBAs) are essential and resource-intensive assessment tools in health professions education. Artificial Intelligence (AI), such as large language models (LLMs), can automate the creation of SBAs; however, evidence comparing the quality of AI-generated and human-created items is still dispersed. The purpose of the study is to compare the psychometric quality, measured by difficulty and discrimination indices, of AIgenerated SBAs with those authored by humans in health professions education. The current study followed PRISMA guidelines. The search was conducted on Scopus, PubMed, and Google Scholar. Studies published through April 25th, 2025, and those that directly compared AI- and human-generated SBAs and reported the mean, standard deviation, and sample size for both difficulty and discrimination indices were included. Two reviewers independently extracted the data. Standardized mean differences (SMDs) were calculated and combined using random-effects models (Jamovi MAJOR module, version 2.6.44-06 March 2025). Heterogeneity and publication bias were assessed. Four studies met the inclusion criteria, providing eight comparison outcomes (4 for difficulty, 4 for discrimination). The combined analysis of both outcomes revealed no statistically significant difference, overall (SMD= -0.084, 95% CI: -0.65 to 0.49, p=0.773); however, the heterogeneity was very high (I²=92.7%). Separate analyses revealed that AI-generated questions were significantly easier than human-generated questions (SMD= +0.541, 95% CI: 0.17 to 0.91, p=0.004; I²=62.3%). Conversely, human-authored questions demonstrated significantly higher discrimination indices than AI-generated questions (SMD= -0.701, 95% CI: -1.33 to -0.08, p=0.028; I²=86.2%). No evidence of publication bias was found. AI-generated items tend to be easier, potentially aiding accessibility, whereas human-authored items currently exhibit superior discriminatory power, which is crucial for robust assessment. High heterogeneity underscores context dependency.


Download File

[img] Text
128045.pdf - Published Version
Available under License Creative Commons Attribution.

Download (987kB)

Additional Metadata

Item Type: Article
Subject: Medicine (all)
Divisions: Faculty of Medicine and Health Science
DOI Number: https://doi.org/10.26641/2307-0404.2026.2.365915
Publisher: Dnipro State Medical University
Keywords: artificial intelligence; item difficulty; item discrimination; large language models; medical education; meta-analysis; psychometrics; single best answer
Sustainable Development Goals (SDGs): SDG 4: Quality Education, SDG 3: Good Health and Well-being, SDG 9: Industry, Innovation and Infrastructure
Depositing User: Ms. Siti Radziah Mohamed@mahmod
Date Deposited: 26 Aug 2026 07:53
Last Modified: 26 Aug 2026 07:53
Altmetrics: http://www.altmetric.com/details.php?domain=psasir.upm.edu.my&doi=10.26641/2307-0404.2026.2.365915
URI: http://psasir.upm.edu.my/id/eprint/128045
Statistic Details: View Download Statistic

Actions (login required)

View Item View Item