Abstract
Single Best Answer Questions (SBAs) are essential and resource-intensive assessment tools in health professions education. Artificial Intelligence (AI), such as large language models (LLMs), can automate the creation of SBAs; however, evidence comparing the quality ofAI-generated and human-created items is still dispersed. The purpose of the study is to compare the psychometric quality, measured by difficulty and discrimination indices, ofAI-generated SBAs with those authored by humans in health professions education. The current study followed PRISMA guidelines. The search was conducted on Scopus, PubMed, and Google Scholar. Studies published through April 25th, 2025, and those that directly compared AI-and human-generated SBAs and reported the mean, standard deviation, and sample size for both difficulty and discrimination indices were included. Two reviewers independently extracted the data. Standardized mean differences (SMDs) were calculated and combined using random-effects models (Jamovi MAJOR module, version 2.6.44-06 March 2025). Heterogeneity and publication bias were assessed. Four studies met the inclusion criteria, providing eight comparison outcomes (4 for difficulty, 4 for discrimination). The combined analysis of both outcomes revealed no statistically significant difference, overall (SMD=-0.084, 95% CI:-0.65 to 0.49, p=0.773); however, the heterogeneity was very high (I2=92.7%). Separate analyses revealed that AI-generated questions were significantly easier than human-generated questions (SMD= +0.541, 95% CI: 0.17 to 0.91, p=0.004; I2=62.3%). Conversely, human-authored questions demonstrated significantly higher discrimination indices than AI-generated questions (SMD=-0.701, 95% CI:-1.33 to-0.08, p=0.028; I2=86.2%). No evidence of publication bias was found. AI-generated items tend to be easier, potentially aiding accessibility, whereas human-authored items currently exhibit superior discriminatory power, which is crucial for robust assessment. High heterogeneity underscores context dependency.
| Original language | English |
|---|---|
| Pages (from-to) | 135-148 |
| Number of pages | 14 |
| Journal | Medychni Perspektyvy |
| Volume | 31 |
| Issue number | 2 |
| DOIs | |
| Publication status | Published - 2026 |
Keywords
- artificial intelligence
- item difficulty
- item discrimination
- large language models
- medical education
- meta-analysis
- psychometrics
- single best answer
- MULTIPLE-CHOICE
Fingerprint
Dive into the research topics of 'Quality Of Ai Vs Human-Generated Single Best Answer Questions: A Systematic Review And Meta-Analysis'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver