Multiple-choice item quality analysis is essential for educational assessment quality assurance, yet conventional psychometric analysis relies on examinee response data that are often unavailable during item development. This study develops an explainable machine learning framework based on simulated responses to analyze multiple-choice item quality using 10,962 CommonsenseQA items. The dataset serves as a computational testbed due to its consistent five-option format, available answer keys, large scale, and plausible distractors, rather than as a substitute for validated educational item banks. Responses from 200 virtual respondents per item were probabilistically simulated based on answer-key weighting and surface-level distractor plausibility. Four psychometric indicators—difficulty index, discrimination index, distractor effectiveness, and distractor entropy—combined with five surface linguistic features were used to train XGBoost to classify item quality as poor, moderate, or good. Five-fold cross-validation yielded a weighted F1-score of 0.9987 ± 0.0011 and accuracy of 0.9987 ± 0.0011. Feature importance and SHAP identified distractor entropy and discrimination index as dominant predictors. The near-perfect performance is not interpreted as external predictive validity because the target and primary predictors derive from the same simulation mechanism. This framework provides a transparent, auditable proof-of-concept for preliminary item-bank evaluation before empirical testing.
Copyrights © 2026