Maulida Nurhidayati
Universitas Islam Negeri Kiai Ageng Muhammad Besari Ponorogo, Indonesia

Published : 1 Documents Claim Missing Document
Claim Missing Document
Check
Articles

Found 1 Documents
Search

Are We Measuring Ability or Guessing?: CTT and IRT Evidence from a Multiple-Choice Assessment in Econometric Test Ajeng Wahyuni; Yunaita Rahmawati; Maulida Nurhidayati; Muhtadin Amri
Hipotenusa: Journal of Mathematical Society Vol. 8 No. 1 (2026): Hipotenusa : Journal of Mathematical Society
Publisher : Program Studi Tadris Matematika Universitas Islam Negeri (UIN) Salatiga

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.18326/hipotenusa.v8i1.7088

Abstract

Multiple-choice tests are widely used in mathematics-related higher education courses because they are practical for assessing broad learning outcomes. Correct responses may not always indicate full conceptual mastery, as students may answer correctly through partial knowledge, distractor elimination, unintended item cues, or pseudo-guessing. This study evaluates the quality of a 20-item four-option multiple-choice econometrics assessment using Classical Test Theory (CTT) and Item Response Theory (IRT). The test was administered to 108 undergraduate students and was designed to measure econometrics competence as an applied mathematics construct, including quantitative and statistical reasoning, regression and model interpretation, hypothesis testing and inference, model assumptions and diagnostics, and data-based decision-making. CTT was used to examine item difficulty, item discrimination, while IRT was used to compare the 1PL, 2PL, and 3PL models and to diagnose pseudo-guessing. The results showed a mean score of 13.85 out of 20, KR-20 of 0.681, and Cronbach’s alpha of 0.677, indicating moderate but not strong internal consistency. CTT identified no difficult items, nine easy items, and three items with poor discrimination. The 1PL model had the lowest BIC and was therefore the most fit model, while the 3PL model was retained diagnostically because it estimates pseudo-guessing. Eight items, namely I01, I05, I06, I12, I15, I16, I18, and I20, had pseudo-guessing parameters above 0.25. These findings suggest that some correct responses may have been influenced by non-mastery factors. This study contributes to mathematics education by demonstrating how integrated CTT and IRT diagnostics can improve the validity of econometrics assessment as a measure of quantitative and statistical reasoning.