Item quality plays a crucial role in ensuring that learning assessments accurately measure student understanding and support valid educational decision-making. However, teachers still face challenges in conducting empirical item analysis, particularly on science questions based on contextual phenomena, resulting in assessment instruments that may not optimally measure student competency. Therefore, this study aimed to analyze the quality of multiple-choice items on light and optics based on contextual phenomena using the test analysis program (TAP) version 14.7.4. A descriptive quantitative method was used involving 54 eighth-grade students from a private junior high school in Sidoarjo. The instrument consisted of 40 multiple-choice items that had been previously validated by two subject matter experts and one science teacher, obtaining an average validation score of 3.5, indicating that the instrument was suitable for use with minor revisions. The analysis showed that 26 items (65%) were valid, while 14 items (35%) were invalid. Reliability was dominated by the very low category (50%), discriminatory power by the satisfactory category (37.5%), difficulty by the moderate category (55%), and distractor effectiveness by the very low category (97.5%). Overall, 12 items (30%) were retained, 11 items (27.5%) required revision, 11 items (27.5%) required major revision or replacement, and 6 items (15%) were identified as potentially problematic. These findings indicate that although some items met acceptable quality standards, improvements are still needed, particularly in terms of validity, reliability, discriminatory power, and distractor effectiveness, to produce a more objective, accurate, consistent, and high-quality science assessment instrument based on contextual phenomena
Copyrights © 2026