The quality of instructional evaluation instruments often does not fully reflect the intended learning objectives. Daily assessment tests used to measure students’ competency achievement may still exhibit weaknesses in terms of content validity, difficulty level, discrimination power, and distractor effectiveness. Such conditions can reduce the accuracy of information obtained by teachers regarding students’ learning achievement, thereby affecting the precision of instructional follow-up. This study aims to analyze the quality of test items and test implementation in daily assessments on the topic of biodiversity. A quantitative method with a descriptive approach was employed. The research sample consisted of 86 twelfth-grade social science students who had previously received instruction on biodiversity. Data analysis techniques involved inferential statistics. The findings indicate that the quality of the test items varied. In terms of content validity, most items aligned with the learning indicators; however, several items did not fully represent the targeted competencies. The analysis of item difficulty revealed an imbalanced proportion of easy, moderate, and difficult items, which limited comprehensive measurement of students’ abilities. Based on the results of validity testing, reliability analysis, item discrimination analysis, and difficulty level analysis, the instrument was deemed suitable for use as a daily assessment tool for evaluating learning outcomes on biodiversity content. The implications of this study suggest that suboptimal item quality may affect the accuracy of information on students’ learning achievement, leading to less effective instructional follow-up.
Copyrights © 2025