Learning outcomes evaluation through summative assessment plays a crucial role in accurately capturing student competency achievement. However, field practices indicate that test instruments are frequently developed without meticulous psychometric testing, thereby triggering measurement distortion. This study aims to map the distribution of valid and invalid test items, as well as to identify the root causes of instrument invalidity across the dimensions of content, construction, and language readability. The research method employed is a descriptive quantitative approach integrated with qualitative content analysis. Data collection was conducted through documentation of the summative test manuscript for the Islamic Religious Education and Character (PAI-BP) subject of Class XI in the Light Vehicle Engineering (TKR) department at SMK Negeri Karangpucung, along with the response data from 35 students. The data were tabulated into a binary matrix (scores of 1 and 0) and analyzed quantitatively using Excel formulas assisted by SPSS software, then compared against an r-table value of 0.334. The results demonstrated that out of the 30 multiple-choice items tested, 14 items were declared valid and 16 items were declared invalid or discarded, including an anomaly of two test items yielding negative (minus) correlation coefficients. The root causes of invalidity were identified as non-homogeneous distractors, biased stems, typographical errors, and overly lengthy text stimuli that induced reading fatigue among vocational students. The implications of this study emphasize the urgency for educators to conduct a total reconstruction of flawed item banks and to formulate more concise, applicable questions to uphold accountability and fairness within the school assessment system.