The quality of examination items plays a crucial role in ensuring the accuracy and fairness of educational assessment. This study aimed to evaluate the quality of Islamic Religious Education (PAI) Final Semester Examination items for Grade X students at SMK Negeri 3 Kolaka Timur using the Classical Test Theory (CTT) framework. Specifically, the study examined item validity, reliability, difficulty level, and discrimination power to determine the effectiveness of the examination instrument in measuring student achievement. A descriptive quantitative research design was employed using data derived from 23 students’ answer sheets, examination blueprints, answer keys, and supporting interview data from the subject teacher. Data were analyzed through Pearson Product-Moment correlation for item validity, Cronbach’s Alpha for reliability, difficulty index analysis, and discrimination index analysis. The findings revealed that 14 out of 20 items (70%) met the validity criteria, while 6 items (30%) were classified as invalid. The instrument demonstrated high reliability with a Cronbach’s Alpha coefficient of 0.828, indicating strong internal consistency. However, the difficulty level analysis showed an imbalanced distribution of items, with 65% categorized as easy and 35% as moderate, while no difficult items were identified. Furthermore, discrimination analysis indicated that 75% of the items exhibited poor discrimination power and only 25% achieved a fair category. These results suggest that although the instrument was reliable, it was less effective in differentiating students according to their achievement levels. The novelty of this study lies in its comprehensive application of Classical Test Theory to evaluate Islamic Religious Education examination items within a vocational secondary school context, an area that remains underexplored in assessment research. This study contributes to the literature on educational measurement by demonstrating that high reliability alone does not guarantee overall test quality and highlighting the importance of integrating validity, difficulty level, and discrimination analyses in the development of evidence-based assessment practices.