This study aims to analyze the quality of Arabic language test items based on difficulty level, discrimination index, and distractor effectiveness in formative assessment. The research employed a descriptive quantitative approach involving 30 students of an intensive Arabic program. The instrument consisted of 20 multiple-choice items covering grammar (qawā‘id) and text comprehension. Data were analyzed using Classical Test Theory with the assistance of Microsoft Excel and SPSS. The results showed that 14 items were valid while 6 items were invalid. The reliability coefficient (Cronbach’s Alpha) was 0.716, indicating good internal consistency. The difficulty level analysis revealed that 80% of the items were categorized as easy and 20% as moderate, with no difficult items found. In terms of discrimination index, most items were classified as good to very good, indicating their effectiveness in distinguishing students’ abilities. Distractor analysis showed that most distractors functioned adequately, although some items had ineffective distractors due to low selection rates. These findings indicate that although the test instrument is generally reliable and has good discrimination power, improvements are needed in balancing item difficulty and enhancing distractor quality. Therefore, the development of more varied and higher-order thinking (HOTS)-based items is recommended to improve the overall quality of Arabic language assessment.