Generative artificial intelligence (AI) has increasingly entered higher education, including chemistry education. Although AI can provide rapid and apparently coherent explanations, its outputs may contain factual, conceptual, representational, and reasoning errors. This study examines pre-service chemistry teachers’ ability to evaluate the conceptual accuracy of AI-generated answers to chemistry problems using the Rasch measurement model. The study was designed as a quantitative descriptive evaluation involving 120 pre-service chemistry teachers and 30 scenario-based items. Each item presented a chemistry problem followed by an AI-generated response containing different levels and types of accuracy. Rasch analysis was used to estimate person ability, item difficulty, reliability, separation, and item fit. The descriptive interpretation focused on the occurrence of response patterns and the characteristics of the most difficult and easiest items. The simulated analysis produced person reliability of 0.87 and item reliability of 0.96, with person and item separation indices of 2.56 and 4.89, respectively. The most difficult items involved organic structure, chemical equilibrium, and acid-base reasoning, whereas factual errors and simple calculation errors were more readily identified. The findings indicate that the ability to use AI should not be equated with the ability to critically evaluate AI outputs. Domain-specific AI literacy in chemistry requires students to mobilize conceptual knowledge and chemical reasoning to verify apparently plausible AI-generated explanations.