This study examines the reliability of an image-native multimodal AI system for automated scoring of handwritten responses to Advanced Placement (AP) Calculus free-response items without requiring optical character recognition (OCR) preprocessing. Using inter-rater agreement indices and test–retest reliability analyses, we found substantial to almost perfect agreement between artificial intelligence (AI)-generated scores and calibrated human ratings, as well as almost perfect stability across repeated scoring sessions. These results suggest that the observed reliability of the AI scoring system warrants further investigation of validity-related evidence and inferences. As a practical implication for assessment practice, we propose a human-in-the-loop teacher–AI collaborative (TAC) framework in which automated scoring operates under teacher oversight. Taken together, these findings provide initial evidence of reliability supporting the responsible use of AI-based scoring as a measurement instrument in high-stakes educational assessment.
Copyrights © 2026