Automated short answer scoring (ASAG) systems based on Sentence-BERT (SBERT) and cosine similarity are commonly evaluated using aggregate metrics such as Mean Absolute Error (MAE), Mean Squared Error (MSE), and Pearson correlation, averaged across an entire item set. Aggregate reporting can conceal substantial variability at the level of individual items, masking failure patterns relevant to instructional reliability. This study re-examines item-level scoring data from a deployed SBERT-cosine similarity ASAG module within the BAJAPRO Basic Java Programming platform, comparing a single-reference synonym-expansion strategy against a multi-reference text-preprocessing strategy across 21 short explanatory items. Rather than reporting averages alone, performance is decomposed per item, revealing MAE ranging from 0.0007 to 0.070 and Pearson correlation ranging from 0.20 to 0.96 within a single evaluation condition, a range invisible in prior aggregate-only reporting. A qualitative failure case further shows SBERT-based similarity failing to distinguish semantically opposite mathematical operations expressed in near-identical sentence structures. The study explicitly reports its inability to verify item-to-concept mappings due to undocumented original evaluation scripts, treating this as a methodological finding on reproducibility practice in ASAG research rather than a limitation to be concealed. The findings argue for routine item-level reporting and improved experimental documentation in future SBERT-based ASAG studies.
Copyrights © 2026