Journal of Evrimata: Engineering and Physics
Vol. 04 No. 01, 2026

Beyond Aggregate Metrics: Item-Level Performance Variability in SBERT-Based Automated Short Answer Scoring for Programming Education

Pramana Yoga Saputra (Department of Informatics Engineering, Information Technology, State Polytechnic of Malang, Malang, Indonesia)
Yan Watequlis Syaifudin (Department of Informatics Engineering, Information Technology, State Polytechnic of Malang, Malang, Indonesia)
Dika Rizky Yunianto (Department of Informatics Engineering, Information Technology, State Polytechnic of Malang, Malang, Indonesia)
Triana Fatmawati (Department of Informatics Engineering, Information Technology, State Polytechnic of Malang, Malang, Indonesia)
Rossa Akmalia (Department of Informatics Engineering, Information Technology, State Polytechnic of Malang, Malang, Indonesia)
Farida Ariany (Department of Informatics Engineering, Information Technology, State Polytechnic of Malang, Malang, Indonesia)
Yuri Ariyanto (Department of Informatics Engineering, Information Technology, State Polytechnic of Malang, Malang, Indonesia)



Article Info

Publish Date
30 Jun 2026

Abstract

Automated short answer scoring (ASAG) systems based on Sentence-BERT (SBERT) and cosine similarity are commonly evaluated using aggregate metrics such as Mean Absolute Error (MAE), Mean Squared Error (MSE), and Pearson correlation, averaged across an entire item set. Aggregate reporting can conceal substantial variability at the level of individual items, masking failure patterns relevant to instructional reliability. This study re-examines item-level scoring data from a deployed SBERT-cosine similarity ASAG module within the BAJAPRO Basic Java Programming platform, comparing a single-reference synonym-expansion strategy against a multi-reference text-preprocessing strategy across 21 short explanatory items. Rather than reporting averages alone, performance is decomposed per item, revealing MAE ranging from 0.0007 to 0.070 and Pearson correlation ranging from 0.20 to 0.96 within a single evaluation condition, a range invisible in prior aggregate-only reporting. A qualitative failure case further shows SBERT-based similarity failing to distinguish semantically opposite mathematical operations expressed in near-identical sentence structures. The study explicitly reports its inability to verify item-to-concept mappings due to undocumented original evaluation scripts, treating this as a methodological finding on reproducibility practice in ASAG research rather than a limitation to be concealed. The findings argue for routine item-level reporting and improved experimental documentation in future SBERT-based ASAG studies.

Copyrights © 2026






Journal Info

Abbrev

EEP

Publisher

Subject

Automotive Engineering Chemical Engineering, Chemistry & Bioengineering Chemistry Civil Engineering, Building, Construction & Architecture Computer Science & IT

Description

- Engineering (miscellaneous) - Civil and Structural Engineering - Electrical and Electronic Engineering - Mechanical Engineering - Chemical Engineering - Physics - Computer Science - ...