This study evaluates the reliability, algorithmic bias, and feasibility of implementing Human-in-the-Loop (HITL) in Large Language Model (LLM)-based Automated Essay Scoring (AES). Two architectures were compared: the Mixture-of-Experts architecture represented by GPT OSS 120B and the Dense architecture represented by Qwen3-32B, using a multiple-run scoring approach on 150 IELTS-like essays written by 50 respondents. Each essay was evaluated five times across four criteria: Grammar, Lexical Resource, Coherence, and Task Achievement. The evaluation employed the Intraclass Correlation Coefficient (ICC), Coefficient of Variation (CV), score range, paired t-test, and confidence-based routing simulation.The results show that GPT OSS 120B achieved higher reliability, with an ICC of 0.94, an average range of 2.40 points, and a CV of 4.69%, while Qwen3-32B obtained an ICC of 0.84, an average range of 5.56 points, and a CV of 9.17%. However, GPT OSS 120B experienced a parsing failure rate of 25.4%, whereas Qwen3-32B demonstrated full format compliance. Comparative analysis also showed that GPT OSS 120B tended to score more strictly, while Qwen3-32B was more lenient. In the Data Report task without visual input, both models exhibited conservatism bias due to contextual limitations. The HITL simulation showed that GPT OSS 120B could automatically approve 55.4% of essays, compared with only 7.8% for Qwen3-32B. These findings highlight the importance of HITL in maintaining the reliability, fairness, and integrity of academic assessment.
Copyrights © 2026