This study aims to analyze scoring variations and test the statistical significance of differences between AI-based evaluation (Large Language Model) and human raters in assessing the oral performance of German proverbs. Using a quantitative, within-subjects design, the study involved 30 students who were evaluated on four parameters: grammar, pronunciation, fluency, and the use of proverbs. The results of a paired-samples t-test reveal highly significant differences (p < 0.001) across all assessment aspects. The AI consistently demonstrates a leniency bias, assigning higher absolute scores than human raters. The largest discrepancy is in the use of proverbs, with a mean difference of 3.73 points. These findings indicate that AI tends to operate at the level of structured linguistic surface features, yet remains limited in capturing cultural nuances, implicit meanings, and pragmatic contextual appropriateness. These are areas that constitute the core strength of human cognitive sensitivity. This study recommends implementing a hybrid app using language to evaluate, in which AI assesses structural-mechanistic aspects, while interpretive sociolinguistic dimensions remain under human raters' control.
Copyrights © 2026