Mutiara Teccalonica Simanjuntak
Institut Teknologi Del

Published : 1 Documents Claim Missing Document
Claim Missing Document
Check
Articles

Found 1 Documents
Search

Evaluating LLM-Based Institutional Information Chatbot Responses Using a Preliminary Human-Scored Analytic Rubric and Automatic Metrics Rosni Lumbantoruan; Arnaldo Marulitua Sinaga; Markus Pardianto Hutagalung; Priskila Christine Natalia Parapat; Mutiara Teccalonica Simanjuntak
Elinvo (Electronics, Informatics, and Vocational Education) Vol. 11 No. 1 (2026): May 2026
Publisher : Department of Electronic and Informatic Engineering Education, Faculty of Engineering, UNY

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.21831/elinvo.v11i1.91571

Abstract

Large language models (LLMs) are increasingly used for answering institutional information enquiries in higher education, yet the quality of responses is not straightforward to evaluate, as factual accuracy alone does not account for interactive qualities such as clarity, conversational flow, error handling and personalisation. This pilot study developed, preliminary examined a human-scored analytic rubric for assessing ChatGPT responses in a higher-education institutional-information setting and explored the alignment of selected rubric scores with reference-based automatic metrics. We designed a literature-informed rubric comprising 15 criteria across five conceptual domains. Of these, 14 criteria were operationalised through 42 rubric questions, and system usability was rated separately using the System Usability Scale. Seventy-five qualified students from one higher-education institution rated the chatbot responses using a four-point scale. Preliminary evidence at the item level was provided by item-total correlations and Cronbach’s alpha, while Pearson and Spearman correlations were used to investigate the alignment between human scores and reference-based metrics namely ROUGE-1, ROUGE-2, ROUGE-L and SacreBLEU for four content-oriented criteria. The results showed positive but partial agreement between human ratings and referenced-based metrics, with stronger agreement for clarity and up-to-date response than for accuracy and relevance. These findings suggest that reference-based metrics can complement, but not replace, human evaluation for insitutional information chatbot assessment. The study was confined to one institution and did not incorporate inter-rater reliability, expert validation or factor analysis. Thus, the rubric should be seen as a preliminary evaluation instrument rather than a fully validated scale.