ELINVO (Electronics, Informatics, and Vocational Education)
Vol. 11 No. 1 (2026): May 2026

Evaluating LLM-Based Institutional Information Chatbot Responses Using a Preliminary Human-Scored Analytic Rubric and Automatic Metrics

Rosni Lumbantoruan (Institut Teknologi Del)
Arnaldo Marulitua Sinaga (Institut Teknologi Del)
Markus Pardianto Hutagalung (Institut Teknologi Del)
Priskila Christine Natalia Parapat (Institut Teknologi Del)
Mutiara Teccalonica Simanjuntak (Institut Teknologi Del)



Article Info

Publish Date
16 Jul 2026

Abstract

Large language models (LLMs) are increasingly used for answering institutional information enquiries in higher education, yet the quality of responses is not straightforward to evaluate, as factual accuracy alone does not account for interactive qualities such as clarity, conversational flow, error handling and personalisation. This pilot study developed, preliminary examined a human-scored analytic rubric for assessing ChatGPT responses in a higher-education institutional-information setting and explored the alignment of selected rubric scores with reference-based automatic metrics. We designed a literature-informed rubric comprising 15 criteria across five conceptual domains. Of these, 14 criteria were operationalised through 42 rubric questions, and system usability was rated separately using the System Usability Scale. Seventy-five qualified students from one higher-education institution rated the chatbot responses using a four-point scale. Preliminary evidence at the item level was provided by item-total correlations and Cronbach’s alpha, while Pearson and Spearman correlations were used to investigate the alignment between human scores and reference-based metrics namely ROUGE-1, ROUGE-2, ROUGE-L and SacreBLEU for four content-oriented criteria. The results showed positive but partial agreement between human ratings and referenced-based metrics, with stronger agreement for clarity and up-to-date response than for accuracy and relevance. These findings suggest that reference-based metrics can complement, but not replace, human evaluation for insitutional information chatbot assessment. The study was confined to one institution and did not incorporate inter-rater reliability, expert validation or factor analysis. Thus, the rubric should be seen as a preliminary evaluation instrument rather than a fully validated scale.

Copyrights © 2026






Journal Info

Abbrev

elinvo

Publisher

Subject

Computer Science & IT Education Electrical & Electronics Engineering

Description

ELINVO (Electronics, Informatics and Vocational Education) is a peer-reviewed journal that publishes high-quality scientific articles in Indonesian language or English in the form of research results (the main priority) and or review studies in the field of electronics and informatics both in terms ...