Concept map-based assessment is a practical approach to measure students’ conceptual understanding, but manual assessment still faces challenges such as subjectivity, inconsistency, and limited scalability. This study proposes the application of Retrieval-Augmented Generation (RAG) as an artificial intelligence-based automated assessment solution in an educational context. The objectives of this study are to compare the effectiveness of two retrieval methods, BM25 and FAISS, and to analyse the trade-off between large-scale generative models (GPT) and Small-LLM in assessing concept map propositions. This study uses a quantitative experimental approach by combining a retriever and a generator in the RAG system. Performance evaluation is carried out using the Macro-F1 and QWK metrics to measure agreement with expert judgment, and the Explanation Relevance Score (ERS) to assess explanation quality. The experimental results show that the FAISS–GPT combination achieves the best performance, with a Macro-F1 of 0.338 and a QWK of 0.146, slightly superior to the BM25–GPT combination. In contrast, the use of Small-LLM, both with BM25 and FAISS, showed lower performance with Macro-F1 values in the range of 0.167–0.221 and QWK close to zero. This finding confirms that semantic-based retrieval plays a vital role in improving the accuracy of automated assessment, while large-scale generative models are more effective in representing conceptual relationships in depth. This study contributes through a comparative analysis of retrievers and generators, and by introducing ERS as an additional metric for RAG-based automated assessment in the field of education.
Copyrights © 2026