The increasing administrative workload of teachers makes manual preparation of assessment questions time-consuming and may lead to the reuse of questions from previous years. This study examines the application of the transformer-based Generative Pre-trained Transformer 2 (GPT-2) model for automatic question-and-answer generation in elementary school Natural and Social Sciences (IPAS). The model was fine-tuned using 239 Indonesian context-question-answer triples, with 90% used for training and 10% for testing. Fine-tuning was conducted using four dataset sizes: 50, 100, 150, and 239 samples, to examine the relationship between training data volume and generation quality. Model performance was evaluated by examining parameter changes between the pre-trained and fine-tuned models and using the ROUGE metric to measure textual similarity between generated and reference questions and answers. The 239-sample dataset produced the most coherent and contextually appropriate questions, with ROUGE-1, ROUGE-2, and ROUGE-L scores of 1.0 for questions and 0.41, 0.24, and 0.36 for answers. These findings suggest that GPT-2 fine-tuning can support automatic question-generation tools when sufficient training data are available.
Copyrights © 2026