The development of Large Language Models (LLMs) has opened new opportunities for the automatic generation of math word problems (MWPs). However, many existing approaches still produce repetitive and template-based problems due to limited variation in context, narrative structure, and semantic relationships. This limitation reduces the effectiveness of such problems in assessing students’ conceptual understanding. This study aims to develop a semantic diversity pipeline based on context-aware generation to produce more varied, meaningful, and curriculum-aligned math word problems for Indonesian elementary education. The proposed method involves building an Indonesian MWP dataset, fine-tuning an LLM using Low-Rank Adaptation (LoRA), and designing a generation pipeline consisting of context retrieval, prompt diversification, semantic evaluation, and a solvability filtering mechanism. Evaluation was conducted using automated metrics, including Self-BLEU, Jaccard Similarity, and Cosine Similarity based on Sentence-BERT, as well as qualitative assessments from mathematics teachers. The results show that the proposed approach successfully improves lexical, semantic, and contextual diversity in generated problems. The Self-BLEU score of 4.79 indicates low repetition, the Jaccard Similarity score of 0.194 reflects high vocabulary variation, and the Cosine Similarity score of 0.424 demonstrates balanced semantic diversity while maintaining mathematical consistency. Teacher evaluations further confirm that the generated problems are relevant to the curriculum, appropriately challenging, and more natural compared to conventional methods. Overall, this research contributes to the development of more adaptive and diverse LLM-based math word problem generation systems for mathematics learning in Indonesia.
Copyrights © 2026