Haikal Abror
Informatic and Computer Engineering Education Study Program, Faculty of Engineering, Universitas Negeri Semarang, Indonesia

Published : 1 Documents Claim Missing Document
Claim Missing Document
Check
Articles

Found 1 Documents
Search

Assessing Generative AI with Context-Augmented Zero-Shot Prompting for HOTS Question Generation Aligned with Bloom’s Taxonomy Saiful Ridlo; Ahmad Sehabuddin; Syahroni Hidayat; Taofan Ali Achmadi; Uswatun Hasanah; Indah Indi Afifah; Haikal Abror
Edu Komputika Journal Vol. 12 No. 2 (2025): Edu Komputika Journal
Publisher : Universitas Negeri Semarang

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.15294/edukom.v12i2.34437

Abstract

This study investigates the use of HOTS-GenAI, a generative AI system employing context-augmented zero-shot prompting, to automatically generate multiple-choice questions aligned with higher-order thinking skills (HOTS) in Bloom’s taxonomy.  A dataset of 200 items for vocational high schools was validated by three experts. The ground truth data demonstrated good quality with an inter-rater reliability of 0.75 (Gregory’s Index). System performance across analysis (C4), evaluation (C5), and creation (C6) levels was evaluated using Content Validity Index (CVI), gap analysis, and confusion-matrix-based metrics. The findings revealed that HOTS-GenAI performed relatively well at the analytical level, where 70% of items met the HOTS threshold, supported by higher expert consensus. However, only 10% of items achieved the threshold for evaluation, and none for creation. CVI results indicated moderate validity overall, with stronger agreement for C4 than for C5 or C6. Confusion matrix analysis further confirmed this imbalance: accuracy and F1-scores were highest for analysis items but dropped sharply for evaluation and creation, where recall and precision were near zero. These results suggest that while HOTS-GenAI has potential in generating analytical questions, its capacity to model evaluative and creative tasks remains underdeveloped. Future research should involve larger datasets, refined prompt design, and more operational rubrics to enhance both validity and reliability in AI-generated HOTS assessments.