This study investigates the use of HOTS-GenAI, a generative AI system employing context-augmented zero-shot prompting, to automatically generate multiple-choice questions aligned with higher-order thinking skills (HOTS) in Bloom’s taxonomy. A dataset of 200 items for vocational high schools was validated by three experts. The ground truth data demonstrated good quality with an inter-rater reliability of 0.75 (Gregory’s Index). System performance across analysis (C4), evaluation (C5), and creation (C6) levels was evaluated using Content Validity Index (CVI), gap analysis, and confusion-matrix-based metrics. The findings revealed that HOTS-GenAI performed relatively well at the analytical level, where 70% of items met the HOTS threshold, supported by higher expert consensus. However, only 10% of items achieved the threshold for evaluation, and none for creation. CVI results indicated moderate validity overall, with stronger agreement for C4 than for C5 or C6. Confusion matrix analysis further confirmed this imbalance: accuracy and F1-scores were highest for analysis items but dropped sharply for evaluation and creation, where recall and precision were near zero. These results suggest that while HOTS-GenAI has potential in generating analytical questions, its capacity to model evaluative and creative tasks remains underdeveloped. Future research should involve larger datasets, refined prompt design, and more operational rubrics to enhance both validity and reliability in AI-generated HOTS assessments.
Copyrights © 2025