Rivaldo Nugraha
STMIK AMIKBANDUNG

Published : 1 Documents Claim Missing Document
Claim Missing Document
Check
Articles

Found 1 Documents
Search

Comparative Evaluation of ChatGPT, Gemini, and Claude Response Quality on Programming Questions Based on Bloom's Taxonomy Kahfi Bintang; Median Zikri; Rivaldo Nugraha; Dhanny Karewur
Journal of Technology and System Information Vol. 3 No. 3 (2026): July
Publisher : Indonesian Journal Publisher

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.47134/jtsi.v3i3.6354

Abstract

The rapid advancement of Large Language Models (LLMs) has significantly influenced programming education and software development by providing automated code generation and problem-solving assistance. However, differences in the quality of responses generated by ChatGPT, Gemini, and Claude require a comprehensive evaluation using multidimensional assessment criteria. This study aims to compare the quality of responses produced by these three LLMs on programming questions based on Bloom's Revised Taxonomy and the ACM/IEEE Computing Curricula assessment rubric. A quantitative comparative approach was employed using 40 programming questions distributed across six Bloom cognitive levels (C1–C6). Each response was evaluated using five indicators: accuracy, completeness, clarity, logical reasoning, and efficiency. The collected data were analyzed using descriptive statistical methods and mean comparison. The results indicate that ChatGPT achieved the highest overall mean score of 21, followed by Claude (20) and Gemini (19). ChatGPT also obtained the highest scores in accuracy (4.3), clarity (4.1), logical reasoning (4.4), and efficiency (4.5), while ChatGPT and Claude shared the highest completeness score (4.1). Furthermore, performance across all models declined as the cognitive level increased from Remember (C1) to Create (C6), although ChatGPT demonstrated the most consistent performance across all Bloom levels. These findings indicate that ChatGPT provides superior and more consistent programming solutions and that integrating Bloom's Taxonomy with multidimensional evaluation criteria offers a comprehensive framework for assessing the quality of LLM-generated programming responses.