Claim Missing Document
Check
Articles

Found 2 Documents
Search

Classification of Pneumonia Using CNN and Vision Transformer Ma`dan Shomsomi; Widhaksa Triawan; Purwadi
Journal of Artificial Intelligence and Engineering Applications (JAIEA) Vol. 5 No. 2 (2026): February 2026
Publisher : Yayasan Kita Menulis

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.59934/jaiea.v5i2.1906

Abstract

Pneumonia remains one of the leading causes of mortality among children worldwide. This study aims to evaluate the performance of two deep learning architectures, Convolutional Neural Network (CNN) and Vision Transformer (ViT), for pneumonia classification using chest X-ray images. Four training scenarios were examined, consisting of MobileNetV2 baseline, MobileNetV2 fine-tuned, ViT baseline, and ViT fine-tuned models. The dataset was obtained from the Chest X-Ray Images (Pneumonia) collection and was processed through augmentation and preprocessing to produce a balanced set of 9,000 images. Baseline models were trained using a feature extraction approach, while fine-tuning was conducted by selectively unfreezing internal layers. Experimental results show that all models achieved accuracy above 95%. The MobileNetV2 baseline reached 97.63%, while its fine-tuned counterpart did not yield further improvement, achieving 97.41%. In contrast, the Vision Transformer demonstrated substantial performance gains, where partial fine-tuning produced the highest accuracy of 98.59% with an f1-score of 0.99. These findings indicate that ViT with targeted fine-tuning is more effective in capturing global representations within X-ray images, making it a strong candidate for computer-aided pneumonia detection systems supported by artificial intelligence.
Comparative Evaluation of MobileNetV2 and Vision Transformer for Pneumonia Classification on Chest X-Ray Images: Fine-Tuning, Focal Loss, and Decision-Threshold Analysis Hendra Marcos; Ali Novian; Ma`dan Ma`dan Shomsomi; Widhaksa Triawan
Indonesian Journal of Computer Science and Engineering Vol. 3 No. 01 (2026): IJCSE Volume 03 Number 01, May 2026
Publisher : CV. Cendekiawan Muda Sriwijaya

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.70656/ijcse.v3i01.857

Abstract

Pneumonia remains a major respiratory infection that imposes a substantial health burden and requires rapid and accurate detection. This study evaluates five deep-learning configurations for binary classification of chest X-ray (CXR) images into Normal and Pneumonia classes: MobileNetV2 with a frozen backbone, fine-tuned MobileNetV2, fine-tuned MobileNetV2 with Focal Loss, Vision Transformer (ViT) without full fine-tuning, and fine-tuned ViT. The public dataset used in the experiments contains 5,856 images, comprising 5,216 training images, 16 validation images, and 624 test images. All metrics in this revised version were recalculated consistently from the experimental confusion matrices. ViT without full fine-tuning achieved the best overall performance at a threshold of 0.5, with 92.63% accuracy, 98.72% sensitivity, 82.48% specificity, a 94.36% F1-score, 90.60% balanced accuracy, and a Matthews correlation coefficient (MCC) of 0.845. MobileNetV2 with Focal Loss achieved 92.15% accuracy, 98.97% sensitivity, 80.77% specificity, and a 94.03% F1-score, providing slightly higher sensitivity with only four false-negative cases. In contrast, fine-tuned ViT achieved 100% sensitivity but only 57.69% specificity at the 0.5 threshold, indicating a shift in the predicted probability distribution and imbalanced generalization. Post hoc threshold analysis showed that changing the decision threshold can improve the error trade-off; however, it was not used to select the primary model because the threshold sweep was evaluated on the test set. These findings demonstrate that fine-tuning strategy and loss function affect model error characteristics differently, while screening-oriented evaluation should consider sensitivity and specificity together with accuracy