Journal of Defense Technology and Engineering
Vol. 2 No. 1 (2026): July, Journal of Defense Technology and Engineering

An explainable deep learning framework for autonomous drone-based military object detection using vision transformers

Jontinus Manullang (IAKN Tarutung, Tarutung, Indonesia)
R. Fanry Siahaan (Universitas HKBP Nommensen, Medan, Indonesia)
Sutrisno Situmorang (Akademi Informatika Medicom Medan, Medan, Indonesia)
Jonson Manurung (Universitas Pertahanan Republik Indonesia, Bogor, Indonesia)



Article Info

Publish Date
30 Jul 2026

Abstract

The rapid expansion of overhead imagery and unmanned aerial sensing has intensified the need for accurate, interpretable, and computationally adaptable object detection models in security-sensitive remote-sensing environments. The main research problem addressed in this study is the limited transparency of deep object detection models when identifying small, dense, and visually ambiguous objects in aerial or satellite imagery. This study aims to design an explainable deep learning framework that combines Vision Transformer representation learning, transfer learning, and Grad-CAM-based visual explanation for autonomous drone-based military object detection using the xView dataset. The proposed research process includes dataset preparation, image tiling, bounding-box conversion, normalization, transfer learning from an ImageNet-pretrained ViT backbone, transformer-based feature extraction, Grad-CAM visualization, and evaluation using accuracy, precision, recall, F1-score, IoU, AP, and mAP. The dataset used in this manuscript is the xView dataset available through a Kaggle mirror and originally introduced as a large-scale overhead imagery benchmark containing more than one million labeled object instances across 60 categories. Experimental results indicate that the proposed ViT + Grad-CAM framework achieves 0.823 accuracy, 0.814 macro precision, 0.806 macro recall, 0.810 macro F1-score, and 0.802 mAP@0.5 on a five-class defense-relevant subset, outperforming SSD, Faster R-CNN, and YOLOv8 baselines in the reported comparison. The study concludes that combining transformer-based global context modeling with explainability can improve both detection performance and model auditability in overhead imagery, although real deployment requires full-dataset validation, edge-device optimization, robustness testing, and strict human oversight for ethical use.

Copyrights © 2026