Karam Abdullah
Department of Computer Science, College of Education for Pure Sciences, University of Mosul, Iraq

Published : 1 Documents Claim Missing Document
Claim Missing Document
Check
Articles

Found 1 Documents
Search

Adaptive Mix-Gated Multi-Attention Fusion in YOLOv11s for Dense Person Detection: A Systematic Evaluation of Dual and Triple Attention Configurations Aya Ahmed; Karam Abdullah
SISTEMASI Vol 15, No 8 (2026): Sistemasi: Jurnal Sistem Informasi
Publisher : Universitas Islam Indragiri

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.32520/stmsi.v15i8.6733

Abstract

Person detection in a crowded environment is very challenging due to excessive occlusion, scale variation, and heavy overlap between individuals. This paper proposes a lightweight Mix-Gated attention framework to enhance the feature representation capacity of YOLOv11s for dense person detection. The proposed framework incorporates two complementary channel-wise and spatial attention mechanisms under a flexible channel-wise gating strategy, while preserving the original YOLOv11s architecture and real-time inference ability. Unlike previous studies, which primarily employed either a single attention mechanism or fixed combinations of multiple attention modules, the proposed framework introduces an adaptive Mix-Gated feature selection mechanism that dynamically learns the contribution of complementary attention modules through trainable channel-wise gating. Furthermore, this study systematically evaluates ten dual and triple attention configurations under identical training and evaluation conditions, providing a comprehensive and fair comparison of their effectiveness. this study evaluated the performance on four popular benchmark datasets - CrowdHuman, WiderPerson, MOT17Det and MOT20Det - using Precision, Recall, F1-score, mAP@50 and mAP@50-95 as evaluation metrics. The experimental results show that all proposed Mix-Gated configurations outperform the baseline YOLOv11s model. Mix-Gated-CBAM+SE+SA has the best detection accuracy and efficient inference performance among Mix-Gated configurations. These results indicate that adaptive fusion of complementary attention mechanisms can improve the feature representation and person detection performance in crowded scenes while not sacrificing computational efficiency.