Face detection and tracking in crowded environments remain challenging due to occlusion, object overlap, and high visual similarity between individuals. In tracking-by-detection systems, detection quality plays a crucial role in tracking stability, yet the relationship between detection performance and identity consistency is not fully explored. This study proposes an integrated framework combining YOLOv9, an attention mechanism, and DeepSORT to enhance feature representation and improve identity tracking in dense environments, where the attention mechanism is embedded in the detection stage to strengthen feature discriminability and enable more stable identity association across frames. The system is evaluated using three dataset partitioning scenarios (90:10, 50:50, and 10:90) to analyze the impact of training data distribution. Experimental results show that the 90:10 configuration achieves the best performance, with precision 0.9398, recall 0.8869, F1-score 0.9126, MOTA 71.0%, and IDF1 73.0%. These findings confirm that improved feature representation significantly enhances detection quality and tracking stability, and demonstrate that feature stability is more critical than adopting newer detection architectures for achieving robust tracking performance in crowded environments.
Copyrights © 2026