Down Syndrome is a multisystem genetic disorder serving as a primary cause of intellectual disability, where delayed diagnosis often obstructs access to crucial developmental interventions. Conventional clinical diagnosis, which relies on the subjective observation of dysmorphic features, is frequently constrained by the scarcity of medical experts, necessitating the development of automated systems based on facial images to support objective and accessible early detection. This study aims to evaluate and compare the performance of single deep learning models with an innovative hybrid architecture to enhance the accuracy of screening systems. The research methodology employs a dataset of 2,666 facial images of toddlers, processed using Convolutional Neural Networks (CNN) through a transfer learning approach. Comprehensive experiments compared InceptionV3 and EfficientNetB3 architectures both as standalone models and within a hybrid ensemble while assessing the efficacy of feature extraction versus fine-tuning strategies. The results demonstrate that fine-tuning significantly outperforms feature extraction, yielding a 10-12% performance increase due to more specific feature adaptation. The hybrid ensemble model utilizing fine-tuning emerged as the superior approach, achieving a peak validation accuracy of 92.32% and an ghF1-Score of 92.33%. This model proved robust against pose and expression variations while effectively minimizing false negatives. Consequently, integrating computational strengths through a hybrid architecture produces rich feature representations, establishing this method as a reliable and precise solution for medical screening.
Copyrights © 2026