Claim Missing Document
Check
Articles

Found 1 Documents
Search

Robust Deep Learning Approach for Automatic Age and Gender Recognition Based on Voice M. Nasyid Yunitian Rizal; Theopilus Bayu Sasongko; Arifiyanto Hadinegoro; Kumara Ari Yuana; Wahid Miftahul Ashari
Intechno Journal : Information Technology Journal Vol. 8 No. 1 (2026): July
Publisher : Universitas Amikom Yogyakarta

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.24076/intechnojournal.2026v8i1.2827

Abstract

Voice-based age and gender recognition plays an important role in biometric authentication, personalized human–computer interaction, and digital forensic applications. However, existing single deep learning architectures often struggle to simultaneously capture local acoustic patterns, temporal dependencies, and long-range contextual information from speech signals. This study proposes a hybrid CNN–BiLSTM–Transformer framework to improve the accuracy and robustness of multi-class age and gender classification. The proposed approach employs Mel-spectrogram representations generated from the Mozilla Common Voice dataset, followed by audio standardization and feature extraction. A Convolutional Neural Network (CNN) enhanced with Squeeze-and-Excitation blocks extracts discriminative spectral features, a Bidirectional Long Short-Term Memory (Bi-LSTM) network models bidirectional temporal dependencies, and a Transformer Encoder captures global contextual relationships through a multi-head self-attention mechanism. The model was evaluated on 40,392 speech samples across 12 age–gender categories. Experimental results achieved an overall classification accuracy of 91%, outperforming standalone CNN and Bi-LSTM models, which obtained accuracies of 84% and 74%, respectively. In addition, the proposed model demonstrated balanced performance with macro-average precision, recall, and F1-scores of 0.91, 0.92, and 0.91. The novelty of this research lies in the integration of complementary spatial, temporal, and global attention mechanisms within a unified architecture for large-scale multi-class voice-based demographic classification, providing an effective and scalable solution for intelligent biometric and speech analysis systems.