Aldi
Universitas Sembilanbelas November Kolaka

Published : 1 Documents Claim Missing Document
Claim Missing Document
Check
Articles

Found 1 Documents
Search

Hybrid CNN-RNN Architecture with MFCC-LFCC Features for Audio Deepfake Detection Muh. Hajar Akbar; Nurfitria Ningsi; Aldi; Muhammad Na’im Al Jum’ah; Ilcham
Journal of Information System and Informatics Vol 8 No 4 (2026): August
Publisher : Asosiasi Doktor Sistem Informasi Indonesia

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.63158/journalisi.v8i4.1723

Abstract

The proliferation of sophisticated audio deepfake technology poses a significant threat to digital voice authentication and forensic verification systems. This research addresses this challenge by developing and evaluating a lightweight hybrid Convolutional Neural Network and Recurrent Neural Network (CNN-RNN) architecture for audio deepfake detection. The proposed model integrates a CNN for spatial feature extraction with a bidirectional RNN for temporal dependency modeling, utilizing an early vertically fused feature set of Mel-Frequency Cepstral Coefficients (MFCC) and Linear Frequency Cepstral Coefficients (LFCC) stabilized via utterance-level Cepstral Mean and Variance Normalization (CMVN). We assessed the proposed framework on the official ASVspoof 2019 Logical Access (LA) evaluation benchmark dataset (71,237 trials). Comprehensive evaluation on the official evaluation set demonstrated promising performance within the evaluated benchmark, achieving a Global Equal Error Rate (EER) of 8.24% and an Area Under the Receiver Operating Characteristic (ROC-AUC) of 0.9680, while maintaining an internal validation EER of 0.33% on known attacks. While showing high sensitivity to bona fide speech and robust resilience against advanced Neural Text-to-Speech synthesis (EER < 0.25% for A07–A10), the framework exhibits notable vulnerability to phase-preserving voice conversion attacks. Consequently, without real-world forensic operational testing, the model serves as an initial diagnostic screening approach rather than a fully operational forensic solution, and still requires further cross-repository and noisy-condition validation.