As voice biometric systems face escalating threats from deepfake audio, securing non-English languages remains a critical vulnerability. This exploratory study aims to develop a localized Indonesian-language spoofing detection baseline, acknowledging constraints of a small, imbalanced dataset. We propose an approach combining AASIST spectro-temporal graph attention networks with a SAMO multi-center one-class learning loss function. Initial evaluations revealed substantial class overlap without acoustic perturbation, yielding a baseline equal error rate of 48.9%, highlighting under-resourced language vulnerabilities. However, an ablation study injecting controlled Gaussian noise acted as an optimal stochastic regularizer, significantly clarifying the decision boundaries between genuine and spoofed speech. This minor perturbation reduced the equal error rate to 44.05%, whereas excessive noise predictably destroyed the fundamental acoustic structure. These findings demonstrate that regularized multi-center geometries isolate synthetic artifacts, establishing a foundational proof-of-concept and highlighting the need for expanded corpora when securing Indonesian voice authentication infrastructures against advanced generative attacks.
Copyrights © 2026