Automated bioacoustic monitoring is increasingly important for biodiversity observation, yet practical deployment is often limited by computational constraints and the scarcity of annotated audio data. This study introduces MelCNN, a custom compact end-to-end architecture designed to mitigate over-parameterization and negative transfer in animal sound classification under small-data conditions. A balanced dataset of 600 audio clips from three animal classes (Cat, Dog, and Cow) was standardized to 4.8 seconds, converted into Log-Mel spectrograms, and evaluated using 5-fold stratified cross-validation. The proposed Lightweight CNN with Global Average Pooling was compared against YAMNet-based transfer learning baselines and three additional modern lightweight CNN baselines, namely MobileNetV3-Small, ShuffleNetV2-1.0x, and EfficientNet-Lite0-style. The proposed model achieved 97.00% ± 2.01% Accuracy and 97.01% ± 2.00% Macro-F1 while retaining a compact parameter size of approximately 0.39 million, outperforming both the strongest transfer learning baselines and the added modern lightweight baselines in the present setting. The model also maintained performance above 90% at 12 kHz and preserved approximately 90.5% Macro-F1 when trained with only 30% of the available labeled data. These findings indicate that a domain-specific lightweight architecture can provide a favorable accuracy-efficiency trade-off for controlled small-data animal sound classification in resource-constrained monitoring scenarios.
Copyrights © 2026