Audio quality can be degraded by noise contamination, while conventional noise reduction methods such as the Wiener Filter may have limitations in preserving audio spectral characteristics. This study develops an audio noise reduction system using a U-Net Convolutional Neural Network (CNN) based on Short-Time Fourier Transform (STFT) spectrograms to reduce white noise. The dataset consisted of seven accompaniment audio files from the Indonesian Robot Dance Art Contest (KRSTI), with five used for training and two for testing. Clean audio was contaminated with white noise at SNR levels of 0, 5, 10, and 15 dB and transformed into STFT spectrograms, while phase information was retained for reconstruction. The U-Net was trained to learn the mapping between noisy and clean spectrograms and compared with the Wiener Filter as a conventional baseline using SNR, MSE, MAE, LSD, and SI-SDR. The results show that U-Net achieves lower average MSE (0.002453), MAE (0.032514), and LSD (9.663 dB), while the Wiener Filter achieves higher average ΔSNR (3.327 dB) and SI-SDR (12.251 dB) than U-Net, which achieves 1.915 dB and 10.073 dB, respectively. These findings indicate that the two methods exhibit different strengths depending on the evaluation metric. Further research should investigate larger and more diverse datasets, particularly real-world noise conditions, to improve model robustness and generalization.
Copyrights © 2026