Cardiovascular diseases require early and reliable screening because manual auscultation may be affected by noise, subjective interpretation, and inter-observer variability. This study aimed to develop an STFT-based deep learning framework for multiclass phonocardiogram (PCG) classification. The proposed framework was designed to provide a reproducible evaluation procedure by combining standardized preprocessing, time–frequency feature extraction, and deep learning-based classification under the same experimental conditions. Unlike approaches that may evaluate segmented signals without clearly preserving recording-level separation, this study emphasizes a leakage-free splitting strategy to reduce the risk of overestimated performance and to provide a more reliable assessment of model generalization. The Yaseen PCG dataset, consisting of 1000 recordings from five classes (AS, MR, MS, MVP, and Normal), was divided using a leakage-free recording-level split before segmentation and spectrogram generation. After preprocessing, 2-second PCG segments with 50% overlap were converted into 128 × 128 STFT spectrograms and classified using BiLSTM and CNN-BiLSTM models. Both models were trained and tested using the same dataset split, preprocessing pipeline, and evaluation metrics, including accuracy, precision, recall, F1-score, specificity, and confusion matrices. The BiLSTM model achieved 92.36% accuracy in the final independent test run, while the CNN-BiLSTM model achieved 95.83%. Across three repeated runs, BiLSTM achieved 93.85% ± 1.47%, whereas CNN-BiLSTM achieved 95.94% ± 0.71%. These results show that CNN-BiLSTM provides higher and more stable classification performance for five-class PCG classification, while BiLSTM remains a simpler alternative for lightweight implementation. Overall, the proposed STFT-based framework provides a reliable approach for automated heart sound classification and may support future computer-aided cardiac screening applications.
Copyrights © 2026