Muhamad Fakhri Khairil Imam
Universitas Muhammadiyah Sukabumi

Published : 1 Documents Claim Missing Document
Claim Missing Document
Check
Articles

Found 1 Documents
Search

IMPLEMENTASI SPEECH EMOTION RECOGNITION UNTUK KLASIFIKASI TINGKAT STRES DARI VOICE NOTE BERBASIS HYBRID CNN-LSTM Muhamad Fakhri Khairil Imam
Jurnal Informatika dan Teknik Elektro Terapan Vol. 14 No. 3 (2026)
Publisher : Universitas Lampung

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.23960/jitet.v14i3.10900

Abstract

Kesehatan mental menjadi isu penting seiring meningkatnya tekanan psikologis akibat gaya hidup digital, namun deteksi stres masih mengandalkan instrumen self-report yang bersifat terjadwal dan subjektif. Penelitian ini mengembangkan sistem Speech Emotion Recognition (SER) berbasis voice note berbahasa Indonesia untuk mengklasifikasikan tingkat stres menggunakan pendekatan hybrid CNN-LSTM (akustik) dan Zero-Shot Classification IndoBERT (linguistik) dengan metodologi CRISP-DM. Model dilatih menggunakan dataset gabungan RAVDESS (1.440 file, pembagian speaker-independent) dan voice note primer berbahasa Indonesia (106 file) yang dipetakan ke tiga kategori Non_Stress, Mild_Stress, dan High_Stress, menghasilkan 1.546 sampel asli. Fitur akustik diekstraksi menggunakan 40 koefisien MFCC dengan panjang tetap 174 frame. Pengujian pada 256 sampel independen menghasilkan akurasi model CNN-LSTM akustik sebesar 62,11% dengan F1-score terbaik pada kelas High_Stress (0,69), sementara kelas Mild_Stress menunjukkan performa terendah (F1 0,41) akibat karakteristik akustik yang ambigu. Untuk mengatasi kesenjangan bahasa antara dataset RAVDESS dan voice note Indonesia, sistem menerapkan fusi hibrida dengan bobot linguistik 90% dan akustik 10%. Model diintegrasikan ke aplikasi web StressVoice berbasis Flask yang dilengkapi transkripsi Whisper AI, respons empatik, cetak surat rujukan PDF, dan integrasi konsultasi WhatsApp ke psikolog. Mental health has become a critical issue amid rising psychological pressure from digital lifestyles, yet stress detection still relies on scheduled, subjective self-report instruments. This study develops a Speech Emotion Recognition (SER) system based on Indonesian voice notes to classify stress levels using a hybrid CNN-LSTM (acoustic) and IndoBERT Zero-Shot Classification (linguistic) approach under the CRISP-DM methodology. The model was trained on a combined RAVDESS dataset (1,440 files, speaker-independent split) and primary Indonesian voice notes (106 files) mapped into three categories, Non_Stress, Mild_Stress, and High_Stress, yielding 1,546 original samples. Acoustic features were extracted using 40 MFCC coefficients with a fixed length of 174 frames. Testing on 256 independent samples produced an acoustic CNN-LSTM accuracy of 62.11%, with the best F1-score on the High_Stress class (0.69), while Mild_Stress showed the lowest performance (F1 0.41) due to ambiguous acoustic characteristics. To bridge the language gap between RAVDESS and Indonesian voice notes, the system applies hybrid fusion weighting linguistic analysis at 90% and acoustic analysis at 10%. The model is integrated into the Flask-based StressVoice web application, equipped with Whisper AI transcription, empathetic responses.