Claim Missing Document
Check
Articles

Found 12 Documents
Search

Ekstraksi Fitur Pengenalan Emosi Berdasarkan Ucapan Menggunakan Linear Predictor Ceptral Coeffecient Dan Mel Frequency Cepstrum Coefficients Helmiyah, Siti; Riadi, Imam; Umar, Rusydi; Hanif, Abdullah
Mobile and Forensics Vol 1, No 2 (2019)
Publisher : Universitas Ahmad Dahlan

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.12928/mf.v1i2.1259

Abstract

Ucapan suara memiliki informasi penting yang dapat diterima oleh otak melalui gelombang suara. Otak menerima gelombang suara melalui alat pendengaran dan menghasilkan suatu informasi berupa pesan, bahasa, dan emosi. Pengenalan emosi wicara merupakan teknologi yang dirancang untuk mengidentifikasi keadaan emosi seseorang dari sinyal ucapannya. Hal tersebut menarik untuk diteliti, karena berkaitan dengan teknologi zaman sekarang yaitu pada penggunaan smartphone di berbagai macam aktivitas sehari-hari. Penelitian ini membandingkan ekstraksi fitur Metode LPC dan Metode MFCC. Kedua metode ekstraksi tersebut diklasifikasi menggunakan Metode Jaringan Syaraf Tiruan (MLP) untuk pengenalan emosi. Masing-masing metode menggunakan data emosi marah, bosan, bahagia, netral, dan sedih. Data dibagi menjadi dua, yaitu data testing dan data data training dengan perbandingan 80:20. Arsitektur jaringan yang digunakan adalah tiga lapisan yaitu lapisan input, lapisan tersembunyi, dan lapisan output. Parameter MLP yang digunakan learning rate = 0.0001, epsilon = 1e-08, epoch = 500, dan Cross Validation = 5. Hasil akurasi pengenalan emosi dengan ekstraksi fitur LPC sebesar adalah 28%. Sedangkan hasil akurasi dengan ekstraksi fitur MFCC sebesar 61,33%. Hasil akurasi ini bisa ditingkatkan dengan menambahkan data yang lebih banyak lagi, terutama untuk data testing. Perlunya pengujian pada nilai parameter jaringan MLP, yaitu dengan mengubah nilai-nilai parameter, karena dapat mempengaruhi tingkat akurasi pengenalan. Selain itu penentuan ekstraksi fitur dan klasifikasi metode yang lain juga dapat digunakan untuk mencari nilai akurasi pengenalan emosi yang lebih baik lagi.
Speech Classification to Recognize Emotion Using Artificial Neural Network Helmiyah, Siti; Riadi, Imam; Umar, Rusydi; Hanif, Abdullah
Khazanah Informatika Vol. 7 No. 1 April 2021
Publisher : Universitas Muhammadiyah Surakarta

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.23917/khif.v7i1.11913

Abstract

This study seeks to identify human emotions using artificial neural networks. Emotions are difficult to understand and hard to measure quantitatively. Emotions may be reflected in facial expressions and voice tone. Voice contains unique physical properties for every speaker. Everyone has different timbres, pitch, tempo, and rhythm. The geographical living area may affect how someone pronounces words and reveals certain emotions. The identification of human emotions is useful in the field of human-computer interaction. It helps develop the interface of software that is applicable in community service centers, banks, education, and others. This research proceeds in three stages, namely data collection, feature extraction, and classification. We obtain data in the form of audio files from the Berlin Emo-DB database. The files contain human voices that express five sets of emotions: angry, bored, happy, neutral, and sad. Feature extraction applies to all audio files using the method of Mel Frequency Cepstrum Coefficient (MFCC). The classification uses Multi-Layer Perceptron (MLP), which is one of the artificial neural network methods. The MLP classification proceeds in two stages, namely the training and the testing phase. MLP classification results in good emotion recognition. Classification using 100 hidden layer nodes gives an average accuracy of 72.80%, an average precision of 68.64%, an average recall of 69.40%, and an average F1-score of 67.44%.This study seeks to identify human emotions using artificial neural networks. Emotions are difficult to understand and hard to measure quantitatively. Emotions may be reflected in facial expressions and voice tone. Voice contains unique physical properties for every speaker. Everyone has different timbres, pitch, tempo, and rhythm. The geographical living area may affect how someone pronounces words and reveals certain emotions. The identification of human emotions is useful in the field of human-computer interaction. It helps develop the interface of software that is applicable in community service centres, banks, and education and others. This research proceeds in three stages, namely data collection, feature extraction, and classification. We obtain data in the form of audio files from the Berlin Emo-DB database. The files contain human voices that express five sets of emotions: angry, bored, happy, neutral and sad. Feature extraction applies to all audio files using the method of Mel Frequency Cepstrum Coefficient (MFCC). The classification uses Multi-Layer Perceptron (MLP), which is one of the artificial neural network methods. The MLP classification proceeds in two stages, namely the training and the testing phase. MLP classification results in good emotion recognition. Classification using 100 hidden layer nodes gives an average accuracy of 72.80%, an average precision of 68.64%, an average recall of 69.40%, and an average F1-score of 67.44%.
Ekstraksi Fitur Pengenalan Emosi Berdasarkan Ucapan Menggunakan Linear Predictor Ceptral Coeffecient Dan Mel Frequency Cepstrum Coefficients Helmiyah, Siti; Riadi, Imam; Umar, Rusydi; Hanif, Abdullah
Mobile and Forensics Vol. 1 No. 2 (2019)
Publisher : Universitas Ahmad Dahlan

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.12928/mf.v1i2.1259

Abstract

Ucapan suara memiliki informasi penting yang dapat diterima oleh otak melalui gelombang suara. Otak menerima gelombang suara melalui alat pendengaran dan menghasilkan suatu informasi berupa pesan, bahasa, dan emosi. Pengenalan emosi wicara merupakan teknologi yang dirancang untuk mengidentifikasi keadaan emosi seseorang dari sinyal ucapannya. Hal tersebut menarik untuk diteliti, karena berkaitan dengan teknologi zaman sekarang yaitu pada penggunaan smartphone di berbagai macam aktivitas sehari-hari. Penelitian ini membandingkan ekstraksi fitur Metode LPC dan Metode MFCC. Kedua metode ekstraksi tersebut diklasifikasi menggunakan Metode Jaringan Syaraf Tiruan (MLP) untuk pengenalan emosi. Masing-masing metode menggunakan data emosi marah, bosan, bahagia, netral, dan sedih. Data dibagi menjadi dua, yaitu data testing dan data data training dengan perbandingan 80:20. Arsitektur jaringan yang digunakan adalah tiga lapisan yaitu lapisan input, lapisan tersembunyi, dan lapisan output. Parameter MLP yang digunakan learning rate = 0.0001, epsilon = 1e-08, epoch = 500, dan Cross Validation = 5. Hasil akurasi pengenalan emosi dengan ekstraksi fitur LPC sebesar adalah 28%. Sedangkan hasil akurasi dengan ekstraksi fitur MFCC sebesar 61,33%. Hasil akurasi ini bisa ditingkatkan dengan menambahkan data yang lebih banyak lagi, terutama untuk data testing. Perlunya pengujian pada nilai parameter jaringan MLP, yaitu dengan mengubah nilai-nilai parameter, karena dapat mempengaruhi tingkat akurasi pengenalan. Selain itu penentuan ekstraksi fitur dan klasifikasi metode yang lain juga dapat digunakan untuk mencari nilai akurasi pengenalan emosi yang lebih baik lagi.
EKSTRAKSI CIRI EMOSI MANUSIA BERDASARKAN UCAPAN MENGGUNAKAN MEL-FREQUENCY CEPSTRAL COEFFICIENTS (MFCC) Siti Helmiyah; Abdul Fadlil; Anton Yudhana
Prosiding SNST Fakultas Teknik Vol 1, No 1 (2018): PROSIDING SEMINAR NASIONAL SAINS DAN TEKNOLOGI 9 2018
Publisher : Prosiding SNST Fakultas Teknik

Show Abstract | Download Original | Original Source | Check in Google Scholar | Full PDF (293.42 KB)

Abstract

Emosi merupakan perilaku manusia yang dapat diungkapkan dengan tingkah laku berupa raut wajah dan suara. Suara adalah suatu gelombang longitudinal yang merambat di udara. Pada kehidupan sehari-hari manusia berkomunikasi dengan menggunakan suara. Baik itu di dunia pendidikan, kantor, dan tempat umum lainnya. Ketika berkomunikasi seringkali seseorang tidak sadar dengan emosi yang sedang dirasakannya. Pengenalan emosi manusia berdasarkan suara ini merupakan permasalahan yang sulit untuk dipecahkan. Hal ini karena penilaian terhadap emosi sulit dikenali oleh peneliti secara objektif. Maka dari itu penelitian ini dilakukan bertujuan untuk mengetahui pengenalan dan klasifikasi emosi berdasarkan suara. Data diambil dari database Berlin Database of Emotional Speech. Data tersebut menggunakan suara aktor yang sudah dibedakan jenis kelamin, umur, emosinya serta kalimatnya. Ekstraksi ciri yang digunakan untuk pengenalan suara adalah Mel-Frequency Cepstral Coefficients (MFCC). Hasil pola emosi dari ucapan dengan frame bloking = 4, pre-emphasis dengan alpha =0.95, filterbank = 20, dan koefesien cepstral = 13 menunjukkan bahwa dengan kalimat yang sama, aktor, jenis kelamin, suara, dan umur yang berbeda tidak begitu terlihat banyak perbedaan polanya. Maka perlu dilakukan proses prepocessing untuk setiap data. Tapi tidak menutup kemungkinan hasil dari pola ini dapat dijadikan sebagai bahan untuk proses klasifikasi selanjutnya. Kata kunci : Ekstraksi Ciri, Emosi, Mel-Frequency Cepstral Coefficients (MFCC), Ucapan
Speech Classification to Recognize Emotion Using Artificial Neural Network Siti Helmiyah; Imam Riadi; Rusydi Umar; Abdullah Hanif
Khazanah Informatika Vol. 7 No. 1 April 2021
Publisher : Department of Informatics, Universitas Muhammadiyah Surakarta, Indonesia

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.23917/khif.v7i1.11913

Abstract

This study seeks to identify human emotions using artificial neural networks. Emotions are difficult to understand and hard to measure quantitatively. Emotions may be reflected in facial expressions and voice tone. Voice contains unique physical properties for every speaker. Everyone has different timbres, pitch, tempo, and rhythm. The geographical living area may affect how someone pronounces words and reveals certain emotions. The identification of human emotions is useful in the field of human-computer interaction. It helps develop the interface of software that is applicable in community service centers, banks, education, and others. This research proceeds in three stages, namely data collection, feature extraction, and classification. We obtain data in the form of audio files from the Berlin Emo-DB database. The files contain human voices that express five sets of emotions: angry, bored, happy, neutral, and sad. Feature extraction applies to all audio files using the method of Mel Frequency Cepstrum Coefficient (MFCC). The classification uses Multi-Layer Perceptron (MLP), which is one of the artificial neural network methods. The MLP classification proceeds in two stages, namely the training and the testing phase. MLP classification results in good emotion recognition. Classification using 100 hidden layer nodes gives an average accuracy of 72.80%, an average precision of 68.64%, an average recall of 69.40%, and an average F1-score of 67.44%.This study seeks to identify human emotions using artificial neural networks. Emotions are difficult to understand and hard to measure quantitatively. Emotions may be reflected in facial expressions and voice tone. Voice contains unique physical properties for every speaker. Everyone has different timbres, pitch, tempo, and rhythm. The geographical living area may affect how someone pronounces words and reveals certain emotions. The identification of human emotions is useful in the field of human-computer interaction. It helps develop the interface of software that is applicable in community service centres, banks, and education and others. This research proceeds in three stages, namely data collection, feature extraction, and classification. We obtain data in the form of audio files from the Berlin Emo-DB database. The files contain human voices that express five sets of emotions: angry, bored, happy, neutral and sad. Feature extraction applies to all audio files using the method of Mel Frequency Cepstrum Coefficient (MFCC). The classification uses Multi-Layer Perceptron (MLP), which is one of the artificial neural network methods. The MLP classification proceeds in two stages, namely the training and the testing phase. MLP classification results in good emotion recognition. Classification using 100 hidden layer nodes gives an average accuracy of 72.80%, an average precision of 68.64%, an average recall of 69.40%, and an average F1-score of 67.44%.
Perbandingan Kinerja Jaringan Syaraf Tiruan Dan Fuzzy Inference System Untuk Prediksi Prestasi Peserta Didik Siti Helmiyah; Shofwatul ‘Uyun
JATISI (Jurnal Teknik Informatika dan Sistem Informasi) Vol 4 No 1 (2017): JATISI SEPTEMBER 2017
Publisher : Lembaga Penelitian dan Pengabdian pada Masyarakat (LPPM) STMIK Global Informatika MDP

Show Abstract | Download Original | Original Source | Check in Google Scholar | Full PDF (200.823 KB) | DOI: 10.35957/jatisi.v4i1.85

Abstract

Achievement is a result of someone who excels in any field. In the educational world, achievement is often associated with academic value that serve as a reference for learners say in academic achievement. Manual data processing takes long time. It is necessary to use the achievements of predictive computing system that can helpful for the prediction process. Data taken from MAN Model Palangkaraya form of eleven subjects of UAS when MTs value and the average value of one semester report cards when the Supreme Court. For the neural network backpropagation data is normalized with small intervals are [0.1, 0.9] and for data fuzzy inference system is the original data is multiplied 10. Then do the testing using neural networks and fuzzy inference system which will compare the results obtained. Based on data have been tested, the percentage of learners' achievements prediction on back propagation neural network in the training and validation process to produce a percentage of 100% with one hidden layer architecture, the optimal parameters MSE = 0.0001, learning rate = 0 , 9, momentum = 0.4. As for the prediction of learners' achievements in the fuzzy inference system mamdani method by using S curve and bell curve (PI curve) to produce a percentage of 83.8%.
Pengenalan Pola Emosi Manusia Berdasarkan Ucapan Menggunakan Ekstraksi Fitur Mel-Frequency Cepstral Coefficients (MFCC) Siti Helmiyah; Abdul Fadlil; Anton Yudhana
CogITo Smart Journal Vol 4, No 2 (2018): CogITo Smart Journal
Publisher : Fakultas Ilmu Komputer, Universitas Klabat

Show Abstract | Download Original | Original Source | Check in Google Scholar | Full PDF (586.366 KB) | DOI: 10.31154/cogito.v4i2.129.372-381

Abstract

Human emotion recognition subject becomes important due to it's usability in daily lifestyle which requires human and computer interraction. Human emotion recognition is a complex problem due to the difference within custom tradition and specific dialect which exists on different ethnic, region and community. This problem also exacerbated due to objectivity assessment for the emotion is difficult since emotion happens unconsciously. This research conducts an experiment to discover pattern of emotion based on feature extracted from speech. Method used for feature extraction on this experiment is Mel-Frequency Cepstral Coefficient (MFCC) which is a method that similar to the human hearing system. Dataset used on this experiment is Berlin Database of Emotional Speech (Emo-DB). Emotions that are used for this experiments are happiness, boredom, neutral, sad and anger.  For each of these emotion, 3 samples from Emo-DB are taken as experimental subject. The emotion patterns are successfully visible using specific values for MFCC parameters such as 25 for frame duration, 10 for frame shift, 0.97 for preemphasis coefficient, 20 for filterbank channel and 12 for ceptral coefficients. MFCC features are then extracted and calculated to find mean values from these parameters. These mean values are then plotted based on timeframe graph to be investigated to find the specific pattern which appears from each emotion. Keywords— Emotion, Speech, Mel-Frequency Cepstral Coefficients (MFCC).
A Comparative Study of Transfer Learning and Fine-Tuning Method on Deep Learning Models for Wayang Dataset Classification Mustafid, Ahmad; Pamuji, Muhammad Murah; Helmiyah, Siti
IJID (International Journal on Informatics for Development) Vol. 9 No. 2 (2020): IJID December
Publisher : Faculty of Science and Technology, UIN Sunan Kalijaga Yogyakarta

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.14421/ijid.2020.09207

Abstract

Deep Learning is an essential technique in the classification problem in machine learning based on artificial neural networks. The general issue in deep learning is data-hungry, which require a plethora of data to train some model. Wayang is a shadow puppet art theater from Indonesia, especially in the Javanese culture. It has several indistinguishable characters. In this paper, We tried proposing some steps and techniques on how to classify the characters and handle the issue on a small wayang dataset by using model selection, transfer learning, and fine-tuning to obtain efficient and precise accuracy on our classification problem. The research used 50 images for each class and a total of 24 wayang characters classes. We collected and implemented various architectures from the initial version of deep learning to the latest proposed model and their state-of-art. The transfer learning and fine-tuning method showed a significant increase in accuracy, validation accuracy. By using Transfer Learning, it was possible to design the deep learning model with good classifiers within a short number of times on a small dataset. It performed 100% on their training on both EfficientNetB0 and MobileNetV3-small. On validation accuracy, gave 98.33% and 98.75%, respectively.
Perancangan Sistem Deteksi Emosi Mahasiswa Pada Jam Perkuliahan Menggunakan Metode Yolo Helmiyah, Siti; Bhuana, Chindu Lintang
JATISI Vol 12 No 1 (2025): JATISI (Jurnal Teknik Informatika dan Sistem Informasi)
Publisher : Universitas Multi Data Palembang

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.35957/jatisi.v12i1.10195

Abstract

This study aims to detect students' emotions during lecture hours using the YOLO (You Only Look Once) method. Emotions influence learning success, where positive emotions can enhance motivation and understanding, while negative emotions can hinder the learning process. This research employs an artificial intelligence-based video analysis approach to recognize students' facial expressions in real-time. The research stages include data acquisition using lecture videos, data preprocessing through annotation and labeling with bounding boxes, and the implementation of the YOLO method to detect three emotion categories: Enthusiastic, Confused, and Bored. Evaluation was conducted using precision, recall, and mean average precision (mAP) metrics. The test results showed that the model achieved an overall accuracy of 91.7%, with the best performance in the Enthusiastic category (97.0% accuracy) and good performance in the Bored category (93.4%). However, the model failed to detect the Confused emotion (0.0% accuracy), indicating the need for additional training data. This study demonstrates that the YOLO method has the potential to assist lecturers in understanding students' emotional states, enabling more adaptive teaching. Further development is needed to improve accuracy across all emotion categories and ensure the system functions optimally.
Identifikasi Emosi Manusia Berdasarkan Ucapan Menggunakan Metode Ekstraksi Ciri LPC dan Metode Euclidean Distance Helmiyah, Siti; Riadi, Imam; Umar, Rusydi; Hanif, Abdullah; Yudhana, Anton; Fadlil, Abdul
Jurnal Teknologi Informasi dan Ilmu Komputer Vol 7 No 6: Desember 2020
Publisher : Fakultas Ilmu Komputer, Universitas Brawijaya

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.25126/jtiik.2020722693

Abstract

Ucapan merupakan sinyal yang memiliki kompleksitas tinggi terdiri dari berbagai informasi. Informasi yang dapat ditangkap dari ucapan dapat berupa pesan terhadap lawan bicara, pembicara, bahasa, bahkan emosi pembicara itu sendiri tanpa disadari oleh si pembicara. Speech Processing adalah cabang dari pemrosesan sinyal digital yang bertujuan untuk terwujudnya interaksi yang natural antar manusia dan mesin. Karakteristik emosional adalah fitur yang terdapat dalam ucapan yang membawa ciri-ciri dari emosi pembicara. Linear Predictive Coding (LPC) adalah sebuah metode untuk mengekstraksi ciri dalam pemrosesan sinyal. Penelitian ini, menggunakan LPC sebagai ekstraksi ciri dan Metode Euclidean Distance untuk identifikasi emosi berdasarkan ciri yang didapatkan dari LPC.  Penelitian ini menggunakan data emosi marah, sedih, bahagia, netral dan bosan. Data yang digunakan diambil dari Berlin Emo DB, dengan menggunakan tiga kalimat berbeda dan aktor yang berbeda juga. Penelitian ini menghasilkan akurasi pada emosi sedih 58,33%, emosi netral 50%, emosi marah 41,67%, emosi bahagia 8,33% dan untuk emosi bosan tidak dapat dikenali. Penggunaan Metode LPC sebagai ekstraksi ciri memberikan hasil yang kurang baik pada penelitian ini karena akurasi rata-rata hanya sebesar 31,67% untuk identifikasi semua emosi. Data suara yang digunakan dengan kalimat, aktor, umur dan aksen yang berbeda dapat mempengaruhi dalam pengenalan emosi, maka dari itu ekstraksi ciri dalam pengenalan pola ucapan emosi manusia sangat penting. Hasil akurasi pada penelitian ini masih sangat kecil dan dapat ditingkatkan dengan menggunakan ekstraksi ciri yang lain seperti prosidis, spektral, dan kualitas suara, penggunaan parameter max, min, mean, median, kurtosis dan skewenes. Selain itu penggunaan metode klasifikasi juga dapat mempengaruhi hasil pengenalan emosi. AbstractSpeech is a signal that has a high complexity consisting of various information. Information that can be captured from speech can be in the form of messages to interlocutor, the speaker, the language, even the speaker's emotions themselves without the speaker realizing it. Speech Processing is a branch of digital signal processing aimed at the realization of natural interactions between humans and machines. Emotional characteristics are features contained in the speech that carry the characteristics of the speaker's emotions. Linear Predictive Coding (LPC) is a method for extracting features in signal processing. This research uses LPC as a feature extraction and Euclidean Distance Method to identify emotions based on features obtained from LPC. This study uses data on emotions of anger, sadness, happiness, neutrality, and boredom. The data used was taken from Berlin Emo DB, using three different sentences and different actors. This research resulted in inaccuracy in sad emotions 58.33%, neutral emotions 50%, angry emotions 41.67%, happy emotions 8.33% and bored emotions could not be recognized. The use of the LPC method as feature extraction gave unfavorable results in this study because the average accuracy was only 31.67% for the identification of all emotions. Voice data used with different sentences, actors, ages, and accents can influence the recognition of emotions, therefore the extraction of features in the recognition of speech patterns of human emotions is very important. Accuracy results in this study are still very small and can be improved by using other feature extractions such as provides, spectral, and sound quality, using parameters max, min, mean, median, kurtosis, and skewness. Besides the use of classification methods can also affect the results of emotional recognition.