Claim Missing Document
Check
Articles

Combining feature selection and hybrid approach redefinition in handling class imbalance and overlapping for multi-class imbalanced Hartono Hartono; Erianto Ongko; Yeni Risyani
Indonesian Journal of Electrical Engineering and Computer Science Vol 21, No 3: March 2021
Publisher : Institute of Advanced Engineering and Science

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.11591/ijeecs.v21.i3.pp1513-1522

Abstract

In the classification process that contains class imbalance problems. In addition to the uneven distribution of instances which causes poor performance, overlapping problems also cause performance degradation. This paper proposes a method that combining feature selection and hybrid approach redefinition (HAR) method in handling class imbalance and overlapping for multi-class imbalanced. HAR was a hybrid ensembles method in handling class imbalance problem. The main contribution of this work is to produce a new method that can overcome the problem of class imbalance and overlapping in the multi-class imbalance problem.  This method must be able to give better results in terms of classifier performance and overlap degrees in multi-class problems. This is achieved by improving an ensemble learning algorithm and a preprocessing technique in HAR using minimizing overlapping selection under SMOTE (MOSS). MOSS was known as a very popular feature selection method in handling overlapping. To validate the accuracy of the proposed method, this research use augmented R-Value, Mean AUC, Mean F-Measure, Mean G-Mean, and Mean Precision. The performance of the model is evaluated against the hybrid method (MBP+CGE) as a popular method in handling class imbalance and overlapping for multi-class imbalanced. It is found that the proposed method is superior when subjected to classifier performance as indicate with better Mean AUC, F-Measure, G-Mean, and precision.
Hybrid Approach-RSMOTE for Handling Class Imbalance with Label Noise Hartono Hartono; Erianto Ongko
Jurnal Ilmiah Teknik Elektro Komputer dan Informatika Vol 8, No 3 (2022): September
Publisher : Universitas Ahmad Dahlan

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.26555/jiteki.v8i3.23684

Abstract

The class imbalance problem is the main problem in classification. This issue arises because real-world datasets frequently exhibit an imbalance as a result of a class with more instances than other classes. In handling class imbalance, a Hybrid Approach that blends data-level and algorithm-level approaches produce good results. However, apart from the class imbalance, which reduces classification accuracy, the complexity of the data also has an effect. The complexity of this data causes a minority noise sample which lies between the minority and the majority. In order to determine how close minority samples are to their homogeneous and heterogeneous nearest neighbors, it is necessary to calculate the relative density. The greater the proximity to the homogeneous nearest neighbors, the greater the relative density, which causes the minority samples to be in a safe state but otherwise be categorized as noisy samples. This research will combine the application of the Hybrid Approach with A self-adaptive Robust SMOTE (RSMOTE), which is an adaptive method from SMOTE that applies the concept of relative density in the over-sampling process on minority samples. The research contribution is to implement the Hybrid Approach-RSMOTE in handling class imbalance with noise by using relative density in over-sampling and also to improve classification performance. The results showed that the Hybrid Approach-RSMOTE and Hybrid Approach-SMOTE had given good results in handling class imbalance. However, the Hybrid Approach-RSMOTE gave better results in the Precision, Recall, F1-Measure, and G-Mean and showed significant differences. Based on the results of the study, it can be stated that the performance of the Hybrid Approach in handling class imbalance is influenced by the selection of the over-sampling method. The results show that RSMOTE can be considered an over-sampling method in the Hybrid Approach.
Klasifikasi Penyakit Daun Pada Tanaman Jagung Menggunakan Algoritma Support Vector Machine, K-Nearest Neighbors dan Multilayer Perceptron Jaka Kusuma; Rubianto; Rika Rosnelly; Hartono; B. Herawan Hayadi
Journal of Applied Computer Science and Technology Vol 4 No 1 (2023): Juni 2023
Publisher : Indonesian Society of Applied Science

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.52158/jacost.v4i1.484

Abstract

Corn is one of the substitute staple foods in Indonesia after rice. Maize crops grown in Indonesia often experience considerable losses due to maize plant diseases. Generally, plant diseases are initially caused by morphological changes in the leaves. Accurate detection and classification of diseases that appear on the leaves will prevent the widespread spread of the disease. This study will compare classification algorithms, namely Support Vector Machine, K-Nearest Neighbors, and Multilayer Perceptron to find the best algorithm in the classification of leaf disease in corn plants, namely, cercospora leaf spot gray, common rust, and northern leaf blight using the VGG-16 deep learning model used as image feature extraction. The results showed that the Multilayer Perceptron algorithm produced the best values with accuracy, precision, and recall of 97.4% each.
Pendampingan Pendidikan dalam Penerapan Metode Average Untuk Perhitungan Barang Habis Pakai Pada SMK Budi Agung Medan Hartono Hartono; Nur Azelina Harahap; Rohima Rohima
Jurnal Pengabdian Masyarakat Berbasis Teknologi Vol 2, No 1 (2021): VOLUME 2. NO 1. APRIL 2021
Publisher : ISB Atma Luhur

Show Abstract | Download Original | Original Source | Check in Google Scholar

Abstract

This average method assumes that any change in inventory, either due to purchases or sales by the company, will immediately get the average value of the remaining inventory. The average value of the remaining goods is obtained by dividing the total value of the remaining inventory by the number of units of the relevant goods. For data recording activities of company consumables. Microsoft Office Excel 2003 is still used, although only used to create input reports manually, these financial statements are still used. limited to LT (annual report), so the possibility of errors in accounting records is very high. The company's source of funds comes from the company's operating profits. Visual Basic comes from the BASIC language developed in 1963. BASIC stands for all-purpose symbol instruction code for beginners. As the name suggests, the purpose of making the BASIC language is to make it easier for users to learn, create and develop computer programs. Visual Basic is a further development of the BASIC language by Microsoft. Visual Basic is designed to be used as a tool to create and develop programs quickly (Rapid Application Development: RAD), especially if you use a windows-based interface (Graphical User Interface: GUI). UML (Unified Modeling Language) is a powerful tool in the field of object-oriented systems development. This is because UML provides a visual modeling language that allows system developers to create blueprints for their vision in a standard form that is easy to understand, and is equipped with an effective mechanism for sharing and communicating their designs with others. UML is built on a 4 + 1 view model. This model is based on the fact that the structure of a system is described in 5 views, one of which is the use case view. SQL Server is a Relational Database Management System (RDBMS) which is used to store data. The data stored in the database can be small or large. With SQL Server 2005, this is not a problem because SQL Server 2005 has resources capable of managing data
Classification of Basurek Batik Using Pre-Trained VGG16 and Support Vector Machine Meli Handayani; Rika Rosnelly; Hartono Hartono
Proceeding of International Conference on Information Science and Technology Innovation (ICoSTEC) Vol. 2 No. 1 (2023): Proceeding of International Conference on Information Science and Technology In
Publisher : Universitas Respati Yogyakarta

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.35842/icostec.v2i1.31

Abstract

By introducing Indonesian batik motifs, we know that the island of Sumatra, especially Bengkulu and Jambi provinces, has a distinctive batik called Basurek batik. This research aims to classify the two batik motifs using the Support Vector Machine (SVM) algorithm. First, we extract the image of the batik motif with a pre-trained VGG-16 model and then use them as a dataset for the SVM classification process. The classification process itself uses linear, polynomial, and sigmoid kernels. We divided the data 90:10 and used 10-fold cross-validation to analyze each training and testing data classification result. The results of this study are the highest values of accuracy, precision, and recall of 76.4%, 76.5%, and 76.4% produced by the linear kernel for the training data classification. For the testing data classification, both the linear and polynomial kernels generate the best accuracy, precision, and recall values of 87.5%, 90%, and 85.5%. On average, incorporating the training and testing classification results, we found that the linear kernel is the best function for classifying the Basurek batik motif using the collected images from the internet.
Sentiment Analysis of IDAHOBIT Celebrations using Naïve Bayes and Decision Tree Algorithms Jaka Kusuma; Hartono Hartono; B. Herawan Hayadi
Proceeding of International Conference on Information Science and Technology Innovation (ICoSTEC) Vol. 2 No. 1 (2023): Proceeding of International Conference on Information Science and Technology In
Publisher : Universitas Respati Yogyakarta

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.35842/icostec.v2i1.33

Abstract

The development of LGBTIQ in Indonesia reflects the shift in culture and the emergence of this phenomenon has attracted the attention of the Indonesian people. The use of NLP, ML, and statistics technology in tweet analysis can be used to identify sentiments contained in tweets. This study compares Naïve Bayes algorithm and Decision Tree in sentiment analysis classification, in which the multilingual sentiment analysis method is used in the labeling process of training data. Naïve Bayes results give the best classification with 100% accuracy, precision, and recall, and the number of positive sentiments is 385, negative sentiments are 3117, and neutral sentiments are 899. It looks that the negative class is the most superior compared to other classes. This proves that the Indonesian people have an unfavorable response to the IDAHOBIT celebration.
Sentiment Classification on Mandalika MotoGP Event Using K-Means Clustering and Random Forest Khairul Fadhli Margolang; Muhammad Zarlis; Hartono Hartono
Proceeding of International Conference on Information Science and Technology Innovation (ICoSTEC) Vol. 2 No. 1 (2023): Proceeding of International Conference on Information Science and Technology In
Publisher : Universitas Respati Yogyakarta

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.35842/icostec.v2i1.35

Abstract

As one the most famous world-class motorcycle racing competition, MotoGP is an event broadcast live on television with millions of viewers on each race. Indonesia, especially the Pertamina Mandalika Circuit, will hold this prestigious racing event in the 19th series of 2022. This event sparks Indonesian netizens' reactions on social media, especially on Twitter. This research aims to analyze the public sentiment and emotional value regarding this event, with the data collected from Twitter social media. With the features of sentiment and emotion values extracted from the contents of this tweet, we use K-means clustering to generate sentiment clusters as targets for the classification using the Random Forest (RF) algorithm. From the evaluation using the 5-fold and 10-fold cross-validation, we get the highest accuracy of 0.99, the highest precision of 0.990175, and the highest recall of 0.99 from the RF model with ten trees configuration. We also get the lowest accuracy, precision, and recall values of 0.96, 0.960934, and 0.96 from the RF models with 15 and 20 trees configuration, with the 10-fold evaluation
Sentiment Classification on Twitter Social Media Using K-Means Clustering, C4.5 and Naive Bayes (Case Study: Blocking Paypal by Kominfo) Muhammad Zulkarnain Lubis; Hartono Hartono; B. Herawan Hayadi
Proceeding of International Conference on Information Science and Technology Innovation (ICoSTEC) Vol. 2 No. 1 (2023): Proceeding of International Conference on Information Science and Technology In
Publisher : Universitas Respati Yogyakarta

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.35842/icostec.v2i1.37

Abstract

Kominfo (Ministry of Communication and Information) requires all PSEs (Electronic System Providers) to register themselves so that their access is not blocked, as shown in the case of Paypal and several other PSEs. The blocking case reaps mixed opinions from netizens, especially Twitter social media users. We use the sentiment values obtained from the content of tweets collected through the crawling process and employ the K-Means Clustering to group them into clusters. Finally, we use these clusters as the target in a dataset and classify them using the C4.5 and Naive Bayes algorithms. Of the 1000 netizen tweets studied, we found that 6.5% of netizens supported the blocking action, 75.4% did not care or felt that the blocking action had no effect on them, and 15.4% did not support the blocking by Kominfo. The classification results in this study resulted in a 98.2% accuracy value, a 95% precision value, and a 95.5% recall value.
Analysis of Machine Learning Algorithms in Predicting the Flood Status of Jakarta City Irwan Daniel; Hartono Hartono; Zakarias Situmorang
Proceeding of International Conference on Information Science and Technology Innovation (ICoSTEC) Vol. 2 No. 1 (2023): Proceeding of International Conference on Information Science and Technology In
Publisher : Universitas Respati Yogyakarta

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.35842/icostec.v2i1.38

Abstract

By mining the information in the dataset, we can solve a prediction problem, especially flood status prediction based on floodgate levels, using machine learning algorithms. This research employs three machine learning algorithms (K-Nearest Neighbor, Naive Bayes, and Support Vector Machine) for predicting the flood status using a dataset containing the data of DKI Jakarta's floodgate levels. Using a 5-fold, 10-fold, and 20-fold cross-validation evaluation, we get the highest accuracy (85.096%), f-score (85.1%), precision (85.641%), and recall (85.096%) from the model using the SVM algorithm with a polynomial kernel. Average performance-wise, the K-NN algorithm performs better than the other algorithm with an average accuracy of 83.147%, an average f-score of 83.156%, an average precision of 83.566%, and an average recall of 83.147%
Predicting Children's Talent Based On Hobby Using C4.5 Algorithm And Random Forest Sugeng Riyadi; Hartono Hartono; Wanayumini Wanayumini
Proceeding of International Conference on Information Science and Technology Innovation (ICoSTEC) Vol. 2 No. 1 (2023): Proceeding of International Conference on Information Science and Technology In
Publisher : Universitas Respati Yogyakarta

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.35842/icostec.v2i1.54

Abstract

A person's talent is closely related to intelligence, hobbies, and interests. These factors are the best features to be used in a dataset to predict a children's talent, such as in an academy, arts, or sports. This research uses the C4.5 and random forest algorithms in 8 different models to predict a children's talent based on a dataset gained from a survey involving 1601 parents. Each model contains four training-testing data ratios, such as 50:50, 60:40, 70:30, and 80:20. We calculate each model prediction performance using 10-fold and 20-fold crossvalidation, with the accuracy, f-score, precision, and recall values as a comparison. The best result for the training evaluation we get is 91.5% for each comparison value from the random forest model (70:30 ratio) using a 20-fold cross-validation. For the testing evaluation, we get 92.7%, 92.8%, 92.8%, and 92.7% from the random forest model (50:50 ratio). The worst testing evaluation we get is 81.7% for each comparison value from the C4.5 model (50:50 ratio) using a 20-fold cross-validation. For the testing evaluation, we get 89.2%, 89.2%, 89.3%, and 89.2% from the C4.5 model (50:50 ratio).
Co-Authors Aditya Pratama, Bayu Ammar Yasir Nasution Andi Rahmadsyah Andik Bintoro Andre Hasudungan Lubis Andriasan Sudarso Arnes Sembiring Asmah Indrawati B. Herawan Hayadi B. Herawan Hayadi Brilliant Handyman Manalu Citra Rahmadhani Cut Ita Erliana Dadan Ramdan Dahlan Abdullah Dedi Sahputra Desniarti Dian Maya Sari Elvie Maria Erianto Ongko Erianto Ongko Erianto Ongko Erianto Ongko Erianto Ongko Erianto Ongko Erianto Ongko Erianto Ongko Erna Budhiarti Nababan Faadhil, Faadhil Finta Aramita Firman Syahputra Firman Syahputra Gio, Prana Ugiana Habib Satria Imelda Maelani Iqbal Giffari Ritonga Irwan Daniel Jaka Kusuma Jaka Kusuma Khairul Fadhli Margolang Limas, Agus Fahmi Maricha Elveny Marischa Elveny, Marischa Martini, Dewi Meli Handayani Mendarissan Aritonang Muhammad Ikhwani Muhammad Khahfi Zuhanda Muhammad Sadikin Muhammad Zarlis Muhammad Zulkarnain Lubis Mulkan Andika Situmorang N. Nazaruddin Nadapdap, Kristanty M. N. Nadra Ideyani Vita Nasution, Mahyuddin K.M Nos Sutrisno Nur Anzelina Nur Azelina Harahap Nursie, Aly Nurwijayanti Opim Salim Sitompul Prana Ugi Rachmat Aulia, Rachmat Rahmad B.Y Syah Rahmad Syah, Rahmad Rahman, Sayuti Rana Fathinah Ananda Retna Astuti Kuswardani Rezzy Eko Caraka Ria Wuri Andary Rika Rosnelly Rika Rosnelly Rika Rosnelly Rika Rosnelly Rika Rosnelly, Rika Rohima Rohima Rubianto Sabina Krisdayanti Samsul A Rahman Sidik Hasibuan Sembiring, Arnes Silvia Lestari Silvia Lestari Siti Aisyah Sitorus, Peniel Sam Putra Sugeng Riyadi suswati suswati Suswati Suswati Syah, Rahmad B.Y Tulus Tulus Wanayumini Wulan Dari Yeni Risyani Yudi Gebri Foenna Zakarias Situmorang