Claim Missing Document
Check
Articles

Found 16 Documents
Search

IMBALANCED DATA HANDLING FOR OPTIMIZING RANDOM FOREST IN SENTIMENT ANALYSIS OF EAST JAVA GUBERNATORIAL ELECTION Rahma Putri Widyaiswari; Anisa Dzulkarnain; Alqis Rausanfita
Jurnal Sistem Informasi dan Informatika (Simika) Vol. 9 No. 1 (2026): Jurnal Sistem Informasi dan Informatika (Simika)
Publisher : Program Studi Sistem Informasi, Universitas Banten Jaya

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.47080/simika.v9i1.4131

Abstract

Social media has become a strategic platform in conveying public opinion, especially at the moment of the Regional Head Election (Pilkada). The large amount of opinion data produced opens up opportunities for the application of sentiment analysis to map public perception. One of the main challenges in the classification of sentiment is the imbalance of distribution between classes, which can degrade the accuracy of the model, especially in recognizing minority classes. This study aims to analyze the impact of the application of data balancing techniques on the performance of the 2024 East Java Regional Election sentiment classification model using the Random Forest algorithm. The series of processes in the study include data preprocessing, manual sentiment labeling, text preprocessing, word weighting with TF-IDF, and model training on three data ratios, namely 90:10, 80:20, and 70:30. Each ratio was tested in three scenarios, namely no balancing (baseline), undersampling using the Tomek Links method, and oversampling using Borderline-SMOTE. Of all scenarios, Borderline-SMOTE gave the highest accuracy of 82.40% at an 80:20 ratio, an increase of 2.19% compared to the unbalanced condition at the same ratio. These results show that data balancing is able to improve the performance of the model in classifying sentiment more proportionally.
Optimalisasi Random Forest untuk Sentimen Pilkada Jawa Timur dengan Chi-Square dan Mutual Information Rahma Putri Widyaiswari; Anisa Dzulkarnain; Alqis Rausanfita
JUITA: Jurnal Informatika JUITA Vol. 13 Issue 3, November 2025
Publisher : Department of Informatics Engineering, Universitas Muhammadiyah Purwokerto

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.30595/juita.v13i3.26778

Abstract

The rise of social media has transformed the way people express opinions, including in political contexts. In the 2024 East Java Gubernatorial Election, social media platform X became a major outlet for public sentiment toward the governor and deputy governor candidates. This study aims to analyse public sentiment toward three candidate pairs by categorizing the data into three sentiment classes: positive, negative, and neutral. Feature selection was conducted by combining Term Frequency-Inverse Document Frequency (TF-IDF) with Chi-Square and Mutual Information (MI) methods to improve feature quality. The Random Forest algorithm was employed as the primary classification model. In addition, several other algorithms were tested for comparison. The results indicate that the TF-IDF and Chi-Square combination with Random Forest achieved the highest accuracy of 82.07%. These findings highlight the importance of feature selection in improving model performance for sentiment classification. The study provides insights into public opinion that can serve as a reference for strategic decision-making in the political and public sectors.
Hybrid GA-GWO with Dual-Vector Encoding for Indonesian School Timetabling Akbar Muhammad Sadat; Alqis Rausanfita; Pima Hani Safitri
Journal of Fuzzy Systems and Control Vol. 4 No. 2 (2026): Vol. 4 No. 2 (2026)
Publisher : Peneliti Teknologi Teknik Indonesia

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.59247/jfsc.v4i2.418

Abstract

School timetabling is a complex combinatorial optimization problem that involves assigning subjects, teachers, and classes to predefined time slots while satisfying numerous institutional constraints. In many Indonesian junior high schools, scheduling is still performed using manual approaches, which are often time-consuming and prone to conflicts. Compared with university timetabling, school timetabling presents additional challenges due to fixed class groups, rigid subject allocations, teacher availability constraints, and institutional regulations. To address these challenges, this study proposes a hybrid optimization framework that combines a Guided Genetic Algorithm (GA) and Grey Wolf Optimizer (GWO) for the school timetable. The proposed framework incorporates dual-vector solution encoding to provide a structured representation of scheduling components and support efficient constraint handling during the optimization process. In addition, a majority-voting and guided mutation strategy is employed to enhance the balance between exploration and exploitation. The proposed method was evaluated using real-world scheduling data from an Indonesian junior high school consisting of 27 classes, 54 teachers, 13 subjects, and 36 time slots. Experimental results show that the proposed hybrid GA-GWO achieved a fitness improvement of 95.84%, reducing the fitness value from 16,120 to 670, compared with improvements of 89.83% and 94.31% obtained by Traditional GA and Guided GA, respectively. Although the proposed method required approximately 28 minutes of execution time, it produced the highest overall timetable quality among the evaluated approaches. These findings demonstrate that the integration of dual-vector encoding, majority voting, and guided mutation within a hybrid GA-GWO framework can effectively improve timetable optimization for real-world Indonesian school scheduling environments.
Penggunaan Metode Naïve Bayes untuk Klasifikasi Topik Tweet Bidang dan Non-Bidang Rektorat Telkom University pada Akun Telyufess Mukhammad Hafiz Bima Ibrahim; Anisa Dzulkarnain; Alqis Rausanfita
Bulletin of Computer Science Research Vol. 5 No. 5 (2025): August 2025
Publisher : Forum Kerjasama Pendidikan Tinggi (FKPT)

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.47065/bulletincsr.v5i5.705

Abstract

Student complaints submitted through the Telyufess account on the social media platform X have not been optimally utilized as input for evaluating campus services at Telkom University. This study aims to classify tweets from the Telyufess account into two categories: domain-related (linked to official university units such as academics, finance, and campus services) and non-domain (general complaints unrelated to specific units). The main issue addressed is the need for an automated mapping system of student complaints to support campus service evaluations. The classification method used is Naïve Bayes, involving manual labeling by the researcher and assistant annotators (with inter-rater validation), text preprocessing (normalization using a standard dictionary and the Sastrawi library, removal of special characters, stop word filtering based on Indonesian language lists augmented with the unique term “telyu!”), tokenization, stemming, TF-IDF weighting, and dataset splitting in ratios of 65:35, 70:30, 80:20, and 90:10. A total of 1,090 tweets were collected between January 1, 2023 and January 1, 2025 using the Tweet Harvest API, based on criteria including complaints, opinions, and suggestions (retweets were excluded). The highest accuracy was achieved at 87.27% with a 90:10 split, followed by 81.74% (80:20), 78.35% (70:30), and 76.96% (65:35). Although the model showed signs of overfitting on training data (accuracy >99%), the results demonstrate that Naïve Bayes is effective for automated tweet classification and contributes to the use of social media as a data source for evaluating campus services.
Content-Based Book Recommendation System Using TF-IDF and Cosine Similarity Gede Satyamahinsa Prastita Uttama; Dharma Wiguna Limmarga; Gerrard Sebastian; Alqis Rausanfita
JOURNAL OF INFORMATICS AND TELECOMMUNICATION ENGINEERING Vol. 10 No. 1 (2026): Issues July 2026
Publisher : Universitas Medan Area

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.31289/jite.v10i1.17131

Abstract

The growth of digital platforms offering a wide variety of content or products often leads to information overload, making it difficult for users to find items that match their preferences. This research aims to design and implement a content-based recommendation system capable of providing personalized recommendations based on the similarity of item characteristics. The methods employed include data pre-processing (case folding and text cleaning), text representation using Term Frequency–Inverse Document Frequency (TF-IDF), and the measurement of similarity between objects using Cosine Similarity. The dataset contains 133,102 book titles, with descriptive attributes converted into numerical vectors to form the basis of the recommendation process. The quality of the recommendations was evaluated using two complementary approaches: intrinsic metrics (average Cosine Similarity and Average Intra-List Similarity) and user-based validation via a questionnaire completed by 31 respondents who assessed 20 sample books, measured using Precision@5, Recall@5 and F1@5 under a leave-one-out protocol. The research results show that the system generates recommendations that are relevant to the reference objects and are confirmed by the preferences of real users. This approach is effective when applied in situations where there is limited user interaction data (cold-start).
Studi Komparasi Metode Klasik dan IndoBERT untuk Analisis Sentimen Berbasis Aspek Pelaku Program MBG di Platform X Alqis Rausanfita; Affifah Mutiara Pertiwi; Vessa Rizky Oktavia
JEPIN (Jurnal Edukasi dan Penelitian Informatika) Vol. 12 No. 2 (2026): Volume 12 No 2
Publisher : Program Studi Informatika

Show Abstract | Download Original | Original Source | Check in Google Scholar

Abstract

Program Makan Bergizi Gratis (MBG) memunculkan beragam respons publik, khususnya terkait kredibilitas dan kinerja pelaku pelaksanaannya. Media sosial X menjadi ruang bagi masyarakat untuk menyampaikan dukungan, kritik, maupun pandangan terhadap pelaksanaan program secara real-time. Penelitian ini bertujuan mengklasifikasikan sentimen publik terhadap aspek pelaku Program MBG serta membandingkan kinerja metode klasifikasi klasik dengan model berbasis Transformer. Data diperoleh melalui scraping platform X menggunakan kata kunci “MBG” dan “Makan Bergizi Gratis” pada periode Januari 2025 hingga April 2026. Sebanyak 795 tweet dikumpulkan dan diperoleh 789 data setelah penghapusan duplikasi, kemudian diklasifikasikan menjadi sentimen positif, negatif, dan netral. Metode yang dibandingkan meliputi Naive Bayes, SVM, KNN, Random Forest, dan IndoBERT, dengan Random Oversampling untuk menyeimbangkan data. IndoBERT memberikan kinerja terbaik dengan akurasi 83,54%, presisi 80,43%, recall 78,41%, dan F1-score 79,28%. Hasil tersebut menunjukkan bahwa representasi bahasa berbasis konteks melalui IndoBERT mampu meningkatkan kinerja klasifikasi sentimen dibandingkan metode klasik. Penelitian ini memberikan gambaran empiris respons publik terhadap aspek pelaku Program MBG serta kontribusi metodologis dengan perbandingan lima pendekatan klasifikasi.