Claim Missing Document
Check
Articles

Found 9 Documents
Search

Exploring the Impact of Back-Translation on BERT's Performance in Sentiment Analysis of Code-Mixed Language Data Setiono, Nisrina Hanifa; Sari, Yunita
IJCCS (Indonesian Journal of Computing and Cybernetics Systems) Vol 19, No 2 (2025): April
Publisher : IndoCEISS in colaboration with Universitas Gadjah Mada, Indonesia.

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.22146/ijccs.104757

Abstract

Social media, particularly Twitter, has become a key platform for communication and opinion-sharing, where code mixing, the blending of multiple languages in a single sentence, is common. In Indonesia, Indonesian-English code mixing is widely used, especially in urban areas. However, sentiment analysis on code-mixed text poses challenges in natural language processing (NLP) due to the informal nature of the data and the limitations of models trained on formal text. This study applies back translation to address these challenges and optimize BERT-based sentiment analysis. The method is tested on the INDONGLISH dataset, consisting of 5,067 labeled tweets. Results show that applying back translation directly to raw tweets yields better performance by preserving original meaning, improving model accuracy. However, when back translation follows monolingual translation, accuracy declines due to semantic distortions. Repeated translation modifies sentence structure and sentiment labels, reducing reliability. These findings indicate that each additional translation step risks decreasing sentiment analysis accuracy, particularly for code-mixed datasets, which are highly sensitive to linguistic shifts. Back translation proves to be an effective approach for formalizing data while maintaining contextual integrity, enhancing sentiment analysis performance on code-mixed text
Penerapan Alat Pilah Olah Sampah Mandiri Terintegrasi (PLASMA-T) untuk Pengelolaan Sampah di Desa Karanglewas Abednego Dwi Septiadi; Yudha Islami Sulistya; Nisrina Hanifa Setiono; Laurensius Windy Octanio Haryanto; Galih Putra Pamungkas
Masyarakat Mandiri : Jurnal Pengabdian dan Pembangunan Lokal Vol. 3 No. 1 (2026): Januari: Masyarakat Mandiri : Jurnal Pengabdian dan Pembangunan Lokal
Publisher : Lembaga Pengembangan Kinerja Dosen

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.62951/masyarakatmandiri.v3i1.2842

Abstract

The waste problem in Karanglewas Village, Kutasari District, Purbalingga Regency, continues to increase alongside agricultural, trade, and MSME activities, while the management system remains conventional, without sorting, resulting in environmental pollution and limited landfill capacity. This community service program aims to implement the Integrated Independent Waste Sorting and Processing Tool (PLASMA-T) as an innovative solution that processes organic waste into compost and melts plastic waste into paving blocks that are useful and economical. The activities are carried out in stages, including problem identification, socialization, technical training, operational trials, evaluation of results, and the handover of the equipment to the Mitra Sejahtera TPS, with ongoing assistance so that the community can operate the equipment independently. The expected outcomes include a reduction in waste volume at the TPS, increased community awareness and skills in waste management, and the creation of new business opportunities through compost and paving block products. Thus, this program not only addresses environmental issues but also strengthens the village's circular economy and has the potential to serve as a model for other regions.
Recommending E-Commerce Platforms for MSMEs: A Sentiment Analysis Approach Adiyana, Imam; Kurniawan, Angga; Rahmatika, Alfilia Hilda; Setiono, Nisrina Hanifa; Gumelar, Satya Fajar
Enthusiastic : International Journal of Applied Statistics and Data Science Volume 5 Issue 2, October 2025
Publisher : Universitas Islam Indonesia

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.20885/enthusiastic.vol5.iss2.art8

Abstract

The rapid growth of e-commerce in Indonesia presents significant opportunities for micro, small, and medium enterprises (MSMEs), yet the diversity of marketplace platforms complicates the selection of an optimal sales channel. This study addressed this challenge by developing a data-driven recommendation system based on sentiment analysis of user reviews. Utilizing a dataset of 80,000 reviews scraped from four major platforms on the Google Play Store (Shopee, Tokopedia, Lazada, and Blibli), two classification approaches were implemented and compared: support vector machine (SVM) and long short-term memory (LSTM). Both models demonstrated a competitive performance, enabling effective sentiment categorization. Furthermore, multinomial logistic regression was employed to analyze the influence of key variables rating, number of likes, and marketplace brand on sentiment outcomes. The analysis revealed that Shopee yielded the highest probability of receiving positive reviews (97.82%) and showed no significant association with negative sentiment. Consequently, this study recommends Shopee as the primary platform for MSMEs to enhance their digital presence and sales performance. The primary contribution lies in integrating machine learning-based sentiment analysis with statistical modelling to generate actionable, evidence-based marketplace recommendations for MSMEs.
Gradient Boosting Teroptimasi untuk Klasifikasi Diabetes dengan Analisis Eksplanabilitas Menggunakan SHAP dan LIME Angga Kurniawan; Nisrina Hanifa Setiono; Afin Muhammad Nurtsani; Andi Hisyam Helmi Faalih Fakhruddin; Muhammad Naufal Farabbia; Nadia Syahda Fitriani
Buletin Sistem Informasi dan Teknologi Islam (BUSITI) Vol 7, No 2 (2026)
Publisher : Universitas Muslim Indonesia

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.33096/busiti.v7i2.3413

Abstract

Penelitian ini mengusulkan kerangka klasifikasi diabetes yang mengintegrasikan tiga algoritma gradient boosting, yaitu XGBoost, LightGBM, dan CatBoost, yang dioptimasi secara otomatis menggunakan Optuna berbasis optimasi Bayesian dengan 50 percobaan dan cross validation 5-folds. Dataset yang digunakan adalah Pima Indians Diabetes Dataset dengan 768 sampel dan 8 fitur klinis. Preprocessing dilakukan menggunakan beberapa metode meliputi penggantian nilai nol sebagai nilai hilang, imputasi median dari data pelatihan, serta penambahan fitur indikator Insulin_missing. Ketidakseimbangan kelas ditangani melalui parameter scale_pos_weight yang dioptimasi bersama hiperparameter model dalam satu ruang pencarian. Evaluasi model menggunakan metrik AUC-ROC, akurasi, F1-score, recall, dan presisi dengan ambang batas klasifikasi 0,6. Analisis eksplanabilitas dilakukan menggunakan SHAP pada level global dan lokal serta LIME pada level instans untuk meningkatkan transparansi model. Hasil optimasi menunjukkan CatBoost mencapai AUC-ROC cross validation tertinggi sebesar 0.8471, diikuti LightGBM sebesar 0.8409 dan XGBoost sebesar 0.8406.  Pada data uji, XGBoost mencapai AUC-ROC 0,8246, akurasi 0,753, F1-Score 0,683, dan recall kelas diabetes 0,759, sedangkan CatBoost mencapai recall tertinggi 0,888 dengan F1-Score 0,690. Analisis SHAP secara konsisten mengidentifikasi Glucose, BMI, Age, dan DiabetesPedigreeFunction sebagai empat prediktor paling berpengaruh di ketiga model, selaras dengan pengetahuan klinis mengenai faktor risiko utama diabetes tipe 2.
Meningkatkan Literasi Digital Siswa SMA/SMK melalui Kurikulum Koding Berbasis Project-Based Learning di Sulawesi Selatan Angga Kurniawan; Mawardi Kudin; Abdul Salam At-Taqwa; Nisrina Hanifa Setiono
Jurnal Masyarakat Madani Indonesia Vol. 5 No. 2 (2026): Mei
Publisher : Alesha Media Digital

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.59025/q3fcw725

Abstract

Program pengabdian ini bertujuan meningkatkan literasi digital siswa SMA/SMK di Sulawesi Selatan melalui kurikulum koding berbasis Project-Based Learning (PjBL) selama 16 sesi. Mitra menghadapi tiga masalah utama: belum ada pembelajaran koding sistematis, rendahnya kemampuan computational thinking, dan kesiapan guru terbatas dalam PjBL. Program dilaksanakan daring via Zoom, melibatkan 86 siswa dari tiga sekolah (SMAN 3 Bone, SMAN 12 Bone, SMAN 9 Bulukumba). Metode PjBL memandu peserta membuat aplikasi (kalkulator, kasir) berbasis Python/web. Hasil evaluasi menunjukkan peningkatan rata-rata kognitif sebesar 40.2 poin (84%) dari pre-test ke post-test, dengan 82.5% peserta menghasilkan proyek akhir berkategori "Memuaskan" hingga "Sangat Memuaskan". Rata-rata kehadiran 85,7%. Program ini terbukti efektif meningkatkan literasi digital dan keterampilan koding aplikatif, serta menghasilkan model replikasi untuk pendidikan literasi digital terukur di daerah lain.
Building a Digital Lexical Resource for Banyumasan Javanese: A Low-Resource Language Approach Nisrina Hanifa Setiono; Angga Kurniawan; Viga Laksa Hardjanto
Jurnal Teknoinfo Vol. 20 No. 2 (2026): Period July 2026
Publisher : Universitas Teknokrat Indonesia

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.33365/teknoinfo.v20i2.1705

Abstract

Banyumasan Javanese, widely recognized through the Ngapak dialect, remains culturally significant but is still underrepresented in reusable computational resources. This study develops a digital Banyumasan-Indonesian lexical corpus and frames it as a reusable research artifact rather than a static appendix. The corpus was constructed from a Banyumasan-Indonesian dictionary, normalized into a structured bilingual dataset, and packaged as an installable Python resource so that it can be used directly in computational experiments. The implemented resource supports dataset loading, Banyumasan lookup, Indonesian lookup, simple translation, structured translation analysis, batch translation, and corpus statistics. The resulting corpus contains 2,000 lexical pairs, 1,996 unique Banyumasan forms, 1,444 unique Indonesian equivalents, and 4 duplicated Banyumasan headwords that preserve lexical ambiguity from the source material. To demonstrate practical utility, the study includes a 100-sentence implementation example in which Banyumasan text is translated with the published banyumasan-corpus package and evaluated against Indonesian ground truth using the Indonesian-focused embedding model LazarusNLP/all-indo-e5-small-v4. The average semantic similarity rises from 0.4833 for direct Banyumasan-versus-ground-truth comparison to 0.6427 after translation, producing an absolute gain of 0.1594 and a relative improvement of approximately 33.0% over the baseline. These findings indicate that a structured lexical corpus, when distributed in a directly reusable computational form, can strengthen both resource accessibility and small-scale downstream experimentation for a low-resource regional language.
Feature Optimization for a Content-Based Music Recommendation System on Spotify Viga Laksa Hardjanto; Nisrina Hanifa Setiono
Jurnal Teknoinfo Vol. 20 No. 2 (2026): Period July 2026
Publisher : Universitas Teknokrat Indonesia

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.33365/teknoinfo.v20i2.1731

Abstract

The growth of music streaming services like Spotify has encouraged users to explore millions of songs, making effective recommendation systems essential. This study examines content-based music recommendation systems by comparing feature configurations and similarity metrics. The system uses 11 Spotify audio features, both with and without genre features encoded using one-hot encoding. Similarity is calculated using cosine similarity and Euclidean distance on normalized features (MinMaxScaler) to generate the top 10 recommendations. Recommendations are considered relevant if the recommended song is by the same artist as the searched song. Performance is measured using Precision@10 and Recall@10 on 100 samples. Using audio features alone yields Precision@10 of 3.80–3.90% and Recall@10 of 6.36–6.45%. The addition of genre features improved performance to Precision@10 of 9.00–9.10% and Recall@10 of 9.06–9.89%. These results show that genre features significantly improve the relevance of recommendations, while both similarity metrics perform similarly when the features have been well normalized. This study contributes by demonstrating that feature representation plays a more critical role than similarity metrics in content-based music recommendation.
KLASIFIKASI GANGGUAN RANTAI PASOK MENGGUNAKAN METODE RANDOM FOREST PADA INDUSTRI LOGISTIK UD SUWARA JAYA Anggita Putri Cahyani; M Yoka Fathoni; Nisrina Hanifa Setiono
Governance IT Adoption and Technology Advance Vol. 1 No. 2 (2026)
Publisher : Telkom University

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.25124/govita.v1i2.11843

Abstract

Supply chain disruptions are one of the factors that can hinder the smoothness of logistics activities, especially in companies that need to maintain distribution accuracy and daily operational stability. Usaha Dagang Suwara Jaya, as a business engaged in chicken distribution, faces operational conditions that may potentially experience disruptions, such as discrepancies between pickup quantities and order quantities, remaining stock, shrinkage, and other operational adjustments. This study aims to develop a classification model capable of categorizing daily operational status into two classes: Disruption and Non-Disruption using the Random Forest algorithm. The disruption labels in this study were obtained based on validation from Usaha Dagang Suwara Jaya as the research partner and domain expert. The research stages include operational data collection, data preprocessing, operational variable formation, training and testing data splitting using 80:20 and 70:30 scenarios, Random Forest model training, model performance evaluation, Feature Importance analysis, and Streamlit dashboard implementation. Model evaluation was conducted using a confusion matrix with accuracy, precision, recall, and F1-score metrics. The results show that the Random Forest model with the 80:20 data split scenario achieved the best performance, with an accuracy of 0.9070, precision of 0.8947, recall of 0.8947, and F1-score of 0.8947. Meanwhile, the 70:30 scenario obtained an accuracy of 0.8462, precision of 0.8800, recall of 0.7586, and F1-score of 0.8148. Based on the Feature Importance results, Remaining Stock, Shrinkage, and average_temperature were the most contributing features to the classification results. This study also produced a Streamlit-based dashboard that can be used to display classification results, prediction probabilities, operational data calculations, new dataset uploads, and prediction history. The developed model and dashboard can be used as a supporting medium for monitoring daily supply chain disruptions at UD Suwara Jaya.
Analisis Sentimen Terhadap Program Makan Bergizi Gratis Menggunakan Pendekatan Pseudo-Labeling dan Arsitektur Transformer Wisnu Aji Sanjaya; Yoka Romadani; Jeremy Marcello Waani; Nisrina Hanifa Setiono
Prosiding Seminar Nasional Teknologi Informasi dan Bisnis Prosiding Seminar Nasional Teknologi Informasi dan Bisnis (SENATIB) 2026
Publisher : Fakultas Ilmu Komputer Universitas Duta Bangsa Surakarta

Show Abstract | Download Original | Original Source | Check in Google Scholar

Abstract

Analisis sentimen terhadap kebijakan publik sering kali terkendala oleh ketimpangan distribusi data (class imbalance) yang ekstrem, sehingga model klasifikasi cenderung mengalami bias dan mengabaikan opini minoritas. Penelitian ini bertujuan untuk membangun model klasifikasi sentimen yang mampu mengurangi kecenderungan bias kelas mayoritas pada studi kasus Program Makan Bergizi Gratis di media sosial Twitter. Metodologi yang diusulkan mengintegrasikan teknik pseudo-labeling menggunakan IndoRoBERTa (Teacher Model) untuk mengekstraksi dan menyaring 18.385 cuitan menjadi korpus Silver Standard. Untuk mengatasi ketimpangan kelas tanpa manipulasi data (oversampling), penelitian ini menerapkan Cost-Sensitive Learning melalui modifikasi bobot Inverse Class Frequency pada fungsi Cross-Entropy Loss saat melakukan fine-tuning pada Student Model (IndoBERT dan XLM-RoBERTa). Hasil eksplorasi data menunjukkan kelas Netral mendominasi (50,4%), sementara kelas Negatif menjadi minoritas ekstrem (21,4%) yang memuat kritik rasional terkait beban anggaran. Hasil komputasi membuktikan bahwa arsitektur XLM-RoBERTa menunjukkan kinerja paling optimal dengan Akurasi global 87% dan Recall makro 87%, lebih tangguh dibandingkan IndoBERT. Evaluasi Confusion Matrix menegaskan bahwa penerapan penalti bobot berhasil menekan rasio kegagalan prediksi (False Negative) pada kelas minoritas hingga menyentuh angka 13,47%. Kesimpulannya, kerangka kerja hibrida ini efektif mendisiplinkan jaringan algoritma untuk mengurangi kecenderungan bias kelas mayoritas, sehingga menghasilkan instrumen kecerdasan buatan yang lebih proporsional untuk mendukung evaluasi opini kebijakan publik.