cover
Contact Name
Huzain
Contact Email
huzain.azis@umi.ac.id
Phone
+628114484875
Journal Mail Official
ijodas.journal@gmail.com
Editorial Address
Jln. Paccerakkang, Kel. Berua, Kec.Biringkanaya, Kota Makassar, Propinsi Sulawesi Selatan, 90241
Location
Unknown,
Unknown
INDONESIA
Indonesian Journal of Data and Science
Published by yocto brain
ISSN : -     EISSN : 27159930     DOI : -
Core Subject : Science, Education,
IJODAS provides online media to publish scientific articles from research in the field of Data Science, Data Mining, Data Communication, Data Security and Data Representation
Articles 191 Documents
The Performance of Support Vector Machine in Classifying Public Sentiment toward Student Suicide Cases I Gusi Gede Bagus Ngurah Sarjana; Made Leo Radhitya; Ni Wayan Suardiati Putri
Indonesian Journal of Data and Science Vol. 7 No. 2 (2026): Indonesian Journal of Data and Science
Publisher : yocto brain

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.56705/ijodas.v7i2.431

Abstract

Introduction: The rapid growth of social media has generated large volumes of user-generated content that can be analyzed to understand public responses to sensitive social issues. This study evaluates the performance of Support Vector Machine (SVM) in classifying public sentiment toward a widely discussed student suicide case based on YouTube comments. Method: A total of 5,000 comments were collected from a video on the Denny Sumargo YouTube channel using the YouTube Data API and categorized into positive and negative sentiments. Text preprocessing included cleaning, normalization, tokenization, stop-word removal, and stemming. Term Frequency-Inverse Document Frequency (TF-IDF) was used for feature extraction, while Synthetic Minority Over-sampling Technique (SMOTE) addressed class imbalance. The dataset was divided into 80% training and 20% testing data, and SVM was applied for binary sentiment classification. Results and Discussion: The SVM model achieved 99.96% training accuracy and 89.25% test accuracy, with precision, recall, and F1-score consistently around 89%. These results indicate that the TF-IDF, SMOTE, and SVM pipeline effectively classified Indonesian social media comments despite the linguistic complexity of discussions surrounding sensitive issues. Conclusion: SVM demonstrates effective and robust performance for classifying public sentiment in Indonesian YouTube comments and provides a useful approach for analyzing public responses to sensitive social phenomena.
Digital Burnout Risk Classification Based on Doomscrolling and Fear of Missing Out Using the Mamdani Fuzzy Method Matelda Yunanta Ambon; Rosmasari; Fahrul Agus
Indonesian Journal of Data and Science Vol. 7 No. 2 (2026): Indonesian Journal of Data and Science
Publisher : yocto brain

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.56705/ijodas.v7i2.442

Abstract

Introduction: Intensive digital technology use among university students may contribute to Digital Burnout, particularly when accompanied by doomscrolling and Fear of Missing Out (FOMO). This study classifies Digital Burnout risk into low, moderate, and high categories using these two digital-behavior factors. Method: A Mamdani Fuzzy Inference System was implemented in MATLAB using Doomscrolling and FOMO as inputs and Digital Burnout as the output. Data were collected through validated self-report questionnaires from 414 students, resulting in 408 valid records. The system employed nine IF–THEN rules, trapezoidal and triangular membership functions within a normalized [0,100] domain, centroid defuzzification, and twelve parameter-adjustment iterations. Results and Discussion: The ground-truth distribution showed that 47.5% of students were categorized as having high Digital Burnout, 30.1% moderate, and 22.3% low. The optimal fuzzy configuration achieved 67.89% accuracy and a macro F1-score of 0.6725. The High category achieved very high precision of 0.9688 but moderate recall of 0.6392, indicating that some high-risk students remained undetected. Classification accuracy was higher among gamers (77.19%) than general respondents (61.18%). Conclusion: The Mamdani Fuzzy approach demonstrates the feasibility of classifying Digital Burnout risk from Doomscrolling and FOMO; however, its moderate accuracy and reliance on self-reported, non-clinical labels indicate that it should be considered an exploratory prototype requiring independent validation and expert-reviewed rules before practical deployment.
Effect of Spatial, Intensity, and Hybrid Augmentation on Kidney CT Image Classification Ardha Ardhana Putra Agustavada; Aji Prasetya Wibawa; Dafa Fadhilah Hilmi; Abdullah Sholum; Felix Andika Dwiyanto
Indonesian Journal of Data and Science Vol. 7 No. 2 (2026): Indonesian Journal of Data and Science
Publisher : yocto brain

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.56705/ijodas.v7i2.443

Abstract

Introduction: Kidney diseases remain a major global health challenge, and computed tomography (CT) provides detailed renal imaging that supports accurate diagnosis. Although data augmentation is commonly used to improve deep learning performance on limited medical datasets, the effects of different augmentation strategies on image characteristics and model learning behavior remain insufficiently understood. Method: This study evaluated spatial, intensity, and hybrid augmentation for kidney CT image classification using a custom CNN, MobileNetV2, and EfficientNet-B0. A public dataset containing 12,446 CT images across Normal, Cyst, Tumor, and Stone classes was partitioned using stratified sampling into training, validation, and test sets. Four scenarios—baseline, spatial, intensity, and hybrid augmentation—were evaluated across three independent random seeds using accuracy, precision, recall, F1-score, and AUC. Results and Discussion: The baseline achieved mean accuracies of 99.96%, 98.32%, and 94.06% for CNN, MobileNetV2, and EfficientNet-B0, respectively. Intensity augmentation slightly improved CNN accuracy to 99.99% and consistently produced smaller performance degradation and more stable convergence than spatial and hybrid augmentation. Spatial and hybrid transformations generally reduced classification performance, indicating that excessive geometric changes may disrupt diagnostically relevant anatomical features. Conclusion: Baseline training provided the best overall performance, while intensity augmentation was the most effective augmentation strategy, demonstrating that preserving anatomically meaningful image characteristics is more important than indiscriminately increasing data diversity.
Weakly Supervised Sentiment Analysis of Gold Price Discussions Using Conventional Machine Learning and IndoBERT M Rizki Hardika; Brina Miftahurrohmah
Indonesian Journal of Data and Science Vol. 7 No. 2 (2026): Indonesian Journal of Data and Science
Publisher : yocto brain

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.56705/ijodas.v7i2.445

Abstract

Introduction: Gold price movements attract substantial public and investor attention because gold serves as both a safe-haven asset and a hedging instrument. This study investigates Indonesian public sentiment toward gold price discussions on Platform X using a weakly supervised sentiment-analysis framework. Method: A total of 7,283 Indonesian-language tweets containing the keyword “harga emas” were collected during 2023–2025, with 4,429 tweets retained after preprocessing. Sentiment labels were generated using a domain-specific lexicon and validated through manual annotation. Naïve Bayes, K-Nearest Neighbor, Support Vector Machine, and IndoBERT were evaluated using the same train–test partition. Results and Discussion: Manual validation achieved a Cohen’s Kappa of 0.8718, indicating almost perfect inter-annotator agreement, while the lexicon-based labels achieved 70.62% accuracy against the manually annotated reference. IndoBERT achieved the highest performance on weakly supervised labels with 98.31% accuracy and a 98.16% macro F1-score, outperforming SVM, Naïve Bayes, and KNN. However, its accuracy decreased to 69.49% when evaluated against manually annotated data, demonstrating that downstream performance remains strongly influenced by weak-label quality. Conclusion: Weak supervision provides an efficient and scalable approach for large-scale Indonesian financial sentiment annotation, while contextual models such as IndoBERT offer superior classification performance; however, reliable manual validation remains essential to mitigate label noise and improve generalizability.
Sentiment Classification of TikTok Comments on The Free Nutritious Meal Program Using Multinomial Naïve Bayes Salom Sefanya Onibala; Rosmasari; Kezia Arum Sary
Indonesian Journal of Data and Science Vol. 7 No. 2 (2026): Indonesian Journal of Data and Science
Publisher : yocto brain

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.56705/ijodas.v7i2.448

Abstract

Introduction: TikTok has become an important platform for public discussion of government programs, including Indonesia’s Free Nutritious Meal (MBG) Program, while the large volume of comments makes manual sentiment analysis inefficient. This study evaluates Multinomial Naïve Bayes for classifying public sentiment toward the MBG Program using TikTok comments. Method: A total of 1,048 comments were collected from the official National Nutrition Agency TikTok account using Apify, with 1,028 comments remaining after preprocessing. Sentiment pseudo-labels were generated using IndoRoBERTa, producing 333 positive and 695 negative comments. Undersampling balanced the dataset to 666 comments, followed by an 80:20 stratified train–test split. TF-IDF feature extraction and Multinomial Naïve Bayes classification were integrated in a Scikit-learn Pipeline, while GridSearchCV with 5-fold cross-validation optimized model parameters. Results and Discussion: The optimized model achieved 80.60% accuracy, 81.04% precision, 80.60% recall, and an 80.53% F1-score on 134 test comments. Negative sentiment dominated the original post-preprocessing dataset at 67.61%. However, these results primarily indicate agreement with IndoRoBERTa-generated pseudo-labels rather than human-verified sentiment labels. Conclusion: The TF-IDF–Multinomial Naïve Bayes pipeline provides a useful baseline for MBG-related TikTok sentiment classification, but manual label validation and comparison with alternative classifiers are required to establish stronger reliability and generalizability.
Comparative Evaluation of Machine Learning Models for Heavy Crude Oil Viscosity Prediction Using Repeated Nested Cross-Validation and Independent Holdout Testing Enggie Hendrawan Saputra; Ilham Ari Elbaith Zaeni; Didik Dwi Prasetya; Azlan Mohd Zain; Welly Antonius; I Made Wirawan
Indonesian Journal of Data and Science Vol. 7 No. 2 (2026): Indonesian Journal of Data and Science
Publisher : yocto brain

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.56705/ijodas.v7i2.455

Abstract

Introduction: Accurate prediction of heavy crude oil viscosity is important for reservoir engineering, production planning, and flow assurance because viscosity strongly affects fluid mobility and transport behavior. This study comparatively evaluates established machine learning models under a rigorous validation protocol rather than proposing a new predictive framework. Method: A published Middle Eastern heavy crude-oil dataset containing 196 development measurements and 47 independent holdout measurements was used. Linear Regression, Support Vector Regression, Random Forest, Gradient Boosting, and the Beggs–Robinson correlation were evaluated using repeated nested cross-validation with five outer folds repeated twice and five inner folds. Preprocessing and hyperparameter selection were embedded within the validation pipeline, while the untouched holdout set was used only for final evaluation. Results and Discussion: Gradient Boosting achieved the best internal performance with R² = 0.99313 and RMSE = 11.41 cP. On the independent holdout set, it achieved R² = 0.99308, RMSE = 8.43 cP, MAE = 6.64 cP, and MAPE = 0.78%, outperforming Random Forest and Support Vector Regression. Residual diagnostics showed no detectable heteroscedasticity, while permutation importance identified temperature and C7+ as the dominant predictors. Conclusion: Gradient Boosting provides highly accurate viscosity predictions within the sampled domain; however, the absence of row-level oil identifiers and external reservoir data limits conclusions regarding oil-disjoint and field-level generalization.
An IoT-Based Precision Hydroponic Monitoring System and Long-Term Characterization of Low-Cost Temperature Sensor Drift Dedy Atmajaya; Abdullah Basalamah; Nia Kurniati; Muhammad Iqbal; Thalita Sherly Putri Jasmin
Indonesian Journal of Data and Science Vol. 7 No. 2 (2026): Indonesian Journal of Data and Science
Publisher : yocto brain

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.56705/ijodas.v7i2.456

Abstract

Introduction: Temperature monitoring is critical in hydroponic cultivation because it influences nutrient solubility, dissolved oxygen, and root uptake, yet low-cost digital sensors commonly used in Internet of Things (IoT) systems may experience accuracy degradation during long-term deployment. This study develops a low-cost IoT monitoring platform for nutrient film technique (NFT) hydroponics and characterizes temperature-sensor drift under continuous operating conditions. Method: The system employed two redundant temperature sensors in the nutrient channel and one ambient sensor, with measurements timestamped, filtered, stored locally at the edge, and visualized through a cloud dashboard. Temperature data were recorded hourly for 40 days, producing 960 observations per sensor. Drift was evaluated from the deviation between the primary and redundant channel sensors using mean absolute error, maximum absolute deviation, standard deviation, drift onset, and estimated drift rate. Results and Discussion: The primary sensor showed a mean absolute deviation of 0.82 °C over the full monitoring period and a maximum deviation of 1.47 °C. Sensor agreement remained close during the first 10 days but diverged after approximately day 12, with late-period MAE increasing to 1.18 °C and an estimated drift rate of about 0.04 °C/day. Conclusion: Long-term drift in low-cost temperature sensors can materially affect hydroponic monitoring accuracy, and redundant sensing provides a practical baseline for future adaptive edge-based calibration methods.
Zero-Shot Detection of IndoT5-Synthesized Indonesian Scientific Abstracts Using mDeBERTa v3 Aldo Syahputra; Aris Wahyu Murdiyanto; Ulfi Saidata Aesyi
Indonesian Journal of Data and Science Vol. 7 No. 2 (2026): Indonesian Journal of Data and Science
Publisher : yocto brain

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.56705/ijodas.v7i2.457

Abstract

 Introduction: Distinguishing human-written scientific abstracts from AI-synthesized text remains challenging, particularly when machine-generated language appears fluent and formally structured. This study evaluates mDeBERTa v3 in a zero-shot Natural Language Inference (NLI) setting for detecting Indonesian scientific abstracts specifically synthesized using IndoT5-base-paraphrase. Method: A balanced dataset of 2,274 abstracts comprising 1,137 human-written abstracts from SINTA 3 journals and 1,137 IndoT5-synthesized counterparts was analyzed. Seven linguistic features were examined using the Mann–Whitney U test, followed by zero-shot mDeBERTa v3 classification using one-, three-, and five-aspect NLI instruction scenarios. A Random Forest classifier using the same linguistic features was included as a supervised baseline. Results and Discussion: All seven linguistic features differed significantly between classes (p < 0.001), with AI texts showing substantially higher sentence-length variation than human texts. The targeted one-aspect NLI scenario achieved the highest recall of 76.52% but only 53.52% accuracy because 790 human abstracts were misclassified as AI. Increasing instruction complexity further reduced recall. In contrast, Random Forest achieved 91.21% accuracy and an F1-score of 0.9130, confirming that the identified linguistic anomalies are strong learnable signals. Conclusion: Zero-shot mDeBERTa v3 can detect generator-specific structural artifacts but remains insufficiently precise for standalone academic-integrity screening and should be supplemented by supervised methods and human review.
An IoT-Based Fuzzy Decision Support System for Rhizobium Inoculation to Improve Soybean Productivity Andi Ulfah Tenripada; Lukman Syafie; Muh Mu&#039;min; Raqhib Ataillah; Muh Nawir
Indonesian Journal of Data and Science Vol. 7 No. 2 (2026): Indonesian Journal of Data and Science
Publisher : yocto brain

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.56705/ijodas.v7i2.458

Abstract

Introduction: Soybean productivity in Indonesia remains below national demand, while the effectiveness of Rhizobium inoculation depends strongly on dynamic soil conditions such as pH, moisture, temperature, and nitrogen availability. This study develops an Internet of Things (IoT)-based decision support system for adaptive Rhizobium inoculation in highland soybean cultivation. Method: The system integrates a Soil NPK RS485 Modbus sensor, ESP32-WROOM-32D microcontroller, MQTT communication, and the SMARTO web dashboard to monitor four environmental parameters in real time. A Mamdani Fuzzy Inference System was implemented using 17 membership functions and 25 expert-derived IF–THEN rules, with centroid defuzzification producing Rhizobium dose recommendations from 0 to 200 g/ha. Results and Discussion: Field readings of pH 7.5, soil moisture 55%, temperature 26°C, and nitrogen 155 mg/kg generated a recommendation of 16.56 g/ha, classified as Very Low, with a pump volume of 33 mL/ha. Expert validation across 30 simulated highland scenarios produced an overall agreement rate of 83.3%, demonstrating satisfactory consistency between system recommendations and agronomic judgment. Conclusion: The proposed IoT-Fuzzy DSS demonstrates the feasibility of location-specific, real-time, and interpretable Rhizobium inoculation support for highland soybean cultivation, providing a practical foundation for precision biological-input management. 
The Performance of the XGBOOST-LSTM and CNN-LSTM Algorithms in the Analysis of Stock Price Prediction Models for the Indonesian Banking Sector M. Zainal Arifin; Filbert Chaitra Bessel Kristianto; Fadia Irsania Putri; Agung Bella Putra Utama
Indonesian Journal of Data and Science Vol. 7 No. 2 (2026): Indonesian Journal of Data and Science
Publisher : yocto brain

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.56705/ijodas.v7i2.459

Abstract

Introduction: Accurate stock price forecasting is important for investors and financial institutions because stock movements reflect dynamic market conditions and influence investment decision-making. This study compares CNN-LSTM and XGBoost-LSTM models for predicting stock prices in the Indonesian banking sector. Method: Historical daily closing prices of PT Bank Central Asia Tbk (BBCA), PT Bank Rakyat Indonesia (Persero) Tbk (BBRI), and PT Bank Mandiri (Persero) Tbk (BMRI) were collected from Yahoo Finance, with 1,000 observations for each stock. Data were normalized using Min-Max scaling, transformed using a four-day sliding window to predict the following day, and chronologically divided into 80% training and 20% testing sets. Both models were evaluated using RMSE, MAE, R², and MAPE, with each experiment repeated ten times. Results and Discussion: CNN-LSTM consistently outperformed XGBoost-LSTM on all three test datasets. For BBCA, BBRI, and BMRI, CNN-LSTM achieved R² values of 0.8589, 0.7746, and 0.7409, respectively, compared with 0.8048, 0.7509, and 0.6260 for XGBoost-LSTM. CNN-LSTM also produced lower RMSE, MAE, and MAPE values across all test sets, indicating stronger and more stable generalization. Conclusion: CNN-LSTM provides more reliable predictive performance than XGBoost-LSTM for the evaluated Indonesian banking stocks, demonstrating the effectiveness of combining local feature extraction with temporal dependency learning for stock price forecasting