Claim Missing Document
Check
Articles

Found 2 Documents
Search
Journal : journal of applied informatics and computing

From Sparse Features to Transformers: A Statistical Evaluation of TF-IDF, FastText, and IndoBERT for Sentiment Classification of Indonesian Travel App Reviews Claudian Tikulimbong Tangdilomban; Syaifullah Yusuf Ramdhan; Muhammad Rizal; Cici Suhaeni; Bagus Sartono
Journal of Applied Informatics and Computing Vol. 10 No. 3 (2026): June 2026
Publisher : Politeknik Negeri Batam

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.30871/jaic.v10i3.12610

Abstract

This study compares three text representation techniques, namely TF-IDF, FastText, and IndoBERT, in the sentiment classification task of Indonesian-language user reviews of travel applications. The dataset consists of 4.000 reviews from Traveloka and Tiket.com, collected through Google Play Store scraping and manually annotated with sentiment labels. Each representation technique was combined with three classification algorithms, namely Support Vector Machine, Logistic Regression, and Random Forest, resulting in nine experimental configurations. The evaluation was conducted using stratified 5-fold cross-validation with macro F1-score as the primary metric, supported by hyperparameter tuning using GridSearchCV, paired t-test statistical analysis, and Cohen’s d effect size measurement. The evaluation results indicate that IndoBERT generally achieved the best performance compared to TF-IDF and FastText. The best configuration was obtained by IndoBERT with Logistic Regression, achieving an F1-score of 0.9261 after tuning. The statistical test showed that the performance differences among text representations were statistically significant, with large effect sizes in the comparison between IndoBERT and TF-IDF (d = −1.36) and between IndoBERT and FastText (d = −1.10). Nevertheless, TF-IDF combined with Logistic Regression and SVM remained competitive, achieving an F1-score of approximately 0.892 after tuning, making it a lightweight and interpretable alternative. This study concludes that the quality of text representation has a more dominant influence on sentiment classification performance than the complexity of the classification algorithm.
Performance Evaluation of Word2Vec and FastText Embeddings in a CNN-BiLSTM Model for Sentiment Classification of the LPDP Alumni Controversy Dwi Erzalianti; Joice Junansi Tandirerung; Cici Suhaeni; Bagus Sartono
Journal of Applied Informatics and Computing Vol. 10 No. 4 (2026): August 2026
Publisher : Politeknik Negeri Batam

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.30871/jaic.v10i4.13298

Abstract

This study aims to analyze public sentiment toward the LPDP alumni controversy on social media using a deep learning approach. The research data consist of YouTube user comments related to the LPDP issue, which were processed through text preprocessing and automatically labeled using IndoBERT into three sentiment classes: negative, neutral, and positive. This study compares two text representation methods, namely Word2Vec and FastText, implemented within a hybrid CNN–BiLSTM architecture. In addition, data imbalance was addressed using class weighting and undersampling scenarios, while TF-IDF-based Logistic Regression was used as the baseline model. The results show that the baseline achieved an accuracy of 0.83 but was strongly biased toward the negative class as the majority class. The CNN–BiLSTM model improved the ability to detect minority classes. Under the class weighting scenario, FastText demonstrated more stable performance with an accuracy of 0.77 and a macro F1-score of 0.62. Under the undersampling scenario, Word2Vec was more stable, achieving an accuracy of 0.68 and a macro F1-score of 0.67. These findings indicate that both text representation and imbalance-handling strategies substantially affect sentiment classification performance.