Claim Missing Document
Check
Articles

Found 1 Documents
Search

The Impact of Signal Deletion in Text Preprocessing on Timestamp Comment Detection under Severe Class Imbalance decha danillo novicahyanto; Albertus Dwivoga Widiantoro
Insearch: Information System Research Journal Vol 6, No 02 (2026): Insearch (Information System Research) Journal
Publisher : Fakultas Sains dan Teknologi UIN Imam Bonjol Padang

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.15548/isrj.v6i02.14780

Abstract

Timestamping in YouTube comment sections is a distinctive form of participatory content curation. This study examines why four text classification algorithms, namely Logistic Regression, Support Vector Machine, Random Forest, and XGBoost, fail completely at detecting timestamp-bearing comments in an informal Indonesian-language corpus. Across 14,149 comments from Ruangguru Clash of Champions videos, the imbalance ratio reaches 73.86:1. Every model returns an accuracy near 98.7% with a minority-class recall of zero, a Balanced Accuracy of 0.50, and Cohen’s Kappa at or below zero, the textbook signature of the Accuracy Paradox. The contribution is diagnostic: class imbalance is not the sole cause. A token-level trace shows that punctuation removal followed by standalone-numeral removal deterministically destroys the pattern defining the positive label, while sparse-term pruning at a 0.99 threshold discards vocabulary occurring in fewer than 1% of documents, above the minority prevalence of 1.34%. Recovering the confusion matrices permits a complete imbalance-robust metric set to be computed, confirming that all four models are indistinguishable from a constant classifier that ignores its input. The study specifies the ablation required to separate these causes, and establishes that before minority-class failure is attributed to imbalance, researchers must verify that preprocessing has not deleted the signal being learned.