This study compares TF-IDF, TF-IDF with chi-square feature selection, and BERT Sentence Transformer representations for clustering Indonesian electric-vehicle discourse in YouTube comments. A dataset of 2,601 comments from nine relevant videos published by official news accounts was preprocessed through punctuation and whitespace removal, case folding, stemming, and edit-distance correction of non-standard words. K-means was evaluated using WCSS, the elbow method, Silhouette, Calinski-Harabasz, and Davies-Bouldin indices. TF-IDF with chi-square achieved the strongest overall internal validity at k=3, with Silhouette 0.9157 and Davies-Bouldin 0.5470.
Copyrights © 2026