Background: The increasing prevalence of Indonesian–English code-mixing in social media reflects sociolinguistic shifts driven by globalization, while challenging language processing systems that assume monolingual input. Objective: This study evaluates the effectiveness of machine learning models in detecting Indonesian–English code-mixing in Twitter data and the role of linguistic features in improving accuracy. Method: A supervised approach was applied using SVM and Random Forest classifiers on annotated tweets, enriched with features such as part-of-speech patterns, token-level language identification, and morphological markers. Results: SVM models outperform baselines with high accuracy and balanced precision–recall, while linguistic features significantly enhance detection, especially for intra-word mixing; errors mainly arise from lexical borrowing, short contexts, and morphologically integrated forms. Implication: These findings emphasize the importance of integrating linguistic knowledge into computational models to improve robustness in multilingual and low-resource settings. Novelty: This study demonstrates that linguistically informed machine learning frameworks enhance both performance and interpretability in detecting Indonesian–English code-mixing.
Copyrights © 2026