The effectiveness of suicide prevention in Indonesia is severely hindered by significant underreporting and social stigma, leading at-risk individuals to express their distress on social media platforms such as Twitter. However, detecting these signals is computationally challenging due to the informal nature of Indonesian slang and the risk of losing emotional context through aggressive pre-processing. This study aims to evaluate the performance of various deep learning and traditional models in detecting suicidal ideation while specifically analyzing the impact of syntactic preservation. We performed a comparative analysis using IndoBERT, Multilingual BERT (mBERT), and three baseline models including Logistic Regression, Support Vector Machine, and Random Forest which were evaluated through two distinct pre-processing strategies, namely Dataset A (fully preprocessed) and Dataset B (partially preprocessed). Experimental results demonstrate that IndoBERT achieved the highest performance on Dataset B with an accuracy of 90.6%, outperforming traditional baselines. The results indicate that omitting stopword removal and stemming is more effective for transformer-based architectures, as retaining original word forms helps preserve semantic integrity and emotional nuances required for contextual understanding. These findings highlight the importance of context-aware solutions to bridge the data gap in public health. By providing a robust computational framework that maintains linguistic integrity, this research has the potential to facilitate scalable, real-time interventions for supporting the identification of high-risk individuals within the Indonesian digital landscape. This study contributes to the development of automated mental health screening tools by demonstrating the potential of NLP-based methods for detecting psychological distress in low-resource, informal languages.
Copyrights © 2026