Research on sarcasm detection in sentences generally focuses on English and the national languages of certain countries, so that regional languages such as Palembang are still underrepresented in NLP research. This study aims to classify sarcasm in Palembang sentences by applying 12 models and evaluating their performance using a dataset containing 1,952 manually annotated samples. The experiment was also conducted using a 5-fold cross-validation scheme, and the statistical significance of the test results was analyzed using the Friedman t-test ( =41.96, FF=12.90). Contrary to common expectations, the Ensemble classifier achieved the best overall performance, attaining a mean F1-score of 0.877 and outperforming larger pre-trained and dialect-specific models. Even though it is still in the same language family, MelayuBERT's performance is very far behind (F1:0.691). Furthermore, the relatively lower performance of IndoBERT-1.5G (0.731) suggests that model scale alone does not guarantee effectiveness in low-resource language settings. These findings highlight the robustness of feature-engineered classical models compared to large-scale or dialect-specific pre-trained models for sarcasm detection in low-resource regional languages. Our research results with datasets sourced from varied domains have exceeded several national benchmarks. This study can be a baseline for further research.
Copyrights © 2026