Introduction: Distinguishing human-written scientific abstracts from AI-synthesized text remains challenging, particularly when machine-generated language appears fluent and formally structured. This study evaluates mDeBERTa v3 in a zero-shot Natural Language Inference (NLI) setting for detecting Indonesian scientific abstracts specifically synthesized using IndoT5-base-paraphrase. Method: A balanced dataset of 2,274 abstracts comprising 1,137 human-written abstracts from SINTA 3 journals and 1,137 IndoT5-synthesized counterparts was analyzed. Seven linguistic features were examined using the Mann–Whitney U test, followed by zero-shot mDeBERTa v3 classification using one-, three-, and five-aspect NLI instruction scenarios. A Random Forest classifier using the same linguistic features was included as a supervised baseline. Results and Discussion: All seven linguistic features differed significantly between classes (p < 0.001), with AI texts showing substantially higher sentence-length variation than human texts. The targeted one-aspect NLI scenario achieved the highest recall of 76.52% but only 53.52% accuracy because 790 human abstracts were misclassified as AI. Increasing instruction complexity further reduced recall. In contrast, Random Forest achieved 91.21% accuracy and an F1-score of 0.9130, confirming that the identified linguistic anomalies are strong learnable signals. Conclusion: Zero-shot mDeBERTa v3 can detect generator-specific structural artifacts but remains insufficiently precise for standalone academic-integrity screening and should be supplemented by supervised methods and human review.
Copyrights © 2026