Anggit Dwi Hartanto
Amikom Yogyakarta University

Published : 2 Documents Claim Missing Document
Claim Missing Document
Check
Articles

Found 2 Documents
Search

Lexicon-Based Indonesian Local Language Abusive Words Dictionary to Detect Hate Speech in Social Media Mardhiya Hayaty; Sumarni Adi; Anggit Dwi Hartanto
Journal of Information Systems Engineering and Business Intelligence Vol. 6 No. 1 (2020): April
Publisher : Universitas Airlangga

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.20473/jisebi.6.1.9-17

Abstract

Background: Hate speech is an expression to someone or a group of people that contain feelings of hate and/or anger at people or groups. On social media users are free to express themselves by writing harsh words and share them with a group of people so that it triggers separations and conflicts between groups. Currently, research has been conducted by several experts to detect hate speech in social media namely machine learning-based and lexicon-based, but the machine learning approach has a weakness namely the manual labelling process by an annotator in separating positive, negative or neutral opinions takes time long and tiringObjective: This study aims to produce a dictionary containing abusive words from local languages in Indonesia. Lexicon-base is very dependent on the language contained in dictionary words. Indonesia has thousands of tribes with 2500 local languages, and 80% of the population of Indonesia use local languages in communication, with the result that a significant challenge to detect hate speech of social media.Methods: Abusive words surveys are conducted by using proportionate stratified random sampling techniques in 4 major tribes on the island of Java, namely Betawi, Sundanese, Javanese, MadureseResults: The experimental results produce 250 abusive words dictionary from 4 major Indonesian tribes to detect hate speech in Indonesian social media by using the lexicon-based approach. Conclusion: A stratified random sampling technique has been conducted in 4 major Indonesian tribes to produce 250 abusive words for hate speech detection using the lexicon-based approach.
ANALISIS PERBANDINGAN KINERJA ALGORITMA STEMMING NAZIEF-ADRIANI DAN IN-IDRIS DALAM PENGOLAHAN TEKS BAHASA INDONESIA Muhammad Iqbal; Ema Utami; Anggit Dwi Hartanto
JIPI (Jurnal Ilmiah Penelitian dan Pembelajaran Informatika) Vol 11, No 1 (2026)
Publisher : STKIP PGRI Tulungagung

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.29100/jipi.v11i1.7723

Abstract

Stemming merupakan salah satu tahap penting dalam pemrosesan bahasa alami (Natural Language Processing/NLP) untuk mengubah kata berim-buhan menjadi bentuk dasarnya. Penelitian ini membandingkan dua algoritma stemming populer dalam Bahasa Indonesia, yaitu Nazief-Adriani yang berbasis kamus dan In-Idris yang berbasis aturan. Evaluasi dilakukan terhadap tiga jenis dokumen dengan karakteristik gaya bahasa berbeda untuk mengukur kinerja masing-masing algoritma berdasarkan tiga parameter: akurasi, durasi pemrosesan, dan Root Mean Square Er-ror (RMSE). Hasil menunjukkan bahwa algoritma In-Idris memiliki tingkat akurasi lebih tinggi (rata-rata 86%) dan RMSE lebih rendah (0.13) dibandingkan Nazief-Adriani (75,3% dan 0.24), sehingga lebih stabil dan akurat dalam mengidentifikasi bentuk dasar kata. Sementara itu, Nazief-Adriani menunjukkan keunggulan dalam efisiensi waktu pemrosesan. Temuan ini menegaskan pentingnya pemilihan algoritma yang disesuaikan dengan kebutuhan spesifik aplikasi NLP, serta mem-buka peluang pengembangan pendekatan hibrida untuk hasil yang lebih optimal.