Muhammad Iqbal Shiddiq
Budi Luhur University

Published : 1 Documents Claim Missing Document
Claim Missing Document
Check
Articles

Found 1 Documents
Search

IMPLEMENTASI METODE NER STATISTIK BERBASIS WEB UNTUK EKSTRAKSI ENTITAS PADA BERITA BENCANA ALAM TVRI Muhammad Iqbal Shiddiq; Indra Indra
SKANIKA: Sistem Komputer dan Teknik Informatika Vol 9 No 2 (2026): Jurnal SKANIKA Juli 2026
Publisher : Universitas Budi Luhur

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.36080/skanika.v9i2.3897

Abstract

Disaster news published on TVRI's portal is typically presented in unstructured text containing important details such as disaster type, location, time, and related organizations, making manual information extraction inefficient. This study implements a web-based Statistical Named Entity Recognition (NER) using the Naive Bayes algorithm to extract disaster, location, time, date, and organization entities from 200 TVRI disaster news articles. Training data was created through expert-validated manual labeling using the BIO (Begin-Inside-Outside) scheme, producing 40,667 labeled tokens converted into CurrentWord, Token Type, CurrentTag, Bef1Tag, and Class features, with Laplace Smoothing applied to address zero probability. Testing was conducted before and after applying Random Undersampling to assess its effect on class balance. Before Random Undersampling, the model achieved 87.91% accuracy, 55.73% precision, 58.66% recall, and 56.63% F1-score, after application, accuracy dropped to 83.17% and recall to 55.50%, while precision rose to 62.74% and F1-score to 57.45%. These results show that Random Undersampling improved the precision-recall balance despite lower accuracy, proving that Naive Bayes-based Statistical NER can automatically extract entities, though the relatively small F1-score improvement (about 0.82 percentage points) indicates a need for further testing.