Claim Missing Document
Check
Articles

Found 1 Documents
Search
Journal : International Journal of Advances in Data and Information Systems

Features Selection for Entity Resolution in Prostitution on Twitter Reisa Permatasari; Nur Aini Rakhmawati
International Journal of Advances in Data and Information Systems Vol. 2 No. 1 (2021): April 2021 - International Journal of Advances in Data and Information Systems
Publisher : Indonesian Scientific Journal

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.25008/ijadis.v2i1.1214

Abstract

Entity resolution is the process of determining whether two references to real-world objects refer to the same or different purposes. This study applies entity resolution on Twitter prostitution dataset based on features with the Regularized Logistic Regression training and determination of Active Learning on Dedupe and based on graphs using Neo4j and Node2Vec. This study found that maximum similarity is 1 when the number of features (personal, location and bio specifications) is complete. The minimum similarity is 0.025662627 when the amount of harmful training data. The most influencing similarity feature is the cellphone number with the lowest starting range from 0.997678459 to 0.999993523.  The parameter - length of walk per source has the effect of achieving the best similarity accuracy reaching 71.4% (prediction 14 and yield 10).