Claim Missing Document
Check
Articles

Found 23 Documents
Search

Toward Real-Time Hoax Detection: Integrating Transformer for News Scraping and Semantic Analysis M. Adnan Nur; Herlinah; Sitti Zuhriyah
Journal of Information Systems Engineering and Business Intelligence Vol. 12 No. 2 (2026): June
Publisher : Universitas Airlangga

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.20473/jisebi.12.2.348-360

Abstract

Background: Social media has become one of the primary sources of information for the public, but it is also vulnerable to fake news (hoaxes). Various supervised learning approaches have been used for hoax detection; however, they generally depend on labeled data availability and require retraining to handle new claims. Objective: This study implements and evaluates a zero-shot inference-based hoax verification pipeline capable of verifying claims in near real-time without model retraining. Methods: The study uses a computational experiment approach consisting of four stages: (1) construction of a test claim dataset through paraphrasing using the NLLB-200, T5, and mT5 models; (2) keyword extraction using KeyBERT with IndoBERT embeddings and parameter optimization; (3) news article summarization using IndoBART; and (4) semantic similarity analysis by comparing Multilingual-E5, BGE, LaBSE, SBERT, IndoBERT, and TF-IDF as the baseline. Results: In the keyword extraction stage, the best balance between relevance, redundancy, and keyword coverage was achieved using the hybrid configuration of Top-k = 12, Top-p = 0.85, MMR = 0.6, and Ctx = 2. In the verification stage, Multilingual-E5 delivered the highest performance, achieving an accuracy of 0.96 and an F1-score of 0.94. In contrast, IndoBERT produced the lowest results, particularly for paraphrased claims. TF-IDF achieved good performance on original claims but experienced a decline when semantic variations were handled. Article summarization helped reduce article length, although this stage removed part of the information relevant to the verification process in some cases. Conclusion: The results of this study show that the zero-shot inference-, retrieval-, and semantic similarity-based approaches are capable of verifying hoaxes with high accuracy while maintaining stable performance across various claim variations. This approach offers greater flexibility for handling new claims than supervised learning methods and is suitable for application in near real-time hoax verification systems.   Keywords: hoax verification, keyword extraction, semantic similarity, summarization, zero-shot inference
Real-Time News Authenticity Verification Using MPNet (Masked and Permuted Pre-training Network)-Based Sentence Embeddings on Digital News Portals Ira Lestari; Herlinah Herlinah; M. Adnan Nur
Journal of Applied Informatics and Computing Vol. 10 No. 3 (2026): June 2026
Publisher : Politeknik Negeri Batam

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.30871/jaic.v10i3.13193

Abstract

The dissemination of fake news (hoaxes) on digital news portals represents a significant challenge in the digital era, as it may mislead the public and reduce trust in circulating information. The rapid and open nature of digital media enables unverified information to spread widely within a short period of time, while manual verification processes require substantial time and effort. This study proposes a semantic similarity-based approach to support real-time news verification using the Multilingual MPNet model. The proposed approach utilizes content text as input, followed by keyword extraction using KeyBERT to represent the core information of the news. The extracted keywords are employed in a news scraping process to obtain comparative news articles from digital news portals. A dataset consisting of 200 Indonesian news articles, including 100 factual news articles and 100 hoax news articles, was used for evaluation. Subsequently, semantic similarity measurement is conducted to evaluate the degree of semantic relevance between the test news and the scraped news. Evaluation metrics were applied to assess the effectiveness of the proposed approach. The findings demonstrate that semantic text representation using Multilingual MPNet effectively supports hoax detection and provides relevant supporting evidence in the form of semantically related news articles, enabling users to access comparative news sources that support the verification process. Experimental results show that the proposed approach achieved an accuracy of 83.5%, precision of 97.18%, recall of 69.0%, F1-score of 80.70%, and an AUC of 0.695, indicating that Multilingual MPNet can effectively support news verification through semantic similarity analysis.
Analysis of Hoax News Propagation Patterns on Threads Using GCN-Based Structural Node Embedding Andi Novia Nuzul Qurani; Sitti Zuhriyah; M. Adnan Nur
TEPIAN Vol. 7 No. 3 (2026): September 2026
Publisher : Politeknik Pertanian Negeri Samarinda

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.51967/tepian.v7i3.3929

Abstract

The dissemination of hoax news on social media has become a significant issue because it can influence public opinion and accelerate the spread of misinformation. Although various approaches have been proposed for hoax detection, most of them rely on textual analysis and do not adequately capture the structural relationships among users. Therefore, this study aims to analyze the propagation patterns of hoax news on the Threads social media platform using the Graph Convolutional Network (GCN) method. Data was collected by crawling posts from the Threads platform using hoax-related keywords and represented as a social graph in which nodes represent user accounts, mentions, and hashtags, while edges represent interactions among users. The constructed graph consists of 425 nodes and 697 edges. The proposed GCN model was evaluated based on degree centrality, betweenness centrality, clustering coefficient, and Spearman correlation analysis. The highest clustering coefficient obtained was 0.8046, while the highest Spearman correlation coefficient between the GCN score and degree centrality reached 0.6118. The results demonstrate that GCN effectively represents user relationships, identifies influential nodes, and provides a comprehensive understanding of hoax propagation patterns on the Threads platform.