Ilham Syarief Roem Mohamad
School of Industrial Engineering, Telkom University, Bandung, Indonesia

Published : 1 Documents Claim Missing Document
Claim Missing Document
Check
Articles

Found 1 Documents
Search

Development and Evaluation of a Local Large Language Model-Based Healthcare Chatbot Using BERTScore and Toxicity Analysis Ilham Syarief Roem Mohamad; Fikri Mochamad Faizal
JUSIFO : Jurnal Sistem Informasi Vol 12 No 1 (2026): June
Publisher : Program Studi Sistem Informasi, Fakultas Sains dan Teknologi, Universitas Islam Negeri Raden Fatah Palembang

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.19109/jusifo.v12i1.35462

Abstract

Large Language Models (LLMs) offer substantial potential for healthcare information support, yet concerns remain regarding privacy, semantic reliability, and harmful language generation. This study developed and evaluated Syarief Med AI, a locally deployed healthcare chatbot integrating a quantized seven-billion-parameter LLM with a Flask backend and React frontend. The system incorporated safety-oriented prompt design, input screening, and post-response filtering, while all inference and conversation storage were performed locally. Evaluation used 30 Indonesian healthcare-related questions, with BERTScore F1 measuring semantic similarity and a multi-attribute toxicity classifier assessing linguistic safety. The chatbot achieved a mean BERTScore F1 of 0.727, indicating measurable semantic correspondence with the reference answers. Toxicity-related scores were generally low, with a maximum overall toxicity score of 0.395403 and severe toxicity remaining near zero. These findings demonstrate the feasibility of combining local LLM deployment with complementary semantic and linguistic-safety evaluation. However, the results do not establish factual accuracy, clinical validity, or medical safety, and further expert validation and user-based evaluation are required.