Large Language Models (LLMs) offer substantial potential for healthcare information support, yet concerns remain regarding privacy, semantic reliability, and harmful language generation. This study developed and evaluated Syarief Med AI, a locally deployed healthcare chatbot integrating a quantized seven-billion-parameter LLM with a Flask backend and React frontend. The system incorporated safety-oriented prompt design, input screening, and post-response filtering, while all inference and conversation storage were performed locally. Evaluation used 30 Indonesian healthcare-related questions, with BERTScore F1 measuring semantic similarity and a multi-attribute toxicity classifier assessing linguistic safety. The chatbot achieved a mean BERTScore F1 of 0.727, indicating measurable semantic correspondence with the reference answers. Toxicity-related scores were generally low, with a maximum overall toxicity score of 0.395403 and severe toxicity remaining near zero. These findings demonstrate the feasibility of combining local LLM deployment with complementary semantic and linguistic-safety evaluation. However, the results do not establish factual accuracy, clinical validity, or medical safety, and further expert validation and user-based evaluation are required.
Copyrights © 2026