Packaged food labels may contain technical and multilingual ingredient terms that complicate preliminary screening for pork-derived non-halal substances. This study develops a web-based pipeline that integrates optical character recognition (OCR), automatic translation, and Long Short-Term Memory (LSTM) text classification. A public dataset of 528,092 labeled ingredient records, comprising 291,920 halal and 236,172 pork-related non-halal records, was used for model development. After text normalization, tokenization, and sequence padding, the data were divided into training, validation, and testing subsets using an 80:10:10 ratio. The final test set contained 52,810 records. The confusion matrix contained 29,182 true negatives, 12 false positives, 41 false negatives, and 23,575 true positives, corresponding to 99.90% accuracy, 99.95% precision, 99.83% recall, and a 99.89% F1-score. The web implementation accepts label images, extracts text, translates non-English content, and applies the trained classifier. The reported metrics evaluate the text classifier rather than the complete OCR-to-classification pipeline; therefore, the system should be treated as a preliminary screening tool and not as a substitute for formal halal certification.
Copyrights © 2026