Claim Missing Document
Check
Articles

Found 2 Documents
Search

The Determining Gender Using Facial Recognition Based On Neural Network With Backpropagation Fauziah Fauziah
Data Science: Journal of Computing and Applied Informatics Vol. 2 No. 1 (2018): Data Science: Journal of Computing and Applied Informatics (JoCAI)
Publisher : Talenta Publisher

Show Abstract | Download Original | Original Source | Check in Google Scholar | Full PDF (1248.839 KB) | DOI: 10.32734/jocai.v2.i1-96

Abstract

One area of science that can apply facial recognition applications is artificial intelligence. The algorithms used in facial recognition are quite numerous and varied, but they all have the same three basic stages, face detection, facial extraction and facial recognition (Face Recognition) . Facial recognition applications using artificial intelligence as a major component, especially artificial neural networks for processing and facial identification are still not widely encountered. Ba ckpropagation is a learning algorithm to minimize the error rate by adjusting the weights based on the desired output and target differences. The test results of 30 images have the average value of mse is 0.14796 and the best value of mse on the test of man number 3 with mse value 0.1488 and mean 0.0047 while for the female number 2 with mse value 0.1497 and niali mean 0.0047.
Image captioning bilingual berbasis BLIP-2 dan MarianMT dengan integrasi text-to-speech untuk literasi visual siswa Rizka Oktaviani; Fauziah Fauziah
Jurnal Ilmiah Teknologi dan Rekayasa Vol. 31 No. 2 (2026)
Publisher : Universitas Gunadarma

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.35760/tr.2026.v31i2.152

Abstract

Visual literacy is the ability to understand and interpret information from visual representations, which involves linguistic processing. In the context of elementary education, this ability is crucial for helping students associate visual objects with bilingual vocabulary. However, most image captioning research remains monolingual, few studies accommodate Indonesian, and none have integrated audio representations into a single multimodal learning system. This study develops a bilingual image captioning system based on BLIP-2 and MarianMT, integrated with Text-to-Speech (TTS) within a FastAPI-based web application. English captions are generated using a pretrained BLIP-2 model, then translated into Indonesian using MarianMT, and converted to audio using gTTS. Additionally, a QLoRA fine-tuning experiment was conducted to compare model performance. The dataset consists of 6400 animal images relevant to the elementary school learning context, with caption quality evaluated using the METEOR metric. The results show that the pretrained BLIP-2 model delivers relatively stable performance with a METEOR score of 0.3765 for English and 0.3295 for Indonesian. These scores indicate that the generated captions are semantically relevant, although the overall performance is still relatively moderate compared to advanced image captioning systems. Functional testing of the prototype involving five elementary school students showed that the system is capable of generating bilingual captions and audio in real time and is easy to use. This multimodal integration supports the visual–verbal association process and has the potential to enrich students’ bilingual vocabulary, although no controlled experimental testing has yet been conducted to quantitatively measure improvements in literacy.