Background: Most lip-reading studies primarily utilize image-based representations of lip movements, offering extensive visual data while imposing significant computational burdens. Lip landmark-based representations are still not well understood, even though they could be a better way to describe lip dynamics in a smaller and more efficient way. This limitation is even more apparent in Bahasa lip-reading research, where there are few studies and computationally efficient solutions remain essential. Objective: This study examines the shortcomings of image-based lip-reading methods by leveraging lip landmarks as a concise and computationally efficient input representation. The proposed method is tested on the IndoLR open dataset, which contains video data of lip-reading in Bahasa. Methods: In this study, video sequences were transformed into coordinate-based landmark data to minimize computational demands while preserving critical information regarding lip dynamics. An attention-based BiLSTM model was trained to group10 word classes and 4 phrase classes using this dataset. Results: The model achieved accuracies of 93.03% for word classification and 95.18% for phrase classification. The approach also maintained high efficiency, with average inference times of 0.000530 and 0.011366 s per sample and computational costs of only 0.01 and 0.14 GFLOPs, respectively. Conclusion: These results show how well lip landmarks can be combined with a lightweight deep learning model with very few resources. This study makes a significant contribution to research on lip-reading in Bahasa and lays the groundwork for future studies that will use larger and more diverse datasets. Keywords: Attention, BiLSTM, Bahasa, Landmark, Lip-Reading
Copyrights © 2026