Bahasa Isyarat Indonesia atau BISINDO merupakan bahasa isyarat yang lahir dari komunitas Tuli yang ada di Indonesia. Gerakan tangan, tubuh, hingga ekspresi wajah merupakan bagian terpenting dalam penggunaan bahasa isyarat. Pada penelitian ini, dilakukan proses rekognisi gerakan bahasa isyarat berdasarkan gerakan tangan, tubuh, dan wajah dengan menggunakan metode LSTM. Dataset yang digunakan merupakan dataset primer yang terdiri dari 100 kosa isyarat yang dibuat berdasarkan dialek Denpasar, Provinsi Bali. Penelitian ini tidak hanya berfokus pada proses rekognisi, tetapi juga mengevaluasi pengaruh variasi jumlah frame terhadap performa model melalui penerapan metode sampling frame. Pemilihan jumlah frame dilakukan dengan memilih representasi frame sebanyak 15, 20, 25, dan 30 frame. Keempat jumlah sampling frame tersebut kemudian dilatih dan diuji pada model LSTM. Hasil penelitian ini menunjukkan hasil sampling frame optimal adalah pada 30 frames yang memperoleh hasil pelatihan dengan nilai loss validasi terkecil pada 7,8% dan hasil pengujian WER terbaik pada 13% jika dibandingkan dengan hasil sampling frame lainnya. Berdasarkan hasil penelitian ini, diharapkan model LSTM yang telah diperoleh dan diimplementasikan ke dalam prototipe sistem dapat digunakan oleh masyarakat umum dalam menjembatani komunikasi masyarakat umum dengan komunitas Tuli serta oleh peneliti selanjutnya dalam pengembangan lebih lanjut proses rekognisi gerakan Bahasa Isyarat Indonesia. Abstract Indonesian Sign Language or BISINDO is a natural sign language that emerged from the Deaf community in Indonesia. Hand movements, body posture, and facial expressions play essential roles in conveying meaning within sign language communication. In this study, sign language gesture recognition is performed based on hand, body, and facial movements using the Long Short-Term Memory (LSTM) method. The dataset used is a primary dataset consisting of 100 sign vocabulary items constructed based on the Denpasar dialect in Bali Province. This study does not solely focus on the recognition process, but also evaluates the effect of varying the number of frames on model performance through the application of a frame sampling method. The selection of frame quantities was conducted by choosing representative samples of 15, 20, 25, and 30 frames. These four sampling configurations were subsequently trained and tested using an LSTM model. The results of this study indicate that the optimal frame sampling configuration is 30 frames, which achieved the lowest validation loss of 7.8% during training and the best Word Error Rate (WER) of 13% during testing, compared to the other sampling configurations. Based on these findings, it is expected that the LSTM model developed and implemented within a prototype system can be utilized by the general public to help bridge communication between the general public and the Deaf community and by future researchers for the further development of Indonesian Sign Language gesture recognition processes.
Copyrights © 2026