Purpose: Sign language recognition systems based on 3D hand keypoints frequently experience generalization issues when trained on limited and homogeneous datasets, particularly under single-subject data collection settings. In BISINDO alphabet recognition, this limitation often leads to significant performance degradation when models are applied to unseen users or different acquisition devices. This study aims to improve cross-domain generalization of BISINDO alphabet recognition models by introducing realistic feature-level augmentation applied directly to 3D hand keypoints. Methods: A realistic 3D keypoint augmentation framework was proposed, consisting of Gaussian Jitter, Anisotropic Scale, Bone Length Scale, and Depth & Tilt Jitter to simulate sensor noise, anatomical variability, and viewpoint changes. Hand keypoints were extracted using MediaPipe Hands and classified using a multilayer perception (MLP). Model performance was evaluated through k-fold cross-validation on a single-subject internal dataset and cross-domain testing on an external dataset involving unseen subjects and different acquisition devices. Results: The experimental results indicate that the proposed augmentation strategy substantially improves generalization performance without degrading in-domain accuracy. The cross-domain F1-score increased from 67.20% in the baseline model to 89.55% after applying realistic 3D keypoint augmentation, while performance variability across validation folds was also reduced, indicating more stable learning behavior. Novelty: This work highlights that controlled geometric manipulation at the 3D keypoint level provides an effective and computationally efficient approach to mitigating overfitting in low-resource BISINDO recognition scenarios. By focusing on feature-level augmentation rather than image-based transformations or algorithm replacement, this study offers a practical strategy for enhancing robustness in real-world sign language recognition systems.
Copyrights © 2026