This study investigates static sign language classification by examining three image-based input representations, namely RGB images, skeleton images, and RGB+skeleton fusion, using two datasets with different visual characteristics. The first dataset reflects a relatively controlled acquisition setting, whereas the second dataset contains more complex background and visual variations. As an initial baseline, eight pretrained convolutional neural network (CNN) architectures were evaluated, and two representative models were subsequently selected for more detailed analysis. The baseline evaluation indicates that ResNet50 and EfficientNetB0 achieve the most competitive performance when RGB images are used as input. Further analysis of input representations shows that skeleton images are highly effective, particularly on the more challenging dataset, while RGB+skeleton fusion does not consistently improve classification performance. The cross-dataset evaluation further reveals a considerable performance drop across all configurations, suggesting the presence of a strong domain shift between the two datasets. In the A-B scenario, EfficientNetB0 with RGB input yields the best results, while in the B-A scenario, EfficientNetB0 with skeleton input shows the most stable performance. These findings indicate that the most effective input representation depends on the direction of domain transfer and that high intra-dataset performance does not necessarily reflect good generalization capability.
Copyrights © 2026