Purpose - This study analyzes the classification performance and computational efficiency of five pretrained Convolutional Neural Network (CNN) architectures for identifying South Kalimantan traditional food images as an expanded benchmark. Design/methods/approach - The models were trained on a curated traditional food image dataset using a two-stage transfer learning strategy consisting of linear probing and full fine-tuning with frozen Batch Normalization, supported by a multi-technique data augmentation pipeline. Evaluation was conducted under a fixed stratified data-splitting scenario across repeated runs with different random seeds to assess model stability and reproducibility. Findings - EfficientNetV2B0 achieved the strongest overall performance among the evaluated architectures and provided the most favorable balance between classification accuracy and computational efficiency. Its performance was comparable to other high-performing CNN architectures while requiring substantially lower training time than deeper residual and hybrid networks. The results indicate that greater architectural complexity does not necessarily translate into better recognition performance for a relatively small traditional food image dataset. Research implications/limitations - The findings provide practical guidance for selecting efficient CNN architectures for traditional food recognition. However, the evaluation was conducted on a curated dataset under controlled conditions, and the absence of an ablation study prevents isolating the individual effects of data augmentation and two-stage fine-tuning. Originality/value - This expanded benchmark highlights the critical trade-off between reliable classification performance and computational cost, offering practical guidance for selecting efficient deep learning models to support the digital preservation of culinary heritage.