Pre-trained vision-language models such as BLIP have achieved remarkable success in general image captioning tasks. However, their performance on domain-specific applications, particularly cultural heritage documentation, remains limited due to the lack of specialized knowledge and the inability to handle multi-label cultural categories. Full fine-tuning of these large models is computationally expensive and risks catastrophic forgetting, while standard adapter-based methods treat all images uniformly without considering domain-specific class characteristics. This study proposes ML-CAA-BLIP (Multi-Label Cultural-Aware Adapter for BLIP), a novel parameter-efficient adaptation method for Balinese carving image captioning. The proposed method introduces class-specific scaling parameters for each cultural motif category (Barong, Punggel, Keketusan, Gajah, Goak, Cina, and Daun) and employs a learned importance-weighted fusion mechanism to handle multi-label inputs where images contain multiple artistic styles. Experiments conducted on the BaliCarving dataset comprising 2,181 images demonstrate that ML-CAA-BLIP achieves the best BLEU-4 score of 0.2718 (+52.4% improvement over Base BLIP) and ROUGE-L score of 0.5835 (+15.2% improvement) while adding only 903 trainable parameters. The model also shows competitive performance on other metrics including METEOR and BERTScore. These results indicate that cultural-aware adaptation significantly improves domain-specific image captioning while maintaining parameter efficiency, contributing to the digital preservation of Balinese cultural heritage
Copyrights © 2026