Sarah S. Mohammed
Ibn Sina University for Medical and Pharmaceutical Sciences

Published : 1 Documents Claim Missing Document
Claim Missing Document
Check
Articles

Found 1 Documents
Search

Video summarization using deep image captioning models Qudes M. B. Aljelawy; Sarah S. Mohammed; Entessar K. Hanoun
Indonesian Journal of Electrical Engineering and Computer Science Vol 43, No 2: August 2026
Publisher : Institute of Advanced Engineering and Science

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.11591/ijeecs.v43.i2.pp547-554

Abstract

This research presents a novel approach for video summarization by leveraging deep image captioning models. A pretrained image captioning model, namely Salesforce's bootstrapped language image pretraining (BLIP), is used to extract keyframes from a video at regular intervals and produce natural language descriptions. These textual descriptions are then filtered for non-repetition and concatenated into a coherent summary, allowing users to understand the video’s content without viewing it in full. The proposed framework aims to improve video browsing, indexing, and retrieval efficiency, particularly for big datasets. The proposed method achieves a significant reduction in redundancy by 40% compared to raw captioning sequences. Evaluation using semantic consistency checks demonstrates that the BLIP-based framework maintains high descriptive accuracy even in complex scenes, providing a scalable solution for large-scale video indexing.