Author

Publication Date

Spring 2026

Degree Type

Master's Project

Degree Name

Master of Science in Computer Science (MSCS)

First Advisor

Navrati Saxena

Second Advisor

Amith Kamath Belman

Third Advisor

Vuthea Chheang

Keywords

bilingual video captioning, computer vision, deep learning, image captioning, multimodal learning, natural language processing, video captioning

Abstract

This report studies automatic caption generation in two settings: static images and short bilingual video clips. For image captioning, we present VIBE-CAP, an evaluation framework that compares pretrained vision-language captioners---GIT-Large (zero-shot), BLIP-Large (fine-tuned), and a late fusion of their outputs---against a retrieval-generation hybrid baseline on the MSCOCO Karpathy split using standard COCOEvalCap metrics. The BLIP-Large and fusion model exceed the baseline on BLEU, METEOR, ROUGE-L, and CIDEr scores, while GIT-Large also improves over the baseline but has small improvements. We then extend this to bilingual video captioning in English and Hindi on the MSR-VTT dataset. In the bilingual video-captioning framework consisting of frozen CLIP frame embeddings, a compact trainable Q-Former bridge, and partially unfrozen multilingual decoders mBART-50 and mT5-large are compared under a shared preprocessing and decoding protocol. Quantitative results of our models show higher scores when compared to a public Hindi LSTM baseline, with a decoder trade-off where mT5-large is stronger on overlap oriented metrics across languages and mBART-50 is more robust on Hindi CIDEr scores. This study shows how modern pretrained multimodal models can outperform classical retrieval heavy pipelines on standard benchmarks while also considering the evaluation of low-resource languages like Hindi.

Available for download on Saturday, May 22, 2027

Share

COinS