VIBE-CAP: Vision-Language Baseline Enhancements for Image Captioning
Publication Date
1-1-2026
Document Type
Conference Proceeding
Publication Title
2026 1st International Conference on Emerging Trends in Advancements and Applications of Computational Intelligence Techniques Etaact 2026
DOI
10.1109/ETAACT69135.2026.11541427
Abstract
Image captioning has made significant progress from traditional methods, such as encoder-decoder frameworks, template-based, and retrieval-based approaches, to pretrained Vision-Language Models (VLMs). In this paper, we introduce VIBE-CAP, which is a unified evaluation framework comparing three pretrained models: GIT-Large, BLIP-Large, and the GIT-BLIP fusion model. This is to determine whether pretrained models can outperform the retrieval-augmented framework on the standard MSCOCO dataset with the Karpathy split. We used standard COCOEvalCap metrics to show that these models outperform the baseline: GIT-Large improves CIDEr from 1.280 to 1.311, and the Fusion model achieves a CIDEr of 2.491. BLIP-Large achieves the highest score among all models, with a BLEU-4 score of 0.992 and a CIDEr score of 2.704. These scores obtained by three of the models have surpassed the baseline scores across all metrics, such as BLEU-1, BLEU-4, CIDEr, METEOR, and ROUGE, proving that pretrained captioning models perform much better than traditional retrieval-based approaches with complex architectures.
Keywords
artificial intelligence, computer vision, deep learning, multilingual captioning, natural language processing, semantic evaluation
Department
Computer Science
Recommended Citation
Gitika Rath and Navrati Saxena. "VIBE-CAP: Vision-Language Baseline Enhancements for Image Captioning" 2026 1st International Conference on Emerging Trends in Advancements and Applications of Computational Intelligence Techniques Etaact 2026 (2026). https://doi.org/10.1109/ETAACT69135.2026.11541427