Image Captioning
33개 벤치마크 · 논문 2,086편 · 이 태스크의 논문 보기 →
Benchmarks
VizWiz 2020 test-dev
COCO Captions
nocaps in-domain
nocaps near-domain
nocaps out-of-domain
nocaps entire
VizWiz 2020 test
nocaps-XD entire
TextCaps 2020
nocaps-XD in-domain
nocaps-XD near-domain
nocaps-XD out-of-domain
nocaps-val-in-domain
nocaps-val-overall
nocaps-val-near-domain
nocaps-val-out-domain
SCICAP
Flickr30k Captions test
WHOOPS!
Object HalBench
nocaps val
COCO Captions test
Conceptual Captions
FlickrStyle10K
Localized Narratives
MS-COCO
AIC-ICC
BanglaLekhaImageCaptions
ChEBI-20
IU X-Ray
MSCOCO
Peir Gross
Most implemented
Show, Attend and Tell: Neural Image Caption Generation with Visual Attention
Show and Tell: A Neural Image Caption Generator
Bottom-Up and Top-Down Attention for Image Captioning and Visual Question Answering
Self-critical Sequence Training for Image Captioning
CutMix: Regularization Strategy to Train Strong Classifiers with Localizable Features
Diverse Beam Search: Decoding Diverse Solutions from Neural Sequence Models
Papers
A Glance Is All You Need: Single-Pass Fine-Grained Image Captioning with SimLoss
An image may be worth a thousand words, but most captioning models describe it in only a few. Modern vision-language models produce fluent high-level captions, yet routinely miss the attributes, counts, textures, materia…
Image CaptioningMapping Written Words to Spoken Words in a Different Language Using Only Visual Grounding
In many low-resource settings, even just eliciting speech for data collection is difficult. One promising approach has been to ask speakers to describe images. But how do we build models from such visually grounded speec…
Visual GroundingImage CaptioningKeyword SpottingOverview of SHROOM-Visions 2026: A Shared Task on Hallucination Detection in Large Vision-Language Models
In 2026, we held the fourth iteration of the SHROOM Shared Task series: SHROOM-Visions (\textbf{S}hared-task on \textbf{H}allucinations and \textbf{R}elated \textbf{O}bservable \textbf{O}vergeneration \textbf{M}istakes i…
Image CaptioningText GenerationRe$^3$Cap: Retrieval-Guided Refinement for Image Captioning Enhancement via Reinforcement Learning
Reinforcement Learning (RL) has demonstrated significant gains in image captioning, yet it is still limited in encouraging Large Vision-Language Models (LVLMs) to explore novel reasoning strategies. This limitation leads…
Reinforcement LearningImage CaptioningIs Multimodal Speculative Decoding Ready for Diffusion-Based Parallel Drafting? A Survey and Empirical Diagnosis
Speculative decoding accelerates autoregressive generation by allowing a lightweight drafter to propose future tokens while a target model verifies them in parallel. Its lossless guarantee has motivated a line of work th…
Visual ReasoningImage CaptioningTowards Clinically Faithful Medical Image Captioning via Enhanced Vision-Language Alignment
Medical image captioning is a technique that accelerates early-stage diagnostic workflows and enhances the interpretability of medical diagnostic AI systems. However, unlike general image captioning, clinically reliable …
Image CaptioningType prediction