paper-with-me

Image Captioning

33개 벤치마크 · 논문 2,086편 · 이 태스크의 논문 보기 →

Benchmarks

VizWiz 2020 test-dev

결과 50개

COCO Captions

결과 41개

nocaps in-domain

결과 41개

nocaps near-domain

결과 40개

nocaps out-of-domain

결과 40개

nocaps entire

결과 39개

VizWiz 2020 test

결과 13개

nocaps-XD entire

결과 12개

TextCaps 2020

결과 11개

nocaps-XD in-domain

결과 11개

nocaps-XD near-domain

결과 11개

nocaps-val-in-domain

결과 11개

nocaps-val-overall

결과 11개

nocaps-val-near-domain

결과 10개

nocaps-val-out-domain

결과 10개

SCICAP

결과 9개

WHOOPS!

결과 6개

Object HalBench

결과 3개

nocaps val

결과 3개

COCO Captions test

결과 2개

Conceptual Captions

결과 2개

FlickrStyle10K

결과 2개

Localized Narratives

결과 2개

MS-COCO

결과 2개

AIC-ICC

결과 1개

ChEBI-20

결과 1개

IU X-Ray

결과 1개

MSCOCO

결과 1개

Peir Gross

결과 1개

Most implemented

Papers

A Glance Is All You Need: Single-Pass Fine-Grained Image Captioning with SimLoss

2026-09-01 · Suryaansh Jain, Rahasya Barkur, Vishal G, Ryan Rossi 외 hf

An image may be worth a thousand words, but most captioning models describe it in only a few. Modern vision-language models produce fluent high-level captions, yet routinely miss the attributes, counts, textures, materia…

Image Captioning

Mapping Written Words to Spoken Words in a Different Language Using Only Visual Grounding

2026-08-27 · Gabriel Pirlogeanu, Dan Oneata, Horia Cucu, Herman Kamper arxiv

In many low-resource settings, even just eliciting speech for data collection is difficult. One promising approach has been to ask speakers to describe images. But how do we build models from such visually grounded speec…

Visual GroundingImage CaptioningKeyword Spotting

Overview of SHROOM-Visions 2026: A Shared Task on Hallucination Detection in Large Vision-Language Models

2026-08-26 · Raúl Vázquez, Aman Sinha, Chuyuan Li, Claudio Savelli 외 arxiv

In 2026, we held the fourth iteration of the SHROOM Shared Task series: SHROOM-Visions (\textbf{S}hared-task on \textbf{H}allucinations and \textbf{R}elated \textbf{O}bservable \textbf{O}vergeneration \textbf{M}istakes i…

Image CaptioningText Generation

Re$^3$Cap: Retrieval-Guided Refinement for Image Captioning Enhancement via Reinforcement Learning

2026-08-21 · Haonan Jia, Shichao Dong, Zenghui Sun, Jiawen Zheng 외 arxiv

Reinforcement Learning (RL) has demonstrated significant gains in image captioning, yet it is still limited in encouraging Large Vision-Language Models (LVLMs) to explore novel reasoning strategies. This limitation leads…

Reinforcement LearningImage Captioning

Is Multimodal Speculative Decoding Ready for Diffusion-Based Parallel Drafting? A Survey and Empirical Diagnosis

2026-08-21 · Yantao Li, Huanlin Gao, Fang Zhao, Chao Tan 외 arxiv

Speculative decoding accelerates autoregressive generation by allowing a lightweight drafter to propose future tokens while a target model verifies them in parallel. Its lossless guarantee has motivated a line of work th…

Visual ReasoningImage Captioning

Towards Clinically Faithful Medical Image Captioning via Enhanced Vision-Language Alignment

2026-08-20 · Yunseo Lee, Hyun Jun Kim, Heeseung Shin, Changwon Lim arxiv

Medical image captioning is a technique that accelerates early-stage diagnostic workflows and enhances the interpretability of medical diagnostic AI systems. However, unlike general image captioning, clinically reliable …

Image CaptioningType prediction

전체 2,086편 보기 →