paper-with-me

Papers

iCap: Interactive Image Captioning with Predictive Text

2020-01-31 · Zhengxiong Jia, Xirong Li

In this paper we study a brand new topic of interactive image captioning with human in the loop. Different from automated image captioning where a given test image is the sole input in the inference stage, we have access to both the test image and a sequence of (incomplete) user-input sentences in the interactive scenario. We formulate the problem as Visually Conditioned Sentence Completion (VCSC). For VCSC, we propose asynchronous bidirectional decoding for image caption completion (ABD-Cap). With ABD-Cap as the core module, we build iCap, a web-based interactive image captioning system capable of predicting new text with respect to live input from a user. A number of experiments covering both automated evaluations and real user studies show the viability of our proposals.

📄 PDF Abstract BibTeX arXiv:2001.11782

Code (0)

등록된 구현이 없습니다.

Tasks

Image CaptioningSentenceSentence Completion

Methods 이 논문이 사용한 방법론

Test 설명 없음

Similar Papers 제목 키워드 기반

Five Years of SciCap: What We Learned and Future Directions for Scientific Figure Captioning

2025-12-25 · Ting-Hao 'Kenneth' Huang, Ryan A. Rossi, Sungchul Kim, Tong Yu 외 arxiv

Between 2021 and 2025, the SciCap project grew from a small seed-funded idea at The Pennsylvania State University (Penn State) into one of the central efforts shaping the scientific figure-captioning landscape. Supported…

SciCap+: A Knowledge Augmented Dataset to Study the Challenges of Scientific Figure Captioning

2023-06-06 · Zhishen Yang, Raj Dabre, Hideki Tanaka, Naoaki Okazaki

In scholarly documents, figures provide a straightforward way of communicating scientific findings to readers. Automating figure caption generation helps move model understandings of scientific documents beyond text and …

Caption GenerationImage CaptioningOptical Character Recognition (OCR)

AFRICAPTION: Establishing a New Paradigm for Image Captioning in African Languages

2025-10-20 · Mardiyyah Oduwole, Prince Mireku, Fatimo Adebanjo, Oluwatosin Olajide 외 arxiv

Multimodal AI research has overwhelmingly focused on high-resource languages, hindering the democratization of advancements in the field. To address this, we present AfriCaption, a comprehensive framework for multilingua…

Image Captioning

RubiCap: Rubric-Guided Reinforcement Learning for Dense Image Captioning

2026-03-10 · Tzu-Heng Huang, Sirajul Salekin, Javier Movellan, Frederic Sala 외 arxiv

Dense image captioning is critical for cross-modal alignment in vision-language pretraining and text-to-image generation, but scaling expert-quality annotations is prohibitively expensive. While synthetic captioning via …

Text-to-Image GenerationReinforcement LearningImage Captioning

OmniCaptioner: One Captioner to Rule Them All

2025-04-09 · Yiting Lu, Jiakang Yuan, Zhen Li, Shitian Zhao 외

We propose OmniCaptioner, a versatile visual captioning framework for generating fine-grained textual descriptions across a wide variety of visual domains. Unlike prior methods limited to specific image types (e.g., natu…

AllImage CaptioningImage GenerationText to Image Generation+2