paper-with-me

홈 › Papers

Cooperative image captioning

2019-07-26 · Gilad Vered, Gal Oren, Yuval Atzmon, Gal Chechik

When describing images with natural language, the descriptions can be made more informative if tuned using downstream tasks. This is often achieved by training two networks: a "speaker network" that generates sentences given an image, and a "listener network" that uses them to perform a task. Unfortunately, training multiple networks jointly to communicate to achieve a joint task, faces two major challenges. First, the descriptions generated by a speaker network are discrete and stochastic, making optimization very hard and inefficient. Second, joint training usually causes the vocabulary used during communication to drift and diverge from natural language. We describe an approach that addresses both challenges. We first develop a new effective optimization based on partial-sampling from a multinomial distribution combined with straight-through gradient updates, which we name PSST for Partial-Sampling Straight-Through. Second, we show that the generated descriptions can be kept close to natural by constraining them to be similar to human descriptions. Together, this approach creates descriptions that are both more discriminative and more natural than previous approaches. Evaluations on the standard COCO benchmark show that PSST Multinomial dramatically improve the recall@10 from 60% to 86% maintaining comparable language naturalness, and human evaluations show that it also increases naturalness while keeping the discriminative power of generated captions.

📄 PDF Abstract BibTeX arXiv:1907.11565

Code (0)

등록된 구현이 없습니다.

Tasks

Image Captioning

Similar Papers 제목 키워드 기반

Alleviating Noisy Data in Image Captioning with Cooperative Distillation

2020-12-21 · Pierre Dognin, Igor Melnyk, Youssef Mroueh, Inkit Padhi 외

Image captioning systems have made substantial progress, largely due to the availability of curated datasets like Microsoft COCO or Vizwiz that have accurate descriptions of their corresponding images. Unfortunately, sca…

Image Captioning

Non-Autoregressive Image Captioning with Counterfactuals-Critical Multi-Agent Learning

2020-05-10 · Longteng Guo, Jing Liu, Xinxin Zhu, Xingjian He 외

Most image captioning models are autoregressive, i.e. they generate each word by conditioning on previously generated words, which leads to heavy latency during inference. Recently, non-autoregressive decoding has been p…

Image CaptioningMachine TranslationMulti-agent Reinforcement LearningSentence+1

Emergent Natural Language with Communication Games for Improving Image Captioning Capabilities without Additional Data

2025-07-11 · Parag Dutta, Ambedkar Dukkipati arxiv

Image captioning is an important problem in developing various AI systems, and these tasks require large volumes of annotated images to train the models. Since all existing labelled datasets are already used for training…

Multi-agent Reinforcement LearningImage Captioning

DRAMA: Joint Risk Localization and Captioning in Driving

2022-09-22 · Srikanth Malla, Chiho Choi, Isht Dwivedi, Joon Hee Choi 외

Considering the functionality of situational awareness in safety-critical automation systems, the perception of risk in driving scenes and its explainability is of particular importance for autonomous and cooperative dri…

Image Captioning

Zero-Resource Neural Machine Translation with Multi-Agent Communication Game

2018-02-09 · Yun Chen, Yang Liu, Victor O. K. Li

While end-to-end neural machine translation (NMT) has achieved notable success in the past years in translating a handful of resource-rich language pairs, it still suffers from the data scarcity problem for low-resource …

DecoderImage CaptioningImage DescriptionMachine Translation+2