paper-with-me

홈 › Papers

Paraphrasing Is All You Need for Novel Object Captioning

2022-09-25 · Cheng-Fu Yang, Yao-Hung Hubert Tsai, Wan-Cyuan Fan, Ruslan Salakhutdinov, Louis-Philippe Morency, Yu-Chiang Frank Wang

Novel object captioning (NOC) aims to describe images containing objects without observing their ground truth captions during training. Due to the absence of caption annotation, captioning models cannot be directly optimized via sequence-to-sequence training or CIDEr optimization. As a result, we present Paraphrasing-to-Captioning (P2C), a two-stage learning framework for NOC, which would heuristically optimize the output captions via paraphrasing. With P2C, the captioning model first learns paraphrasing from a language model pre-trained on text-only corpus, allowing expansion of the word bank for improving linguistic fluency. To further enforce the output caption sufficiently describing the visual content of the input image, we perform self-paraphrasing for the captioning model with fidelity and adequacy objectives introduced. Since no ground truth captions are available for novel object images during training, our P2C leverages cross-modality (image-text) association modules to ensure the above caption characteristics can be properly preserved. In the experiments, we not only show that our P2C achieves state-of-the-art performances on nocaps and COCO Caption datasets, we also verify the effectiveness and flexibility of our learning framework by replacing language and cross-modality association models for NOC. Implementation details and code are available in the supplementary materials.

📄 PDF Abstract BibTeX arXiv:2209.12343

Code (0)

등록된 구현이 없습니다.

Tasks

AllLanguage ModellingObject

Similar Papers 제목 키워드 기반

iParaphrasing: Extracting Visually Grounded Paraphrases via an Image

2018-06-12 · COLING 2018 8 · Chenhui Chu, Mayu Otani, Yuta Nakashima

A paraphrase is a restatement of the meaning of a text in other words. Paraphrases have been studied to enhance the performance of many natural language processing tasks. In this paper, we propose a novel task iParaphras…

Image CaptioningQuestion AnsweringVisual Question AnsweringVisual Question Answering (VQA)

Visual Information Guided Zero-Shot Paraphrase Generation

2022-01-22 · COLING 2022 10 · Zhe Lin, Xiaojun Wan

Zero-shot paraphrase generation has drawn much attention as the large-scale high-quality paraphrase corpus is limited. Back-translation, also known as the pivot-based method, is typical to this end. Several works leverag…

DiversityImage CaptioningParaphrase GenerationTranslation

SLAM-AAC: Enhancing Audio Captioning with Paraphrasing Augmentation and CLAP-Refine through LLMs

2024-10-12 · Wenxi Chen, Ziyang Ma, Xiquan Li, Xuenan Xu 외

Automated Audio Captioning (AAC) aims to generate natural textual descriptions for input audio signals. Recent progress in audio pre-trained models and large language models (LLMs) has significantly enhanced audio unders…

AudioCapsAudio captioningCaption GenerationMachine Translation+4

ProtAugment: Intent Detection Meta-Learning through Unsupervised Diverse Paraphrasing

2021-08-01 · ACL 2021 5 · Thomas Dopierre, Christophe Gravier, Wilfried Logerais

Recent research considers few-shot intent detection as a meta-learning problem: the model is learning to learn from a consecutive set of small tasks named episodes. In this work, we propose ProtAugment, a meta-learning a…

DiversityIntent DetectionLanguage ModelingLanguage Modelling+1

ProtAugment: Unsupervised diverse short-texts paraphrasing for intent detection meta-learning

2021-05-27 · Thomas Dopierre, Christophe Gravier, Wilfried Logerais

Recent research considers few-shot intent detection as a meta-learning problem: the model is learning to learn from a consecutive set of small tasks named episodes. In this work, we propose ProtAugment, a meta-learning a…

DiversityIntent DetectionLanguage ModelingLanguage Modelling+1