paper-with-me

Papers

Crowd Video Captioning

2019-11-13 · Liqi Yan, Mingjian Zhu, Changbin Yu

Describing a video automatically with natural language is a challenging task in the area of computer vision. In most cases, the on-site situation of great events is reported in news, but the situation of the off-site spectators in the entrance and exit is neglected which also arouses people's interest. Since the deployment of reporters in the entrance and exit costs lots of manpower, how to automatically describe the behavior of a crowd of off-site spectators is significant and remains a problem. To tackle this problem, we propose a new task called crowd video captioning (CVC) which aims to describe the crowd of spectators. We also provide baseline methods for this task and evaluate them on the dataset WorldExpo'10. Our experimental results show that captioning models have a fairly deep understanding of the crowd in video and perform satisfactorily in the CVC task.

📄 PDF Abstract BibTeX arXiv:1911.05449

Code (0)

등록된 구현이 없습니다.

Tasks

Video Captioning

Similar Papers 제목 키워드 기반

Dense-Captioning Events in Videos

2017-05-02 · ICCV 2017 10 · Ranjay Krishna, Kenji Hata, Frederic Ren, Li Fei-Fei 외

Most natural videos contain numerous events. For example, in a video of a "man playing a piano", the video might also contain "another man dancing" or "a crowd clapping". We introduce the task of dense-captioning events,…

Dense CaptioningRetrievalVideo Retrieval

Evaluation of Automatic Video Captioning Using Direct Assessment

2017-10-29 · Yvette Graham, George Awad, Alan Smeaton

We present Direct Assessment, a method for manually assessing the quality of automatically-generated captions for video. Evaluating the accuracy of video captions is particularly difficult because for any given video cli…

Caption GenerationMachine TranslationTranslationVideo Captioning

On the effectiveness of task granularity for transfer learning

2018-04-24 · Farzaneh Mahdisoltani, Guillaume Berger, Waseem Gharbieh, David Fleet 외

We describe a DNN for video classification and captioning, trained end-to-end, with shared features, to solve tasks at different levels of granularity, exploring the link between granularity in a source task and the qual…

ClassificationDiversityGeneral ClassificationTransfer Learning+1

Crowdsourcing and Evaluating Text-Based Audio Retrieval Relevances

2023-06-16 · Huang Xie, Khazar Khorrami, Okko Räsänen, Tuomas Virtanen

This paper explores grading text-based audio retrieval relevances with crowdsourcing assessments. Given a free-form text (e.g., a caption) as a query, crowdworkers are asked to grade audio clips using numeric scores (bet…

Audio captioningContrastive LearningRetrieval

Text Alignment for Real-Time Crowd Captioning

2013-06-01 · NAACL 2013 6 · Iftekhar Naim, Daniel Gildea, Walter Lasecki, Jeffrey P. Bigham
Speech Recognition