paper-with-me

Papers

STAIR Captions: Constructing a Large-Scale Japanese Image Caption Dataset

2017-05-02 · ACL 2017 7 · Yuya Yoshikawa, Yutaro Shigeto, Akikazu Takeuchi

In recent years, automatic generation of image descriptions (captions), that is, image captioning, has attracted a great deal of attention. In this paper, we particularly consider generating Japanese captions for images. Since most available caption datasets have been constructed for English language, there are few datasets for Japanese. To tackle this problem, we construct a large-scale Japanese image caption dataset based on images from MS-COCO, which is called STAIR Captions. STAIR Captions consists of 820,310 Japanese captions for 164,062 images. In the experiment, we show that a neural network trained using STAIR Captions can generate more natural and better Japanese captions, compared to those generated using English-Japanese machine translation after generating English captions.

📄 PDF Abstract BibTeX arXiv:1705.00823

Code (1)

William-N-Havard/VGS-dataset-metadata

Tasks

Image CaptioningMachine TranslationTranslation

Similar Papers 제목 키워드 기반

JaSPICE: Automatic Evaluation Metric Using Predicate-Argument Structures for Image Captioning Models

2023-11-07 · Yuiga Wada, Kanta Kaneda, Komei Sugiura

Image captioning studies heavily rely on automatic evaluation metrics such as BLEU and METEOR. However, such n-gram-based metrics have been shown to correlate poorly with human evaluation, leading to the proposal of alte…

Image Captioning

Video Caption Dataset for Describing Human Actions in Japanese

2020-03-10 · LREC 2020 5 · Yutaro Shigeto, Yuya Yoshikawa, Jiaqing Lin, Akikazu Takeuchi

In recent years, automatic video caption generation has attracted considerable attention. This paper focuses on the generation of Japanese captions for describing human actions. While most currently available video capti…

Caption Generation

Cascaded Multilingual Audio-Visual Learning from Videos

2021-11-08 · Andrew Rouditchenko, Angie Boggust, David Harwath, Samuel Thomas 외

In this paper, we explore self-supervised audio-visual models that learn from instructional videos. Prior work has shown that these models can relate spoken words and sounds to visual content after training on a large-sc…

audio-visual learningRetrieval

Extended Abstract: Improving Vision-and-Language Navigation with Image-Text Pairs from the Web

2020-06-12 · ICML Workshop LaReL 2020 7 · Arjun Majumdar, Ayush Shrivastava, Stefan Lee, Peter Anderson 외

Following a navigation instruction such as 'Walk down the stairs and stop near the sofa' requires an agent to ground scene elements referenced via language (e.g.'stairs') to visual content in the environment (pixels corr…

Vision and Language Navigation

STAIR Actions: A Video Dataset of Everyday Home Actions

2018-04-12 · Yuya Yoshikawa, Jiaqing Lin, Akikazu Takeuchi

A new large-scale video dataset for human action recognition, called STAIR Actions is introduced. STAIR Actions contains 100 categories of action labels representing fine-grained everyday home actions so that it can be a…

Action RecognitionTemporal Action Localization