paper-with-me

Papers

Varying image description tasks: spoken versus written descriptions

2018-08-01 · COLING 2018 8 · Emiel van Miltenburg, Ruud Koolen, Emiel Krahmer

Automatic image description systems are commonly trained and evaluated on written image descriptions. At the same time, these systems are often used to provide spoken descriptions (e.g. for visually impaired users) through apps like TapTapSee or Seeing AI. This is not a problem, as long as spoken and written descriptions are very similar. However, linguistic research suggests that spoken language often differs from written language. These differences are not regular, and vary from context to context. Therefore, this paper investigates whether there are differences between written and spoken image descriptions, even if they are elicited through similar tasks. We compare descriptions produced in two languages (English and Dutch), and in both languages observe substantial differences between spoken and written descriptions. Future research should see if users prefer the spoken over the written style and, if so, aim to emulate spoken descriptions.

📄 PDF Abstract BibTeX

Code (1)

cltl/Spoken-versus-Written 공식 구현

Tasks

Image Description

Similar Papers 제목 키워드 기반

Show and Speak: Directly Synthesize Spoken Description of Images

2020-10-23 · Xinsheng Wang, Siyuan Feng, Jihua Zhu, Mark Hasegawa-Johnson 외

This paper proposes a new model, referred to as the show and speak (SAS) model that, for the first time, is able to directly synthesize spoken descriptions of images, bypassing the need for any text or phonemes. The basi…

Decoder

Transcription-Enriched Joint Embeddings for Spoken Descriptions of Images and Videos

2020-06-01 · Benet Oriol, Jordi Luque, Ferran Diego, Xavier Giro-i-Nieto

In this work, we propose an effective approach for training unique embedding representations by combining three simultaneous modalities: image and spoken and textual narratives. The proposed methodology departs from a ba…

Retrieval

DIDEC: The Dutch Image Description and Eye-tracking Corpus

2018-08-01 · COLING 2018 8 · Emiel van Miltenburg, {\'A}kos K{\'a}d{\'a}r, Ruud Koolen, Emiel Krahmer

We present a corpus of spoken Dutch image descriptions, paired with two sets of eye-tracking data: Free viewing, where participants look at images without any particular purpose, and Description viewing, where we track e…

Image DescriptionSpecificityTask 2

Learning Words by Drawing Images

2019-06-01 · CVPR 2019 6 · Didac Suris, Adria Recasens, David Bau, David Harwath 외

We propose a framework for learning through drawing. Our goal is to learn the correspondence between spoken words and abstract visual attributes, from a dataset of spoken descriptions of images. Building upon recent find…

Triplet

Speech-Based Visual Question Answering

2017-05-01 · Ted Zhang, Dengxin Dai, Tinne Tuytelaars, Marie-Francine Moens 외

This paper introduces speech-based visual question answering (VQA), the task of generating an answer given an image and a spoken question. Two methods are studied: an end-to-end, deep neural network that directly uses au…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Question Answeringspeech-recognition+3