paper-with-me

홈 › Papers

The Color of the Cat is Gray: 1 Million Full-Sentences Visual Question Answering (FSVQA)

2016-09-21 · Andrew Shin, Yoshitaka Ushiku, Tatsuya Harada

Visual Question Answering (VQA) task has showcased a new stage of interaction between language and vision, two of the most pivotal components of artificial intelligence. However, it has mostly focused on generating short and repetitive answers, mostly single words, which fall short of rich linguistic capabilities of humans. We introduce Full-Sentence Visual Question Answering (FSVQA) dataset, consisting of nearly 1 million pairs of questions and full-sentence answers for images, built by applying a number of rule-based natural language processing techniques to original VQA dataset and captions in the MS COCO dataset. This poses many additional complexities to conventional VQA task, and we provide a baseline for approaching and evaluating the task, on top of which we invite the research community to build further improvements.

📄 PDF Abstract BibTeX arXiv:1609.06657

Code (0)

등록된 구현이 없습니다.

Tasks

Question AnsweringSentenceVisual Question AnsweringVisual Question Answering (VQA)

Similar Papers 제목 키워드 기반

Gnuastro: visualizing the full dynamic range in color images

2024-01-08 · Raúl Infante-Sainz, Mohammad Akhlaghi

Color plays a crucial role in the visualization, interpretation, and analysis of multi-wavelength astronomical images. However, generating color images that accurately represent the full dynamic range of astronomical sou…

Invertible Grayscale via Dual Features Ensemble

2020-05-22 · TAIZHONG YE, Yong Du, JUNJIE DENG, AND SHENGFENG HE

Grayscale image colorization is known as an ill-posed problem because of the imbalanced matching between intensity and color values. Even given prior hints about the original color image, existing colorization methods …

ColorizationImage Colorization

Saliency Preservation in Low-Resolution Grayscale Images

2017-12-06 · ECCV 2018 9 · Shivanthan A. C. Yohanandan, Adrian G. Dyer, DaCheng Tao, Andy Song

Visual salience detection originated over 500 million years ago and is one of nature's most efficient mechanisms. In contrast, many state-of-the-art computational saliency models are complex and inefficient. Most salienc…

One Channel to Rule Them All: Rethinking Input Representation for Visual Place Recognition

2026-05-31 · Timur Ismagilov, Shakaiba Majeed, Michael Milford, Tan Viet Tuyen Nguyen 외 arxiv

Visual Place Recognition (VPR) is fundamental to long-term robot localization and SLAM, yet current systems overwhelmingly rely on RGB input, implicitly assuming color is necessary for global place recognition. We challe…

Visual Place Recognition

SCGAN: Saliency Map-guided Colorization with Generative Adversarial Network

2020-11-23 · Yuzhi Zhao, Lai-Man Po, Kwok-Wai Cheung, Wing-Yin Yu 외

Given a grayscale photograph, the colorization system estimates a visually plausible colorful image. Conventional methods often use semantics to colorize grayscale images. However, in these methods, only classification s…

ColorizationDecoderGenerative Adversarial Network