paper-with-me

홈 › Papers

Visuo-Linguistic Question Answering (VLQA) Challenge

2020-05-01 · Findings of the Association for Computational Linguistics 2020 · Shailaja Keyur Sampat, Yezhou Yang, Chitta Baral

Understanding images and text together is an important aspect of cognition and building advanced Artificial Intelligence (AI) systems. As a community, we have achieved good benchmarks over language and vision domains separately, however joint reasoning is still a challenge for state-of-the-art computer vision and natural language processing (NLP) systems. We propose a novel task to derive joint inference about a given image-text modality and compile the Visuo-Linguistic Question Answering (VLQA) challenge corpus in a question answering setting. Each dataset item consists of an image and a reading passage, where questions are designed to combine both visual and textual information i.e., ignoring either modality would make the question unanswerable. We first explore the best existing vision-language architectures to solve VLQA subsets and show that they are unable to reason well. We then develop a modular method with slightly better baseline performance, but it is still far behind human performance. We believe that VLQA will be a good benchmark for reasoning over a visuo-linguistic context. The dataset, code and leaderboard is available at https://shailaja183.github.io/vlqa/.

📄 PDF Abstract BibTeX arXiv:2005.00330

Code (1)

shailaja183/vlqa pytorch

Tasks

Question AnsweringReading ComprehensionVisual Question Answering (VQA)

Similar Papers 제목 키워드 기반

VLQA: The First Comprehensive, Large, and High-Quality Vietnamese Dataset for Legal Question Answering

2025-07-26 · Tan-Minh Nguyen, Hoang-Trung Nguyen, Trong-Khoi Dao, Xuan-Hieu Phan 외 arxiv

The advent of large language models (LLMs) has led to significant achievements in various domains, including legal text processing. Leveraging LLMs for legal tasks is a natural evolution and an increasingly compelling ch…

Information RetrievalQuestion Answering

Research on Vision-Language Question Answering Models for Industrial Robots

2026-05-02 · Ping Li, Bartlomiej Brzozka arxiv

A hierarchical cross-modal fusion model is proposed for vision-language question answering (VLQA) in industrial robotics, targeting the challenges of semantic ambiguity, complex environmental layouts, and domain-specific…

Question AnsweringAnomaly DetectionObject Detection

iVQA: Inverse Visual Question Answering

2017-10-10 · CVPR 2018 6 · Feng Liu, Tao Xiang, Timothy M. Hospedales, Wankou Yang 외

We propose the inverse problem of Visual question answering (iVQA), and explore its suitability as a benchmark for visuo-linguistic understanding. The iVQA task is to generate a question that corresponds to a given image…

Question AnsweringQuestion GenerationQuestion-GenerationVisual Question Answering+1

Towards Retrieval Augmented Generation over Large Video Libraries

2024-06-21 · Yannis Tevissen, Khalil Guetari, Frédéric Petitpont

Video content creators need efficient tools to repurpose content, a task that often requires complex manual or automated searches. Crafting a new video from large video libraries remains a challenge. In this paper we int…

Answer GenerationQuestion AnsweringRAGRetrieval+1

Solution for SMART-101 Challenge of ICCV Multi-modal Algorithmic Reasoning Task 2023

2023-10-10 · Xiangyu Wu, Yang Yang, Shengdong Xu, Yifeng Wu 외

In this paper, we present our solution to a Multi-modal Algorithmic Reasoning Task: SMART-101 Challenge. Different from the traditional visual question-answering datasets, this challenge evaluates the abstraction, deduct…

Decoderobject-detectionObject DetectionOptical Character Recognition (OCR)+2