paper-with-me

홈 › Papers

ChiQA: A Large Scale Image-based Real-World Question Answering Dataset for Multi-Modal Understanding

2022-08-05 · Bingning Wang, Feiyang Lv, Ting Yao, Yiming Yuan, Jin Ma, Yu Luo, Haijin Liang

Visual question answering is an important task in both natural language and vision understanding. However, in most of the public visual question answering datasets such as VQA, CLEVR, the questions are human generated that specific to the given image, such as `What color are her eyes?'. The human generated crowdsourcing questions are relatively simple and sometimes have the bias toward certain entities or attributes. In this paper, we introduce a new question answering dataset based on image-ChiQA. It contains the real-world queries issued by internet users, combined with several related open-domain images. The system should determine whether the image could answer the question or not. Different from previous VQA datasets, the questions are real-world image-independent queries that are more various and unbiased. Compared with previous image-retrieval or image-caption datasets, the ChiQA not only measures the relatedness but also measures the answerability, which demands more fine-grained vision and language reasoning. ChiQA contains more than 40K questions and more than 200K question-images pairs. A three-level 2/1/0 label is assigned to each pair indicating perfect answer, partially answer and irrelevant. Data analysis shows ChiQA requires a deep understanding of both language and vision, including grounding, comparisons, and reading. We evaluate several state-of-the-art visual-language models such as ALBEF, demonstrating that there is still a large room for improvements on ChiQA.

📄 PDF Abstract BibTeX arXiv:2208.03030

Code (1)

benywon/ChiQA 공식 구현 pytorch

Tasks

Image RetrievalQuestion AnsweringRetrievalVisual Question AnsweringVisual Question Answering (VQA)

Methods 이 논문이 사용한 방법론

ALBEF ALBEF introduces a contrastive loss to align the image and text representations before fusing them through cross-modal attention. This enables more grounded vision and language…

Similar Papers 제목 키워드 기반

Exploring Rich Subjective Quality Information for Image Quality Assessment in the Wild

2024-09-09 · Xiongkuo Min, Yixuan Gao, Yuqin Cao, Guangtao Zhai 외

Traditional in the wild image quality assessment (IQA) models are generally trained with the quality labels of mean opinion score (MOS), while missing the rich subjective quality information contained in the quality rati…

Image Quality Assessment

RealRestorer: Towards Generalizable Real-World Image Restoration with Large-Scale Image Editing Models

2026-03-26 · Yufeng Yang, Xianfang Zeng, Zhangqi Jiang, Fukun Yin 외 arxiv

Image restoration under real-world degradations is critical for downstream tasks such as autonomous driving and object detection. However, existing restoration models are often limited by the scale and distribution of th…

Autonomous DrivingImage RestorationObject DetectionImage Editing

AirBirds: A Large-scale Challenging Dataset for Bird Strike Prevention in Real-world Airports

2023-04-23 · Hongyu Sun, Yongcai Wang, Xudong Cai, Peng Wang 외

One fundamental limitation to the research of bird strike prevention is the lack of a large-scale dataset taken directly from real-world airports. Existing relevant datasets are either small in size or not dedicated for …

Time Series

Flow-Anything: Learning Real-World Optical Flow Estimation from Large-Scale Single-view Images

2025-06-09 · Yingping Liang, Ying Fu, Yutao Hu, Wenqi Shao 외

Optical flow estimation is a crucial subfield of computer vision, serving as a foundation for video tasks. However, the real-world robustness is limited by animated synthetic datasets for training. This introduces domain…

Depth EstimationMonocular Depth EstimationOptical Flow Estimation

Boosting Zero-shot Stereo Matching using Large-scale Mixed Images Sources in the Real World

2025-05-13 · Yuran Wang, Yingping Liang, Ying Fu

Stereo matching methods rely on dense pixel-wise ground truth labels, which are laborious to obtain, especially for real-world datasets. The scarcity of labeled data and domain gaps between synthetic and real-world image…

Depth EstimationMonocular Depth EstimationStereo Matching