paper-with-me

홈 › Papers

A Novel Framework for Robustness Analysis of Visual QA Models

2017-11-16 · Jia-Hong Huang, Cuong Duc Dao, Modar Alfadly, Bernard Ghanem

Deep neural networks have been playing an essential role in many computer vision tasks including Visual Question Answering (VQA). Until recently, the study of their accuracy was the main focus of research but now there is a trend toward assessing the robustness of these models against adversarial attacks by evaluating their tolerance to varying noise levels. In VQA, adversarial attacks can target the image and/or the proposed main question and yet there is a lack of proper analysis of the later. In this work, we propose a flexible framework that focuses on the language part of VQA that uses semantically relevant questions, dubbed basic questions, acting as controllable noise to evaluate the robustness of VQA models. We hypothesize that the level of noise is positively correlated to the similarity of a basic question to the main question. Hence, to apply noise on any given main question, we rank a pool of basic questions based on their similarity by casting this ranking task as a LASSO optimization problem. Then, we propose a novel robustness measure, R_score, and two large-scale basic question datasets (BQDs) in order to standardize robustness analysis for VQA models.

📄 PDF Abstract BibTeX arXiv:1711.06232

Code (0)

등록된 구현이 없습니다.

Tasks

Question AnsweringVisual Question AnsweringVisual Question Answering (VQA)

Similar Papers 제목 키워드 기반

V$^2$R-Bench: Holistically Evaluating LVLM Robustness to Fundamental Visual Variations

2025-04-23 · Zhiyuan Fan, Yumeng Wang, Sandeep Polisetty, Yi R. Fung

Large Vision Language Models (LVLMs) excel in various vision-language tasks. Yet, their robustness to visual variations in position, scale, orientation, and context that objects in natural scenes inevitably exhibit due t…

Dataset GenerationObject RecognitionPosition

Evaluating Robustness of Visual Representations for Object Assembly Task Requiring Spatio-Geometrical Reasoning

2023-10-15 · Chahyon Ku, Carl Winge, Ryan Diaz, Wentao Yuan 외

This paper primarily focuses on evaluating and benchmarking the robustness of visual representations in the context of object assembly tasks. Specifically, it investigates the alignment and insertion of objects with geom…

BenchmarkingSpatial Reasoning

A Robust Point Cloud Analysis Framework Inspired By Primary Visual Cortex

2026-06-12 · Jisheng Dang, Dengyue Pan, Delin Deng, Yifan Zhang 외 arxiv

Despite significant advancements in point cloud analysis, reducing energy consumption and improving robustness remain understudied, largely due to the inherent limitations of Convolutional Neural Networks (CNNs). To addr…

Visual CoT Makes VLMs Smarter but More Fragile

2025-09-28 · Chunxue Xu, Yiwei Wang, Yujun Cai, Bryan Hooi 외 arxiv

Chain-of-Thought (CoT) techniques have significantly enhanced reasoning in Vision-Language Models (VLMs). Extending this paradigm, Visual CoT integrates explicit visual edits, such as cropping or annotating regions of in…

Visual Question Answering

Can Text-to-image Model Assist Multi-modal Learning for Visual Recognition with Visual Modality Missing?

2024-02-14 · Tiantian Feng, Daniel Yang, Digbalay Bose, Shrikanth Narayanan

Multi-modal learning has emerged as an increasingly promising avenue in vision recognition, driving innovations across diverse domains ranging from media and education to healthcare and transportation. Despite its succes…