paper-with-me

홈 › Papers

$\text{R}^2$-Bench: Benchmarking the Robustness of Referring Perception Models under Perturbations

2024-03-07 · Xiang Li, Kai Qiu, Jinglu Wang, Xiaohao Xu, Rita Singh, Kashu Yamazak, Hao Chen, Xiaonan Huang, Bhiksha Raj

Referring perception, which aims at grounding visual objects with multimodal referring guidance, is essential for bridging the gap between humans, who provide instructions, and the environment where intelligent systems perceive. Despite progress in this field, the robustness of referring perception models (RPMs) against disruptive perturbations is not well explored. This work thoroughly assesses the resilience of RPMs against various perturbations in both general and specific contexts. Recognizing the complex nature of referring perception tasks, we present a comprehensive taxonomy of perturbations, and then develop a versatile toolbox for synthesizing and evaluating the effects of composite disturbances. Employing this toolbox, we construct $\text{R}^2$-Bench, a benchmark for assessing the Robustness of Referring perception models under noisy conditions across five key tasks. Moreover, we propose the $\text{R}^2$-Agent, an LLM-based agent that simplifies and automates model evaluation via natural language instructions. Our investigation uncovers the vulnerabilities of current RPMs to various perturbations and provides tools for assessing model robustness, potentially promoting the safe and resilient integration of intelligent systems into complex real-world scenarios.

📄 PDF Abstract BibTeX arXiv:2403.04924

Code (2)

lxa9867/r2bench 공식 구현 pytorch
lxa9867/R2VOS pytorch

Tasks

Benchmarking

Similar Papers 제목 키워드 기반

CD-RMOT-Bench: Benchmarking the Cross-Domain Referring Multi-Object Tracking

2026-07-28 · Xiangqun Zhang, Likai Wang, Zekun Qian, Ruize Han 외 arxiv

Referring multi-object tracking (RMOT) extends tracking from category-driven perception to language-guided understanding by grounding object trajectories in natural-language expressions. Despite recent progress, existing…

Multi-Object TrackingObject Detection

An Open-Source Benchmark and Baseline for Multi-temporal Referring Segmentation

2026-05-31 · Bingyu Li, Da Zhang, Tao Huo, Zhiyuan Zhao 외 arxiv

Large Vision-Language Models (LVLMs) have shown strong visual understanding and language-guided grounding abilities, yet their capacity for multi-temporal visual reasoning remains underexplored. To bridge this gap, we in…

Change DetectionVisual Reasoning

RefHCM: A Unified Model for Referring Perceptions in Human-Centric Scenarios

2024-12-19 · Jie Huang, Ruibing Hou, Jiahe Zhao, Hong Chang 외

Human-centric perceptions play a crucial role in real-world applications. While recent human-centric works have achieved impressive progress, these efforts are often constrained to the visual domain and lack interaction …

Transfer Learning

RefBench-PRO: Perceptual and Reasoning Oriented Benchmark for Referring Expression Comprehension

2025-12-06 · Tianyi Gao, Hao Li, Han Fang, Xin Wei 외 arxiv

Referring Expression Comprehension (REC) is a vision-language task that localizes a specific image region based on a textual description. Existing REC benchmarks primarily evaluate perceptual capabilities and lack interp…

Referring Expression

Multimodal Referring Segmentation: A Survey

2025-08-01 · Henghui Ding, Song Tang, Shuting He, Chang Liu 외 arxiv

Multimodal referring segmentation aims to segment target objects in visual scenes, such as images, videos, and 3D scenes, based on referring expressions in text or audio format. This task plays a crucial role in practica…

Referring Expression